<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Inside Riya's Mind]]></title><description><![CDATA[Inside Riya's Mind]]></description><link>https://riyajaiswal25.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Wed, 30 Sep 2026 06:38:36 GMT</lastBuildDate><atom:link href="https://riyajaiswal25.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Getting started at data augmentation]]></title><description><![CDATA[Data augmentation is the process in which we enlarge the dataset using existing dataset.
It is basically set of techniques which are applied on the dataset, it might be minor or major changes, to improve model accuracy and give best results.
Classic ...]]></description><link>https://riyajaiswal25.hashnode.dev/getting-started-at-data-augmentation</link><guid isPermaLink="true">https://riyajaiswal25.hashnode.dev/getting-started-at-data-augmentation</guid><category><![CDATA[TensorFlow]]></category><category><![CDATA[Deep Learning]]></category><dc:creator><![CDATA[Riya Jaiswal]]></dc:creator><pubDate>Sat, 15 Jul 2023 13:08:07 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1689424923186/a81a646e-d7ef-430c-8629-9a406e88bf6f.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Data augmentation is the process in which we enlarge the dataset using existing dataset.</p>
<p>It is basically set of techniques which are applied on the dataset, it might be minor or major changes, to improve model accuracy and give best results.</p>
<p>Classic image processing activities for data augmentation are:</p>
<ul>
<li><p>padding</p>
</li>
<li><p>random rotating</p>
</li>
<li><p>re-scaling,</p>
</li>
<li><p>vertical and horizontal flipping</p>
</li>
<li><p>translation ( image is moved along X, Y direction)</p>
</li>
<li><p>cropping</p>
</li>
<li><p>zooming</p>
</li>
<li><p>darkening &amp; brightening/color modification</p>
</li>
<li><p>grayscaling</p>
</li>
<li><p>changing contrast</p>
</li>
<li><p>adding <a target="_blank" href="https://en.wikipedia.org/wiki/Image_noise">noise</a></p>
</li>
<li><p>random erasing</p>
<p>  These activities can be performed on the dataset and it can produce several new datasets upon performing several transformations mentioned above and so improving the chances of accurate predictions.</p>
</li>
</ul>
<p><img src="https://research.aimultiple.com/wp-content/uploads/2021/04/dataaugmention_image_alletranitons.png" alt="Seven examples of image augmentation: rotation, blur, contrast, scaling, illumination &amp; projective." /></p>
<p><em>Source of image(medium)</em></p>
<p>The above image illustrates the various transformations that can be applied on a single image, thus enlarging and generating such images which contribute to better predictions and thus imporoving model's performance.</p>
<p>Benefits of data augmentation include:</p>
<ul>
<li><p>Improving model prediction accuracy</p>
<ul>
<li><p>adding more training data into the models</p>
</li>
<li><p>preventing data scarcity for better models</p>
</li>
<li><p>reducing data overfitting ( i.e. an error in statistics, it means a function corresponds too closely to a limited set of data points) and creating variability in data</p>
</li>
<li><p>increasing generalization ability of the models</p>
</li>
<li><p>helping resolve class imbalance issues in classification</p>
</li>
</ul>
</li>
<li><p>Reducing costs of collecting and labeling data</p>
</li>
<li><p>Enables rare event prediction</p>
</li>
<li><p>Prevents data privacy problems</p>
</li>
</ul>
<p>I personally created a model, one without data augmentation, which gave the accuracy of 68% and after applying data augmentation on the same dataset(Flipping, Rotation and Zoom) the accuracy improved to 73% with just three transformations.</p>
<p>It also eliminated the overfitting of model and enhanced its accuracy drastically. The above article only specified data augmentation on images, this also can be performed on text and audio.</p>
]]></content:encoded></item><item><title><![CDATA[Understanding Cluster Analysis]]></title><description><![CDATA[Basically, cluster analysis is the distribution of data points into various clusters. The criteria of distributing them into various clusters is such that there must be similarity between datapoints in same cluster and difference in data points betwe...]]></description><link>https://riyajaiswal25.hashnode.dev/understanding-cluster-analysis</link><guid isPermaLink="true">https://riyajaiswal25.hashnode.dev/understanding-cluster-analysis</guid><category><![CDATA[Data Mining]]></category><category><![CDATA[clustering]]></category><category><![CDATA[clusters]]></category><dc:creator><![CDATA[Riya Jaiswal]]></dc:creator><pubDate>Fri, 14 Jul 2023 14:45:02 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1689318943611/2f5e4494-e56a-4973-9822-1ff396b76d15.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Basically, cluster analysis is the distribution of data points into various clusters. The criteria of distributing them into various clusters is such that there must be similarity between datapoints in same cluster and difference in data points between different clusters.</p>
<p>The more the similarity between data points between same cluster and the more difference in betweeen data points of different clusters, the more distinct is the clustering.</p>
<p>Clustering: Identifying objects of similar types</p>
<p>Types of clustering:</p>
<ol>
<li><p><strong><mark>Partitional Clustering:</mark></strong> Dataset is divided into clusters, i.e., set of groups, separate k value can be taken for centroid based method. We need to pre-specify the no. of clusters.</p>
</li>
<li><p><strong><mark>Hierarchial Clustering:</mark></strong> Set of nested clustering, organized by representation of a tree. Usually visualized by a dendrogram.</p>
</li>
<li><p><strong><mark>Exclusive Clustering:</mark></strong> Assigning each object to a single group, i.e, non-overlapping clustering.</p>
</li>
<li><p><strong><mark>Non-exclusive Clustering:</mark></strong> Objects in one group can also be present in other groups, i.e, objects can simultaneously belong to one or more than one group.</p>
<p> It is overlapping clustering.</p>
</li>
<li><p><strong><mark>Fuzzy Clustering:</mark></strong> Based on membership weight concept.</p>
<p> Every object should have a minimum weight between 0 and 1.</p>
<p> Clusters are treated as fuzzy sets.</p>
</li>
<li><p><strong><mark>Complete Clustering:</mark></strong> Assign every object to a cluster, no object should be left.</p>
<p> Every object is desired.</p>
</li>
<li><p><strong><mark>Partial Clustering:</mark></strong> Some objects does not have groups or are not clustered properly. It does not have desired object.</p>
</li>
<li><p><strong><mark>Well-separated Clustering: </mark></strong> Threshold used to specify that all the objects in a cluster must be sufficiently close (or similar) to one another.</p>
<p> The distance between any two points in different groups is larger than the distance between any two points within a group.</p>
</li>
<li><p><strong><mark>Prototype based Clustering:</mark></strong> For data with <strong>continuous attributes</strong>, the <strong>prototype</strong> of a cluster is often a <strong>centroid</strong>, i.e., the <strong>average (mean)</strong> of all the points in the cluster. </p>
<p> <strong>Prototype</strong> is a <strong>medoid</strong> (the most representative point of a cluster) for <strong>categorical attributes</strong>.</p>
<p> For many types of data, the <strong>prototype</strong> is the <strong>most central point</strong>, and commonly referred as <strong>center-based clusters</strong>. Such clusters are globular.</p>
</li>
<li><p><strong><mark>Graph Based Clustering:</mark></strong> If the <strong>data</strong> is represented as a <strong>graph,</strong> where the <strong>nodes are objects</strong> and the <strong>links represent connections among objects</strong> then a <strong>cluster</strong> can be defined as a <strong>connected component</strong>; i.e., a group of objects that are connected to one another, but that have no connection to objects outside the group.</p>
</li>
<li><p><strong><mark>Density based:</mark></strong> A cluster is a dense region of objects that is surrounded by a region of low density. A density-based definition of a cluster employed when the clusters are irregular or intertwined, and when noise and outliers are present. E.g., DBSCAN Algorithm</p>
</li>
</ol>
]]></content:encoded></item></channel></rss>