Predictive Modelling Archives - DexLab Analytics | Credit Risk | Market Risk | SAS Python Machine Learning Modeling

Credit Risk in Indian Banking: RBI Data Analysis

Credit Risk in Indian Banking: What RBI’s Data Actually Shows

Every risk professional in Indian banking eventually asks the same question: is credit risk actually improving, or does it just look that way in aggregate numbers? Based on the Reserve Bank of India’s own published data, the answer is both. System-wide asset quality has genuinely strengthened over the past five years. But the composition of that risk is shifting in a direction that deserves closer attention.

This piece works through RBI’s Financial Stability Reports (FSR), sectoral credit data, and the newly finalized Expected Credit Loss (ECL) framework. Together they show what’s really happening with credit risk in Indian banking between 2020 and 2025. It does not rely on a proprietary survey or projected estimates dressed up as findings. In fact, every figure below is sourced directly to a named RBI report. That distinction matters: in a domain where regulators, auditors, and rating agencies check your numbers, credibility is the entire product.

Three questions structure the analysis. How has aggregate asset quality moved since the 2020 pandemic shock? Where is risk concentrating today, even as headline numbers improve? And what does the incoming ECL regime signal about how Indian banks will need to manage credit risk going forward?

Methodology and Data Sources

This analysis draws on primary RBI publications, cross-checked across multiple reporting periods for consistency.

RBI Financial Stability Reports (FSR): Published twice yearly. They consolidate Gross NPA (GNPA) and Net NPA (NNPA) ratios of Scheduled Commercial Banks (SCBs), capital adequacy (CRAR), bank-group-wise asset quality, and stress test results. Specifically, this piece uses FSR editions from January 2021 through December 2025.

RBI Sectoral Deployment of Bank Credit data: Monthly data on credit growth to industry, services, agriculture, and personal loans. It’s sourced from 41 banks representing roughly 95% of non-food credit.

RBI’s ECL framework releases: The draft ECL directions (October 2025) and final directions (April 2026). These describe the shift from incurred-loss to forward-looking PD/LGD/EAD-based provisioning, effective April 1, 2027.

One scope note: RBI does not publish a standardized “default rate by loan product” table. Where this article cites loan-category or bank-group figures, they are GNPA ratios: the share of gross advances classified as non-performing. It’s the metric RBI itself uses, and the one directly comparable across periods.

Finding 1: Asset Quality Has Improved for Five Consecutive Years

SCB GNPA: from 8% to 2.1% in five years

 

PeriodGNPA RatioNNPA RatioSource
March 20208.4%RBI FSR, Jan 2021
September 20207.5%RBI FSR, Jan 2021
March 20242.8%0.6%RBI FSR, Jun 2024
March 20252.3%0.5%RBI FSR, Jun 2025
September 20252.1%–2.2%RBI FSR, Dec 2025
March 2027 (projected, baseline)1.9%RBI FSR, Dec 2025

 

RBI’s January 2021 report recorded a September 2020 GNPA ratio of 7.5%, down from 8.4% in March 2020. That was a system still absorbing the pandemic shock. GNPA had fallen to 2.8% by June 2024, then to 2.3% by March 2025. It touched a multi-decade low of 2.1% by September 2025, and RBI projects further improvement to 1.9% by March 2027 under its baseline scenario.

In practice, this reflects five years of balance sheet cleanup: post-IBC resolution of legacy corporate stress, tighter underwriting after the 2018–2020 NBFC stress episode, and stronger capital buffers overall. Meanwhile, system-wide CRAR remains comfortably above regulatory minimums, with public sector banks at 16% and private banks at 18.1% as of September 2025.

In short, aggregate GNPA is a lagging confirmation of underwriting discipline, not a leading indicator. A PD model trained mainly on 2020–2022 stressed data will overstate current default risk. One trained only on 2023–2025 benign data risks understating tail risk in the next downturn.

Explore our Credit Risk Modeling Certification Training for a structured approach to PD estimation across credit cycles.

Finding 2: Improvement Isn’t Even Across Bank Groups

PSBs are catching up fast

For instance, PSB GNPA fell sharply from 3.7% in March 2024 to 2.8% in March 2025. Meanwhile, private bank GNPA held roughly stable at 2.8% over the same period, and foreign banks improved from 1.2% to 0.9%.

Even so, this convergence matters. For most of the post-2015 asset-quality-review era, PSB asset quality lagged private banks significantly, largely on corporate exposures. That gap has now nearly closed at the aggregate level. However, remaining risk differs by bank group, which leads to the more consequential finding below.

Finding 3: Unsecured Retail Is Where New Risk Concentrates

The retail risk hiding inside a good headline number

This is the most important finding for practitioners, because it sits underneath the reassuring headline number. According to RBI’s December 2025 FSR, roughly 53.1% of retail loan slippages now originate from unsecured products like personal loans and credit cards. At private banks, unsecured loans account for nearly 76% of fresh slippages. GNPA on unsecured retail loans stood at 1.8%, versus 1.1% for overall retail advances.

In other words, the 2.1% aggregate GNPA figure blends a very clean secured/corporate book with a smaller, faster-deteriorating unsecured retail book. RBI flagged this as a fintech-adjacent risk, tied to fast credit growth in small-ticket personal loans to borrowers under 35 through digital lending channels.

This pattern, in fact, tracks with operational experience. Unsecured lending has weaker recovery mechanics (no collateral to liquidate, higher LGD), shorter behavioral history on new-to-credit borrowers, and faster origination cycles that compress underwriting review. Moreover, it is the segment where forward-looking provisioning matters most, since unsecured risk builds up quietly between formal NPA recognition points.

As a result, portfolio-level GNPA alone is no longer sufficient. Overall, segment-level GNPA and vintage curves for unsecured retail belong alongside the aggregate number in any board-level risk dashboard.

Finding 4: ECL Will Formalize This Shift

Why the 2027 ECL shift matters here

RBI has issued directions introducing forward-looking ECL provisioning, replacing the incurred-loss model. It takes effect April 1, 2027, for scheduled commercial banks excluding RRBs, Small Finance Banks, and payments banks. ECL provisioning must be based on a bank’s own historical PD and LGD data spanning at least five years, subject to RBI-specified floors. Accounts 30–90 days past due move into Stage 2, a materially earlier trigger than the current framework.

Overall, the shift aligns India’s prudential norms with global IFRS 9 standards. In addition, it requires closer integration between finance and risk functions, as forward-looking macroeconomic scenarios become a formal input to provisioning.

Indeed, this is a direct regulatory response to Finding 3. An incurred-loss model recognizes impairment only after default has effectively occurred. ECL requires estimating expected loss, via PD, LGD, and EAD, well before that point, catching unsecured deterioration earlier in the cycle.

Even so, for banks building this capability, it isn’t a compliance task to fully outsource. RBI has explicitly made a bank’s board and senior management responsible for the adequacy of the ECL framework. Consequently, internal teams need working fluency in PD/LGD/EAD construction, not just the ability to read vendor output. However, it’s worth noting that the standard formula, Expected Loss = PD × LGD × EAD, assumes independence between the three components. In practice they’re correlated: LGD tends to rise in the same downturns that push PD higher. That’s why RBI’s stress tests apply adverse scenarios jointly rather than multiplying baseline figures in isolation.

What This Means for Banks and Risk Teams

  • First, aggregate GNPA improvement is real but incomplete. Segment-level monitoring, especially for unsecured retail, deserves as much attention as the headline ratio.
  • PD/LGD model recency matters. RBI’s own five-year minimum spans both a stressed period (2020–2021) and a benign one (2023–2025). Models need to represent both.
  • Collateral still matters, but isn’t the whole story. Unsecured products drive a disproportionate share of new slippages. In turn, this argues for tighter underwriting in that segment, not a wholesale retreat from unsecured lending.
  • Finally, the 2027 ECL deadline is closer than it looks. In practice, building five years of clean PD/LGD data and validation capability is a multi-year undertaking. Banks starting in 2026 are already behind institutions that began in 2024–2025.
  • Recovery rate discipline matters for LGD. LGD = 1 − Recovery Rate only holds up when ‘recovery rate’ is the economic, discounted, net-of-cost rate, not the nominal amount eventually collected.

Explore our Credit Risk Modeling Certification Training to build PD, LGD, and EAD modeling skills ahead of the 2027 ECL transition, or see Understanding Credit Risk: Definition and Types for foundational concepts referenced throughout.

FAQ

What is the current GNPA ratio of Indian banks?

As of September 2025, SCB GNPA stood at 2.1%, a multi-decade low, per RBI’s December 2025 Financial Stability Report.

Is unsecured lending riskier than secured lending right now?

Yes, and the gap is widening. Unsecured retail GNPA was 1.8% versus 1.1% for overall retail advances, and unsecured products drove over half of all retail slippages.

When does RBI’s ECL framework take effect?

RBI’s ECL Directions were issued 27 April 2026 and take effect April 1, 2027. They apply to commercial banks, excluding small finance banks, payments banks, and local area banks.

Does EL = PD × LGD × EAD fully capture expected loss?

It’s the standard starting formula, but it assumes PD, LGD, and EAD move independently. In stress, they’re correlated — which is why RBI applies adverse scenarios jointly rather than multiplying baseline values.

Conclusion

The data supports a measured conclusion, not a triumphant one. Indeed, Indian banking’s asset quality genuinely improved for five straight years, and RBI’s own numbers back that up without embellishment. However, the same data shows risk isn’t disappearing. Instead, it’s relocating toward unsecured retail lending, addressed through a regulatory shift that will demand more rigorous PD, LGD, and EAD modeling capability than most institutions currently have in-house. For risk analysts, credit officers, and model validators, that combination is telling: improving headline numbers alongside a harder compliance mandate. It’s exactly why 2025–2027 is a build-capability window, not a wait-and-see one.

This analysis is based on RBI’s Financial Stability Reports, Sectoral Deployment of Bank Credit data, and RBI’s ECL Directions (2025–2026). Figures are reported as published at the cited dates; readers should consult original RBI releases for the most current data.

 

Ready to Build These Skills Hands-On?

Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.

Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.

 


.

Introduction to MongoDB

MongoDB is a document based database program which was developed by MongoDB Inc. and is licensed under server side public license (SSPL). It can be used across platforms and is a non-relational database also known as NoSQL, where NoSQL means that the data is not stored in the conventional tabular format and is used for unstructured data as compared to SQL and that is the major difference between NoSQL and SQL.
MongoDB stores document in JSON or BSON format. JSON also known as JavaScript Object notation is a format where data is stored in a key value pair or array format which is readable for a normal human being whereas BSON is nothing but the JSON file encoded in the binary format which is quite hard for a human being to understand.
Structure of MongoDB which uses a query language MQL(Mongodb query language):-
Databases:- Databases is a group of collections.
Collections:- Collection is a group fields.
Fields:- Fields are nothing but key value pairs
Just for an example look at the image given below:-

Here I am using MongoDB Compass a tool to connect to Atlas which is a cloud based platform which can help us write our queries and start performing all sort of data extraction and deployment techniques. You can download MongoDB Compass via the given link https://www.mongodb.com/try/download/compass

In the above image in the red box we have our databases and if we click on the “sample_training” database we will see a list of collections similar to the tables in sql.

Now lets write our first query and see what data in “companies” collection looks like but before that select the “companies” collection.

Now in our filter cell we can write the following query:-

In the above query “name” and “category_code” are the key values also known as fields and “Wetpaint” and “web” are the pair values on the basis of which we want to filter the data.
What is cluster and how to create it on Atlas?
MongoDB cluster also know as sharded cluster is created where each collection is divided into shards (small portions of the original data) which is a replica set of the original collection. In case you want to use Atlas there is an unpaid version available with approximately 512 mb space which is free to use. There is a pre-existing cluster in MongoDB named Sandbox , which currently I am using and you can use it too by following the given steps:-
1. Create a free account or sign in using your Google account on
https://www.mongodb.com/cloud/atlas/lp/try2-in?utm_source=google&utm_campaign=gs_apac_india_search_brand_atlas_desktop&utm_term=mongodb%20atlas&utm_medium=cpc_paid_search&utm_ad=e&utm_ad_campaign_id=6501677905&gclid=CjwKCAiAr6-ABhAfEiwADO4sfaMDS6YRyBKaciG97RoCgBimOEq9jU2E5N4Jc4ErkuJXYcVpPd47-xoCkL8QAvD_BwE
2. Click on “Create an Organization”.
3. Write the organization name “MDBU”.
4. Click on “Create Organization”.
5. Click on “New Project”.
6. Name your project M001 and click “Next”.
7. Click on “Build a Cluster”.
8. Click on “Create a Cluster” an option under which free is written.
9. Click on the region closest to you and at the bottom change the name of the cluster to “Sandbox”.
10. Now click on connect and click on “Allow access from anywhere”.
11. Create a Database User and then click on “Create Database User”.
username: m001-student
password: m001-mongodb-basics
12. Click on “Close” and now load your sample as given below :

Loading may take a while….
13. Click on collections once the sample is loaded and now you can start using the filter option in a similar way as in MongoDB Compass
In my next blog I’ll be sharing with you how to connect Atlas with MongoDB Compass and we will also learn few ways in which we can write query using MQL.

So, with that we come to the end of the discussion on the MongoDB. Hopefully it helped you understand the topic, for more information you can also watch the video tutorial attached down this blog. The blog is designed and prepared by Niharika Rai, Analytics Consultant, DexLab Analytics DexLab Analytics offers machine learning courses in Gurgaon. To keep on learning more, follow DexLab Analytics blog.


.

Time Series Analysis Part I

 

A time series is a sequence of numerical data in which each item is associated with a particular instant in time. Many sets of data appear as time series: a monthly sequence of the quantity of goods shipped from a factory, a weekly series of the number of road accidents, daily rainfall amounts, hourly observations made on the yield of a chemical process, and so on. Examples of time series abound in such fields as economics, business, engineering, the natural sciences (especially geophysics and meteorology), and the social sciences.

  • Univariate time series analysis- When we have a single sequence of data observed over time then it is called univariate time series analysis.
  • Multivariate time series analysis – When we have several sets of data for the same sequence of time periods to observe then it is called multivariate time series analysis.

The data used in time series analysis is a random variable (Yt) where t is denoted as time and such a collection of random variables ordered in time is called random or stochastic process.

Stationary: A time series is said to be stationary when all the moments of its probability distribution i.e. mean, variance , covariance etc. are invariant over time. It becomes quite easy forecast data in this kind of situation as the hidden patterns are recognizable which make predictions easy.

Non-stationary: A non-stationary time series will have a time varying mean or time varying variance or both, which makes it impossible to generalize the time series over other time periods.

Non stationary processes can further be explained with the help of a term called Random walk models. This term or theory usually is used in stock market which assumes that stock prices are independent of each other over time. Now there are two types of random walks:
Random walk with drift : When the observation that is to be predicted at a time ‘t’ is equal to last period’s value plus a constant or a drift (α) and the residual term (ε). It can be written as
Yt= α + Yt-1 + εt
The equation shows that Yt drifts upwards or downwards depending upon α being positive or negative and the mean and the variance also increases over time.
Random walk without drift: The random walk without a drift model observes that the values to be predicted at time ‘t’ is equal to last past period’s value plus a random shock.
Yt= Yt-1 + εt
Consider that the effect in one unit shock then the process started at some time 0 with a value of Y0
When t=1
Y1= Y0 + ε1
When t=2
Y2= Y1+ ε2= Y0 + ε1+ ε2
In general,
Yt= Y0+∑ εt
In this case as t increases the variance increases indefinitely whereas the mean value of Y is equal to its initial or starting value. Therefore the random walk model without drift is a non-stationary process.

So, with that we come to the end of the discussion on the Time Series. Hopefully it helped you understand time Series, for more information you can also watch the video tutorial attached down this blog. DexLab Analytics offers machine learning courses in delhi. To keep on learning more, follow DexLab Analytics blog.


.

Top 5 Industry Use Cases of Predictive Analytics

Top 5 Industry Use Cases of Predictive Analytics

Predictive analytics is an effective in-hand tool crafted for data scientists. Thanks to its quick computing and on-point forecasting abilities! Not only data scientists, but also insurance claim analysts, retail managers and healthcare professionals enjoy the perks of predictive analytics modeling – want to know how?

Below, we’ve enumerated a few real-life use cases, existing across industries, threaded with the power of data science and predictive analytics. Ask us, if you have any queries for your next data science project! Our data science courses in Delhi might be of some help.

Customer Retention

Losing customers is awful. For businesses. They have to gain new customers to make up for the loss in revenue. But, it can cost more, winning new customers is usually hailed more costly than retaining older ones.

Predictive analytics is the answer. It can prevent reduction in the customer base. How? By foretelling you the signs of customer dissatisfaction and identifying the customers that are most likely to leave. In this way, you would know how to keep your customers satisfied and content, and control revenue slip offs.

Customer Lifetime Value

Marketing a product is the crux of the matter. Identifying customers willing to spend a large part of their money, consistently for a long period of time is difficult to find. But once cracked, it helps companies optimize their marketing efforts and enhance their customer lifetime value.

2

Quality Control

Quality Control is significant. Over time, shoddy quality control measures will affect customer satisfaction ratio, purchasing behavior, thus impacting revenue generation and market share.

Further, low quality control results in more customer support expenses, repairs and warranty challenges and less systematic manufacturing. Predictive analytics help provide insights on potential quality issues, before they turn into crucial company growth hindrances.  

Risk Modeling

Risk can originate from a plethora of source, and it can be any form. Predictive analytics can address critical aspects of risk – it collects a huge number of data points from many organizations and sort through them to determine the potential areas of concern.

What’s more, the trends in the data hint towards unfavorable circumstances that might impact businesses and bottom line in an adverse way. A concoction of these analytics and a sound risk management approach is what companies truly need to quantify the risk challenges and devise a perfect course of action that’s indeed the need of the hour.

Sentiment Analysis

It’s impossible to be everywhere, especially when being online. Similarly, it’s very difficult to oversee everything that’s said about your company.

Nevertheless, if you amalgamate web search and a few crawling tools with customer feedback and posts, you’d be able to develop analytics that’d present you an overview of the organization’s reputation along with its key market demographics and more. Recommendation system helps!

All hail Predictive Analytics! Now, maneuver beyond fuss-free reactive operations and let predictive analytics help you plan for a successful future, evaluating newer areas of business scopes and capabilities.

Interested in data science certification? Look up to the experts at DexLab Analytics.

The blog has been sourced fromxmpro.com/10-predictive-analytics-use-cases-by-industry

Interested in a career in Data Analyst?

To learn more about Data Analyst with Advanced excel course – Enrol Now.
To learn more about Data Analyst with R Course – Enrol Now.
To learn more about Big Data Course – Enrol Now.

To learn more about Machine Learning Using Python and Spark – Enrol Now.
To learn more about Data Analyst with SAS Course – Enrol Now.
To learn more about Data Analyst with Apache Spark Course – Enrol Now.
To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.

Predictive Analytics: The Key to Enhance the Process of Debt Collection

Predictive Analytics: The Key to Enhance the Process of Debt Collection

A wide array of industries has already engaged in some kind of predictive analytics – numerical analysis of debt collection is relatively a recent addition. Financial analysts are now found harnessing the power of predictive analytics to cull better results out for their clients, and measure the effectiveness of their strategies and collections.

Let’s see how predictive analytics is used in debt collection process:

2

Understanding Client Scoring (Risk Assessment)

Since the late 1980’s, FICO score is regarded as the golden standard for determining creditworthiness and loan application. But, however, machine learning, particularly predictive analytics can replace it, and develop an encompassing portrait of a client, taking into effect more than his mere credit history and present debts. It can also include his social media feeds and spending trajectory.

Evaluating Payment Patterns

The survival models evaluate each client’s probability of becoming a potential loss. If the account shows a continuous downward trend, then it should be regarded soon as a potential risk. Predictive analytics can help identify spending patterns, indicating the struggles of each client. A system can be developed which self-triggers whenever any unwanted pattern transpires. It could ask the client if they need any help or if they are going through a financial distress, so that it can help before the situation turns beyond repairs.

For R predictive modeling training courses, visit DexLab Analytics.

Cash Flow Predictions

Businesses are keen to know about future cash flows – what they can expect! Financial institutions are no different. Predictive analytics helps in making more appropriate predictions, especially when it comes to receivables.

Debt collector’s business models are subject to the ability to forecast the success of collection operations, and ascertaining results at the end of each month, before the billing cycle initiates. As a result, the workforce of the company is able to shift their focus from the potential payers to those who would not be able to meet their obligations. This shift in focus helps!

Better Client Relationship

Predictive analytics weave wonders; not only it has the ability to point which clients are the highest risks for your company, but also predict the best time to contact them to reap maximum results. What you need to do is just visit the logs of past conversations.

Challenges

Last, but not the least, all big data models face a common challenge – data cleaning. As it’s a process of wastage in and out, before starting with prediction, company should deal with this problem at first to construct a pipeline, for feeding in the data, clean it and use it for neural network training.

In a concluding statement, predictive analytics is the best bet for debt and revenue collection – it boosts conversion rates at the right time with the right people. If you want to study more about predictive analytics, and its varying uses in different segments of industry, enroll in R Predictive Modelling Certification training at DexLab Analytics. They provide superior knowledge-intensive training to interested individuals with added benefit of placement assistance. For more, visit their website.

 

The blog has been sourced fromdataconomy.com/2018/09/improving-debt-collection-with-predictive-models

 

Interested in a career in Data Analyst?

To learn more about Data Analyst with Advanced excel course – Enrol Now.
To learn more about Data Analyst with R Course – Enrol Now.
To learn more about Big Data Course – Enrol Now.

To learn more about Machine Learning Using Python and Spark – Enrol Now.
To learn more about Data Analyst with SAS Course – Enrol Now.
To learn more about Data Analyst with Apache Spark Course – Enrol Now.
To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.

Bringing Back Science into “Data Science”

Bringing Back Science into “Data Science”

Far from the conventional science disciplines, like physics or mathematics, Data Science is a budding discipline: which means there are no proper definition to explain what data science is and what role it does play.

Nevertheless, the internet is full of working definitions of data science. As per Wikipedia, Data Science is

(an) interdisciplinary field about processes and systems to extract knowledge or insights from data in various forms, either structured or unstructured, which is a continuation of some of the data analysis fields such as statistics, data mining, and predictive analytics.

To that note, a very important aspect is left behind in this explanation: Data Science is a science first, which means a proper scientific method should be devised to tackle different data science practices. By scientific method, we mean a healthy process of asking questions, collecting information, framing hypothesis and analyzing the results to draw conclusions thereafter.

Go below, the process breakup is as follows..

Ask questions

Start by asking what is the business problem? How to leverage maximum gains? What ways to implement to increase return on investment? The finance industry takes help from data science for myriad reasons. One of the most striking reasons is to enhance the return on investment out of marketing campaigns.

What Sets Apart Data Science from Big Data and Data Analytics – @Dexlabanalytics.

Collect data

A predictive modeling analyst has access to vast data resources, which eventually makes the entire research and gathering data process much less complex. However, it is only in theory, because rarely data is stored in the desired format an analyst wants, making his job easier.

Data Science – then and now! – @Dexlabanalytics.

Devise a hypothesis

After getting to the heart and soul of the problem, we start to develop hypotheses. For example, you believe your firm’s profit is leveraged by an optimistic customer reaction towards your product quality and positive advertising capabilities of your firm. Through this example, we explained a nomological network, where you are in a position to infer casualties and correlations. While dealing in Data Science, assessing customer perception is very crucial, and so is the analysis of financial datasets.

Data Science: Is It the Right Answer? – @Dexlabanalytics.

Testing and experiments

Formulating a hypothesis is not enough; a predictive modeler relies on statistical modeling techniques to forecast the future in a probabilistic manner. Keep a note, this doesn’t result in indicating “X will occur”, instead it refers “Given Y, the probability of X occurring is 75%.”

Any proper experiment includes control groups and test, meaning a modeler when preparing a predictive model should divide the dataset so as to ensure availability of few data for testing predictive equation.

Now, if we talk about marketing – consider logistic regression. It offers a probability whether a binary event of interest will take place or not.

Enroll in an R Predictive Modelling Certification program to go through the mechanics of this problem. Reach us at DexLab Analytics.

Tracing Success in the New Age of Data Science – @Dexlabanalytics.

Evaluate results and infer conclusions

Now is the time to make a decision: do you prefer the quantitative approach? As social media is totally unstructured, the qualitative approach needs to be implemented using Natural Language Processing, which can be a tad difficult. Now, how about making a longitudinal analysis, while transforming data into time series? Do all these questions rake your mind? Yes? Then you are on the right track.

Keep Pace with Automation: Emerging Data Science Jobs in India – @Dexlabanalytics.

Reporting of results

This is the final battle scene for all predictive modelers. It calls for all the documents, based on which a modeler made his decision during the development process. All the assumptions taken have to be identified and highlighted beside the results.

And with it comes the end of our Science in Data Science process!

For more interesting updates and blogs, follow us at DexLab Analytics. Opt for our impressive Data Science Courses in gurgaon and lead the road of success!

 

Interested in a career in Data Analyst?

To learn more about Data Analyst with Advanced excel course – Enrol Now.
To learn more about Data Analyst with R Course – Enrol Now.
To learn more about Big Data Course – Enrol Now.

To learn more about Machine Learning Using Python and Spark – Enrol Now.
To learn more about Data Analyst with SAS Course – Enrol Now.
To learn more about Data Analyst with Apache Spark Course – Enrol Now.
To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.

How Predictive Analysis Works With Data Mining

We know that you have probably heard many times that predictive analysis will further optimize and accentuate your marketing campaigns. But it is hard to envision that in more concrete terms what it will achieve. This makes it harder to choose and direct analytics technology.

 

How Predictive Analysis Works With Data Mining

 

Wondering how you can get a functional value for marketing, sales and product directions without being an expert? The solution to all your problems lies in how predictive analytics may offer with benefits for the current marketing operations. But to use it you must learn a few specifics about how it works.

Continue reading “How Predictive Analysis Works With Data Mining”

Making a Histogram With Basic R Programming Skills

DexLab Analytics over the course of next few weeks will cover the basics of various data analysis techniques like creating your own histogram in R programming. We will explore three options for this: R commands, ggplot2 and ggvis. These posts are for users of R programming who are in the beginner or intermediate level and who require accessible and easy to understand resources.

 

Making a Histogram With Basic R Programming Skills

 

Seeking more information? Then take up our R language training course in Gurgaon from DexLab Analytics.

 

What is a histogram?

A histogram is a category of visual representation of a dataset distribution. As such the shape of a histogram is its most common feature for identification. With a histogram one will be able to see which factor has the relatively higher amount of data and which factors or segments have the least.

 

Or put in simpler terms, one can see where the middle or median is in a data distribution, and how close or farther away the data would lie around the middle and where would the possible outliers be found. And precisely because of all this histograms will be the best way to understand your data.

 

But what can a specific shape of a histogram tell us? In short a typical histogram consists of an x-axis and a y-axis and a few bars of varying heights. The y-axis will exhibit how frequently the values on the x-axis are occurring in the data. The y-axis showcases the frequency of the values on the x-axis where the data occurs, the bar group ranges of either values or continuous categories on the x-axis. And the latter explains why the histograms do not have any gaps between the bars.

 

Let’s Take Your Data Dreams to the Next Level

How can one make a histogram with basic R?

Step 1: Get your eyes on the data:

As histograms require some amount of data to be plotted initially, you can carry that out by importing a dataset or simply using one which is built into the system of R. In this tutorial we will make use of 2 datasets the built-in R dataset AirPassengers and another dataset called as chol, which is stored into a .txt file and is available for download.

Step 2: Acquaint yourself with The Hist () function:

One can make a histogram in R by opting the easy way where they use The Hist () function, which automatically computes a histogram of the given data values. One would put the name of their dataset in between parentheses to use this function.

Here is how to use the function:

hist(AirPassengers)

 

But if in case, you want to select a certain column of a data frame like for instance in chol, for making a histogram. The hist function should be used with the dataset name in combination with a $ symbol, which should be followed by the column name:

 

2

Here is a specimen showing the same:

hist(chol$AGE) #computes a histogram of the data values in the column AGE of the dataframe named “chol”

Step 3: Up the level of the hist () function:

You may find that the histograms created with the previous features seem a little dull. That is because the default visualizations do not contribute much to the understanding of the histograms. One may need to take one more step to reach a better and easier understanding of their histograms. Fortunately, this is not too difficult to accomplish, R has several allowances for easy and fast ways to optimize the visualizations of the diagrams while still making use of the hist () function.

To adapt your histogram you will only need to add more arguments to the hist () function, in this way:

hist(AirPassengers,
     main="Histogram for Air Passengers",
     xlab="Passengers",
     border="blue",
     col="green",
     xlim=c(100,700),
     las=1,
     breaks=5)

This code will help to compute a histogram of data values from the dataset AirPassengers, with the name “Histogram for Air Passengers” as the title. The x-axis would be labelled as ‘Passengers’ and will have a blue border with a green colour to the bins, while limiting the x-axis with a range of 100 to 700 and rotating the printed values on the y-axis by 1 while changing  the bin width by 5.

We know what you are thinking – this is a humungous string of code. But do not worry, let us break it down into smaller pieces to see what each component holds. 

Name/colours:

You can alter the title of the histogram by adding main as an argument to the hist () function.

This is how:

hist(AirPassengers, main=”Histogram for Air Passengers”) #Histogram of the AirPassengers dataset with title “Histogram for Air Passengers”

For adjusting the label of the x-axis you can add xlab as the feature. Similarly one can also use ylab to label the y-axis.

This code would work:

hist(AirPassengers, xlab=”Passengers”, ylab=”Frequency of Passengers”) #Histogram of the AirPassengers dataset with changed labels on the x-and y-axes hist(AirPassengers, xlab=”Passengers”, ylab=”Frequency of Passengers”) #Histogram of the AirPassengers dataset with changed labels on the x-and y-axes

If in case you would want to change the colours of the default histogram you can simply choose to add the arguments border or col. Adjusting would be easy, as the name itself kind of gives away the borders and the colours of the histogram.

hist(AirPassengers, border=”blue”, col=”green”) #Histogram of the AirPassengers dataset with blue-border bins with green filling

Note: you must not forget to put the names and the colours within “ ”.

For x and y axes:

To change the range of the x and y axes one can use the xlim and the ylim as arguments to the hist function ():

The code to be used is:

hist(AirPassengers, xlim=c(100,700), ylim=c(0,30)) #Histogram of the AirPassengers dataset with the x-axis limited to values 100 to 700 and the y-axis limited to values 0 to 30

Point to be noted in this case, is the c() function is used for delimiting the values on the axes when one is suing the xlim and ylim functions. It takes 2 values the first being the begin value and the second being the end value.

Make sure to rotate the labels on the y-axis by adding 1as=1 as the argument, the argument 1as can be 0, 1, 2 or 3.

The code to be used:

hist(AirPassengers, las=1) #Histogram of the AirPassengers dataset with the y-values projected horizontally

 

Depending on the option one chooses the placement of the label will vary: like for instance, if you choose 0 the label will always be parallel to the axis (the one that is the default). And if one chooses 1, The label will be horizontally put. If you want the label to be perpendicular to the axis then pick 2 and for placing it vertically select 3.

For bins:

One can alter the bin width by including breaks as an argument, in combination with the number of breakpoints which one wants to have.

This is the code to be used:

hist(AirPassengers, breaks=5) #Histogram of the AirPassengers dataset with 5 breakpoints

If one wants to have increased control over the breakpoints in between the bins, then they can enrich the breaks arguments by adding in it vector of breakpoints, one can also do this by making use of the c() function.

hist(AirPassengers, breaks=c(100, 300, 500, 700)) #Compute a histogram for the data values in AirPassengers, and set the bins such that they run from 100 to 300, 300 to 500 and 500 to 700.

But the c () function can help to make your code very messy at times, which is why we recommend using add = seq(x,y,z) instead. The values of x, y and z are determined by the user and represented in a specific order of appearance, the starting number of x-axis  and the last number of the same as well as the intervals in which these numbers are to appear.

A noteworthy point to be mentioned here is that one can combine both the functions:

hist(AirPassengers, breaks=c(100, seq(200,700, 150))) #Make a histogram for the AirPassengers dataset, start at 100 on the x-axis, and from values 200 to 700, make the bins 150 wide

Here is the histogram of AirPassengers:

Here is the histogram of AirPassengers:
How to Make a Histogram with Basic R – (Image Courtesy  r-bloggers)

Please note that this is the first blog tranche in a list of 3 posts on creating histograms using R programming.

For more information regarding R language training and other interesting news and articles follow our regular uploads at all our channels. 

This post originally appeared onwww.r-bloggers.com/how-to-make-a-histogram-with-basic-r

Interested in a career in Data Analyst?

To learn more about Machine Learning Using Python and Spark – click here.

To learn more about Data Analyst with Advanced excel course – click here.
To learn more about Data Analyst with SAS Course – click here.
To learn more about Data Analyst with R Course – click here.
To learn more about Big Data Course – click here.

We Are Training Snapdeal on Data Science with R

With the Big Data boom within the IT industry worldwide, more and more online retailers are using it to create better shopping experience for their customers through a boost in customer satisfaction to generate better revenue for themselves.

 

We Are Training Snapdeal on Data Science with R
Dexlab Analytics is Conducting Training for Snapdeal in Data Science and R Programming

 

The funny news about Target knowing about a young lady’s pregnancy even before the father could was a viral content that sent the internet crazy. But how did they know this?

 
The answer lies in the wizardry of data analysis, as when a lady starts searching to buy products like nutritional supplements, unscented beauty products and cotton balls then there is a good chance that she is pregnant.

 

For More Information Visit Now www.prlog.org at Dexlab Analytics is Conducting Training for Snapdeal in Data Science and R Programming

Continue reading “We Are Training Snapdeal on Data Science with R”

Call us to know more