Why is loss given default important? Because LGD determines the actual financial loss a bank absorbs when a borrower defaults. It feeds directly into capital requirements under Basel III, provisioning under IFRS 9 and RBI’s ECL framework, and loan pricing decisions. As a result, getting LGD wrong means a bank either holds too little capital against real losses, or misprices credit risk across its portfolio.
What Loss Given Default Measures in Credit Risk
LGD measures the share of an exposure a lender doesn’t recover after a borrower defaults. In other words, it’s the loss severity component of credit risk. That makes it distinct from Probability of Default (PD), which measures how likely a default is in the first place.
Together with Exposure at Default (EAD), these three components combine into the standard credit loss formula:
Expected Loss = PD × LGD × EAD
This formula is really the starting point for understanding why is loss given default important in practice. PD tells a bank how often defaults happen. LGD, on the other hand, tells the bank how much each one costs. So a portfolio with a low default rate but high LGD can generate the same expected loss as a portfolio with a high default rate but low LGD. That’s exactly why risk teams can’t focus on PD alone and call the job done. For a deeper walkthrough, see our companion guide on LGD recovery rates and estimation.
Why Loss Given Default Is Important for Basel III Capital Requirements
Basel III uses LGD as a direct input into risk-weighted assets (RWA). This, in turn, determines how much regulatory capital a bank must hold against a given exposure. Under the Foundation IRB approach, for instance, supervisors prescribe standard LGD values: 45% for senior unsecured corporate exposures, and 75% for subordinated claims. Under the Advanced IRB approach, however, banks estimate their own LGD from internal workout data, subject to regulatory floors.
Here’s why LGD is important practically: a higher LGD assumption produces a higher risk weight, which in turn produces a higher capital charge. Consequently, a bank that under-models LGD on a large unsecured portfolio ends up holding less capital than the actual loss potential warrants. This is exactly the gap regulators scrutinize during model validation and supervisory review.
Why LGD Is Important for IFRS 9 and RBI ECL Provisioning
IFRS 9 requires banks to provision for expected credit losses using the same PD × LGD × EAD structure. This structure applies across Stage 1, 2, and 3 assets, depending on how much credit quality has deteriorated. As a result, LGD estimates need to reflect current conditions and forward-looking scenarios, not just historical averages.
In India, the Reserve Bank of India’s Expected Credit Loss (ECL) framework, effective April 1, 2027, mirrors this approach. It shifts Indian banks away from incurred-loss provisioning toward a forward-looking model, built on internally estimated PD, LGD, and EAD. Under both frameworks, therefore, an inaccurate LGD estimate translates directly into an inaccurate provision. This either overstates earnings through under-provisioning, or unnecessarily constrains lending capacity through over-provisioning.
Why Loss Given Default Is Important in Loan Pricing Decisions
LGD shapes loan pricing long before any default happens. A lender pricing a loan factors in expected loss, and since LGD drives half of that calculation, it directly affects the credit spread a borrower pays.
This is why collateral matters so much in commercial lending. For example, a borrower offering strong, liquid collateral lowers the bank’s expected LGD, which typically earns a lower interest rate. A borrower with weak or no collateral, on the other hand, pushes LGD higher, and the pricing reflects that. Ultimately, risk-based pricing frameworks that ignore LGD, or rely on rough approximations, systematically mis-price credit — undercharging high-severity borrowers and overcharging low-severity ones.
Why LGD Is Harder to Get Right Than PD
PD modeling benefits from decades of relatively clean default data and a statistical distribution that behaves well. LGD, however, doesn’t have either advantage.
Recovery outcomes depend on collateral type, legal jurisdiction, the efficiency of the workout process, and the state of the economy when the default happens. Because of this, recovery data takes years to mature; workouts on defaulted loans can run for years before resolution. In addition, LGD outcomes tend to cluster near the extremes — many loans get almost fully recovered, while others get almost fully lost. This pattern breaks the assumptions behind standard linear regression.
This is exactly why LGD estimation, not PD estimation, tends to generate the most disagreement between banks, auditors, and regulators during model reviews.
Real-World Example: Why LGD Mattered in the Lehman Brothers Collapse
Lehman Brothers’ 2008 collapse shows how badly recovery expectations can move. It also shows why LGD estimates need real stress-scenario grounding, not just long-run averages.
On the day Lehman filed for bankruptcy in September 2008, market prices on its senior bonds implied a recovery rate of roughly 30% for senior creditors. A month later, however, that implied recovery rate had collapsed to around 9%, matching the outcome of Lehman’s credit default swap auction. Anyone modeling LGD off pre-crisis assumptions, therefore, would have badly underestimated the loss severity in that window.
The story didn’t end there. Roughly two and a half years after the filing, Lehman’s estate projected eventual creditor recovery at just 16%. Over the following decade, though, as the estate liquidated assets and resolved legal claims, actual recoveries for senior unsecured creditors climbed well above that early projection, reaching close to 40% by the time of the final distributions — according to research published by the Federal Reserve Bank of New York’s Liberty Street Economics.
That range — a market-implied 9% at the depth of the crisis, versus roughly 40% a decade later — illustrates exactly why Basel’s Advanced IRB framework requires downturn LGD estimation. A model calibrated only on calm-market recovery assumptions would have missed the loss severity precisely when it mattered most, in the weeks immediately following the default.
Frequently Asked Questions
Q1. What is the difference between LGD and PD?
PD measures the likelihood that a borrower defaults. LGD, by contrast, measures how much a lender loses if that default happens. Both feed into the Expected Loss formula, but they capture different dimensions of credit risk.
Q2. Why is loss given default important for bank capital requirements?
Basel III uses LGD as a direct input into risk-weighted assets. So a higher LGD assumption increases the capital charge on an exposure, meaning inaccurate LGD estimates lead directly to mis-calibrated capital requirements.
Q3. Does LGD matter for retail loans, or only corporate lending?
It matters for both. Retail mortgages typically carry lower LGD because of strong collateral coverage. Unsecured retail products like personal loans and credit cards, however, carry higher LGD. Basel and IFRS 9 frameworks require LGD estimation across both segments.
Q4. How often should banks update their LGD estimates?
Regulatory expectations under Basel and IFRS 9/RBI’s ECL framework call for periodic recalibration. This is especially true after economic downturns or shifts in collateral markets, rather than relying on estimates built during a single benign period.
Conclusion: Why Loss Given Default Is Important Going Forward
Why is loss given default important, in the end? Because it converts the probability of a default into an actual number a bank has to plan for: in capital held, provisions booked, and prices charged. Basel III, IFRS 9, and RBI’s ECL framework all treat it as a core input, not a secondary adjustment. Understanding LGD in depth, therefore, including how to estimate it properly and where standard modeling approaches fall short, is foundational for anyone working in credit risk.
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
Loss Given Default (LGD) Modeling and Recovery Rates: A Practical Estimation Guide
Introduction
Loss Given Default, or LGD, decides what a default actually costs a bank. Probability of Default (PD) tells a lender how likely a default is. LGD tells them what it will cost when it happens. LGD is often the harder number to pin down. PD stays fairly stable across a cycle. Recovery outcomes don’t — they swing with collateral quality, legal jurisdiction, workout efficiency, and the state of the economy at the time of default.
This matters directly for capital and provisioning. Both the Internal Ratings-Based (IRB) approach to regulatory capital and the IFRS 9 Expected Credit Loss (ECL) framework use LGD as a direct multiplier: Expected Loss = PD × LGD × EAD. Underestimate LGD by even a few percentage points on a large secured portfolio, and a bank can materially understate its capital requirement or provisioning charge. Risk practitioners often find that LGD, not PD, generates the biggest disagreements between banks, auditors, and regulators. Recovery data is sparser. Workout periods run longer. And the underlying distribution of outcomes looks nothing like the neat bell curve that PD models assume.
Risk analysts, credit modelers, and finance professionals preparing for FRM, PRM, or model validation roles need to understand LGD well. That means knowing how recovery rates behave across secured and unsecured exposures, how collateral gets valued under stress, which statistical techniques actually fit LGD’s unusual distribution, and what regulators — from the Basel Committee to the Reserve Bank of India — expect from a defensible LGD framework. This guide walks through each building block with the technical depth a working risk professional needs.
What Is Loss Given Default (LGD)?
Definition. Loss Given Default is the share of an exposure a lender expects to lose if a borrower defaults. The calculation nets out recoveries from collateral liquidation, guarantees, insurance, and the workout process, then subtracts the costs of collecting. Analysts express LGD as a percentage of Exposure at Default (EAD).
The foundational relationship is simple:
LGD = 1 − Recovery Rate
The Recovery Rate is the present value of everything a lender ultimately recovers, divided by the exposure at the point of default. If a bank recovers 65% of a defaulted loan’s value after liquidating collateral and pursuing legal remedies, LGD works out to 35%.
Exposure at Default (EAD): The outstanding balance — principal, accrued interest, and undrawn commitments where applicable — at the moment of default.
Recoveries: Cash or asset value a bank collects through collateral sale, guarantee invocation, litigation settlement, or restructuring, often over a multi-year workout period.
Discounting: Recoveries arrive over time, not instantly, so analysts discount them back to the default date. The discount rate typically reflects the risk of the recovery cash flows — often the original effective interest rate or a risk-adjusted rate set by internal policy.
Workout costs: Legal fees, collateral maintenance costs, recovery agent fees, and administrative overhead all reduce net recovery and push LGD higher.
Why this matters for capital and provisioning: LGD isn’t a static number sitting in a spreadsheet. It flows directly into risk-weighted asset calculations under Basel and into Stage 1, 2, and 3 provisioning under IFRS 9. A retail mortgage with strong collateral coverage might carry an LGD of 10–20%. An unsecured personal loan or trade finance exposure might sit at 45–75%. That gap — often three to five times higher for unsecured versus well-secured exposures — shows why collateral structure, not just borrower creditworthiness, drives so much of a bank’s capital planning.
A simple example: A bank holds a ₹10 crore secured corporate loan. The borrower defaults. The bank spends ₹40 lakh liquidating the pledged collateral and litigating a personal guarantee, and recovers ₹7 crore in present value terms after 18 months. Net recovery works out to ₹7 crore minus ₹0.40 crore, or ₹6.6 crore. That’s a 66% recovery rate — and a 34% LGD.
Recovery Rate Analysis: Secured vs Unsecured Exposures
Recovery rates are the mirror image of LGD. The single biggest driver of variation in recovery outcomes is whether — and how well — an exposure is secured.
Secured exposures
When collateral backs a loan, the lender holds a claim on a specific asset it can liquidate or repossess. Recovery rates on secured exposures run generally higher and less volatile, but they still vary enormously by collateral type:
Real estate (commercial and residential): Usually the most stable collateral class. Recovery rates often land in the 60–85% range, depending on property type, location, and market liquidity at the time of sale. Residential mortgages in developed markets tend to sit at the higher end; commercial real estate swings more with the cycle.
Financial collateral (cash, government securities, listed equities): Produces the highest and most predictable recoveries, often near 90–100% after haircuts, because valuation and liquidation happen fast and transparently.
Receivables and inventory: More volatile. Recovery depends heavily on the quality of the underlying receivables book and how quickly a bank can sell inventory without a distressed-sale discount.
Plant, machinery, and other physical collateral: Recovery rates run lower and more dispersed, since specialized equipment often has a thin secondary market.
Unsecured exposures
With no specific asset pledged, the lender’s claim ranks alongside other unsecured creditors in insolvency proceedings. Recovery depends on the borrower’s overall asset pool, the seniority of the claim, and the jurisdiction’s insolvency regime. Unsecured corporate exposures commonly show recovery rates of 20–40%. That’s precisely why Basel’s standard supervisory LGD for senior unsecured corporate exposures sits at 40–45% under the Foundation IRB approach, and why subordinated unsecured claims carry an even higher LGD of 75% — implying a recovery rate assumption of just 25%.
Guarantees
A third category deserves a mention: guaranteed exposures. A credible third-party guarantee — sovereign, bank, or corporate — can push recovery closer to the guarantor’s own credit strength rather than the original borrower’s. But this substitution effect only holds if the guarantee is legally robust and the guarantor itself stays out of distress. Many banks relearned this lesson during systemic downturns, when correlated defaults hit both borrower and guarantor at the same time.
The practical takeaway: recovery rate is never a single number for a portfolio. It’s a distribution shaped by collateral type, seniority, jurisdiction, and workout timing. Any LGD model that ignores this segmentation will systematically misestimate risk on the tails.
Collateral Valuation Methods for LGD Estimation
Collateral quality drives most of recovery, which makes collateral valuation one of the most consequential — and most frequently under-scrutinized — inputs into LGD estimation. Banks use three broad approaches, often in combination.
1. Market Approach
Analysts derive value from observed transaction prices for comparable assets — recent sales of similar properties, equipment, or securities. This method works best when an active, liquid market exists, such as listed securities or standard residential property in a liquid city market, because it reflects what the collateral would actually fetch if sold today. The limitation: during systemic stress, comparable transactions become scarce or reflect distressed-sale pricing themselves. That can understate true economic value, or overstate it if valuations lag a falling market.
2. Income Approach
For income-generating collateral such as commercial real estate or business assets, this method values collateral based on the present value of expected future cash flows — rental income or operating cash flows — discounted at an appropriate capitalization rate. It looks further forward than the market approach, but it’s sensitive to assumptions about occupancy, rental growth, and discount rates. All of those need independent validation to avoid overly optimistic collateral values.
3. Adjusted / Haircut Values
Regulatory and internal risk frameworks rarely accept raw market or income valuations at face value. Instead, they apply haircuts — percentage reductions that account for valuation uncertainty, price volatility, currency mismatch between the loan and collateral, and the time and cost required to liquidate. Basel’s standardised and IRB frameworks specify minimum haircuts by collateral class: small haircuts of a few percentage points for financial collateral, and larger haircuts — often 15–40% — for real estate and other physical collateral, reflecting realization risk. Banks then build internal haircut schedules on top of these regulatory floors, calibrated to their own historical liquidation experience.
Practical valuation governance issues risk teams should watch for
Stale valuations: collateral appraised at origination but never revalued through the credit cycle
Correlated collateral risk: collateral value and borrower default probability moving together — for example, a real estate developer whose loan is secured by real estate, where collateral value falls exactly when default risk rises
Valuation model independence: appraisers or valuation models controlled by the same business unit that originated the loan, which creates conflict-of-interest risk
Robust LGD models need a documented, periodically revalidated collateral valuation policy, not a one-time appraisal frozen at origination. Collateral value at the point of default — not at origination — determines actual recovery.
LGD Calculation Approaches: Market, Workout, and Statistical Methods
Banks use several distinct methodologies to derive LGD estimates, often blending them depending on data availability and asset class.
1. Market LGD
This approach uses observed market prices of defaulted debt instruments shortly after default — common for traded corporate bonds and syndicated loans. If a bond trades at 35 cents on the dollar right after default, the implied market LGD comes out to roughly 65%. The method is fast and market-based, but it only works where a liquid secondary market for distressed debt exists. That rules it out for most retail and SME lending.
2. Workout LGD
Banks use this approach most often for loan portfolios, especially retail and commercial lending without a tradable market. It tracks the actual cash flows recovered over the full workout period following default — collateral liquidation proceeds, guarantee payments, restructuring recoveries, litigation settlements — discounts them back to the default date, and nets off workout costs. Workout LGD demands a lot of data: it needs long, clean histories of defaulted accounts tracked to final resolution, which can take years. But it produces the most economically grounded estimate, because it reflects a bank’s own actual recovery experience.
3. Implied Market / Historical LGD
Analysts derive this indirectly from observed credit spreads or historical loss rates on a portfolio, back-solving for an implied LGD given known PD and observed loss experience. It works as a cross-check, or for asset classes with limited direct workout data, but it’s less precise than a direct workout LGD study.
4. Statistical / Regression-Based LGD Models
Once a bank assembles a workout LGD dataset — actual LGD outcomes for a sample of resolved defaults — it typically builds a statistical model to predict LGD for the broader portfolio. Explanatory variables include collateral type and loan-to-value ratio, industry, exposure size, time to resolution, macroeconomic conditions at default, and borrower or facility characteristics. This is where the modeling choice becomes critical, because LGD’s statistical properties are unusual and standard linear regression handles them poorly — the next section explains why.
Segmentation matters as much as method. Whichever base approach a bank uses, best practice segments LGD estimation by product type, collateral class, and — where sample size allows — geography or industry, rather than fitting one LGD model across a heterogeneous portfolio. A single blended LGD for “all secured corporate loans” hides enormous variation. A loan secured by liquid financial collateral behaves very differently from one secured by specialized machinery.
Beta Regression for LGD Modeling
Why Ordinary Least Squares (OLS) regression fails for LGD. LGD data has three statistical properties that violate the core assumptions of OLS:
Bounded support: LGD sits between 0 and 1 (it can occasionally exceed 1 in cases of negative recovery, where workout costs surpass recoveries, though banks typically cap this). OLS assumes an unbounded continuous response, so it will happily predict LGD values below 0% or above 100% for extreme covariate combinations — numbers that make no economic sense.
Non-normal, often bimodal distribution: LGD outcomes frequently cluster near the two extremes: many defaulted loans get fully recovered (LGD near 0), while others are almost entirely lost (LGD near 1), with fewer observations in between. This U-shaped or bimodal pattern breaks OLS’s assumption of normally distributed residuals.
Heteroscedasticity: The variance of LGD outcomes isn’t constant across the range of predicted values — it typically differs by collateral type and loan segment. That further violates OLS assumptions and produces unreliable standard errors and confidence intervals.
Why Beta regression fits better. The Beta distribution lives naturally on the (0, 1) interval and can take on a wide range of shapes — including the U-shaped, bimodal pattern common in LGD data — by adjusting its two shape parameters. A Beta regression model links the mean of this Beta-distributed response to explanatory variables through a logit-type link function. That guarantees predicted LGD values stay within the valid 0–1 range, something OLS can’t promise.
Practical advantages for LGD modeling
Predictions stay automatically bounded, which removes the need for ad hoc capping or flooring rules on OLS outputs
The model can separately parameterize the mean and the dispersion of LGD, so analysts can model not just expected LGD but how much it varies by segment — useful for stress testing and downturn LGD estimation
Tail behavior calibrates better, which matters because LGD’s tails — very low or very high loss — are exactly where capital and provisioning are most sensitive
A common workaround worth knowing: LGD data often includes exact zeros and ones, which creates a practical problem since pure Beta distributions require values strictly between 0 and 1. Analysts typically handle this with a small transformation (the Smithson-Verkuilen adjustment, for example) or a zero-and-one-inflated Beta regression, which models the probability mass at the boundaries separately from the continuous distribution in between. This usually fits better than forcing all observations into the open interval.
Analysts also use other techniques alongside or instead of Beta regression: Tobit models for censored LGD data, fractional response regression, and machine learning approaches such as gradient boosting or random forests. The machine learning methods can capture non-linear interactions between collateral, borrower, and macroeconomic variables, but they demand more careful validation and explainability work before regulators will accept them.
Basel III and RBI LGD Standards
Basel Foundation IRB (F-IRB) supervisory LGD values
Under Basel III, banks using the Foundation IRB approach don’t estimate their own LGD for corporate exposures — they apply supervisor-prescribed values instead. Senior unsecured claims on corporates, sovereigns, and banks get a standard LGD of 45%, recalibrated to 40% for senior unsecured exposures specifically to non-financial corporates under the Basel III finalisation reforms. All subordinated claims carry a 75% LGD. For secured exposures, a prescribed formula adjusts LGD downward using collateral-specific haircuts, subject to input floors that stop banks from modeling implausibly low LGDs even with strong collateral coverage.
Advanced IRB (A-IRB)
Banks approved for the Advanced IRB approach can use their own internal LGD estimates, built from workout data as described earlier. These estimates still face regulatory floors, and they must reflect downturn conditions — meaning LGD estimates need to capture the possibility that recoveries fall during an economic downturn, when collateral values drop and workout costs rise, not just long-run average experience. This “downturn LGD” requirement ranks among the more technically demanding parts of IRB model validation, since a bank needs historical data spanning at least one full economic cycle, including a stress period, to calibrate it credibly.
RBI’s regulatory position
India has taken a notably different path from many global peers. Rather than adopt the IRB approach for credit risk capital, the Reserve Bank of India issued the Commercial Banks – Capital Charge for Credit Risk (Standardised Approach) Directions, 2026, mandating the Standardised Approach for all commercial banks under its jurisdiction — excluding small finance banks, payments banks, and local area banks — effective April 1, 2027. No Indian bank currently operates under IRB for credit risk capital purposes.
LGD becomes central for Indian banks instead through the RBI’s Expected Credit Loss (ECL) framework, finalized in April 2026 and also effective April 1, 2027. It requires scheduled commercial banks to move from the current incurred-loss provisioning model to a forward-looking ECL approach built on internally estimated PD, LGD, and EAD — closely mirroring IFRS 9. Because IFRS 9 and RBI’s ECL directions stay principle-based rather than prescriptive on LGD methodology, RBI has added product-wise prudential floors for Stage 1 and Stage 2 provisioning as a regulatory backstop, along with specific expectations around collateral valuation techniques for LGD and the distressed value of collateral for retail exposures. Individual banks carry the burden of building defensible, well-governed internal LGD models — the exact skill set this guide covers — even though India hasn’t adopted IRB for capital purposes.
Case Study: Recovery Rates in Practice
A widely studied example from the 2008–2009 global financial crisis shows why segmentation and downturn calibration matter so much in LGD modeling. Post-crisis studies of European and US bank loan portfolios found that recovery rates on defaulted bank loans and bonds fell sharply during the crisis compared to long-run historical averages. Unsecured and subordinated corporate exposures took the biggest hit. Recoveries partially rebounded in later years as distressed-asset markets stabilized and collateral prices recovered.
Secured exposures with liquid financial or real estate collateral held up better than unsecured corporate exposures over the same period. Even real-estate-backed recoveries took a hit, though, dragged down by the simultaneous collapse in property prices — a textbook case of the collateral correlation risk covered earlier, where the value of the security and the borrower’s default risk deteriorate together.
Risk teams draw a consistent lesson from this and similar downturn episodes: an LGD model calibrated purely on a benign credit cycle will understate loss severity exactly when it matters most. That’s why Basel’s Advanced IRB framework mandates downturn LGD estimation rather than allowing long-run average LGD alone. It’s also why banks building internal LGD models — for IRB or for IFRS 9/ECL purposes — need workout data spanning a genuine stress period, not just a benign multi-year average, before regulators or internal capital committees will accept the model.
Frequently Asked Questions
Q1. What is the difference between LGD and recovery rate?
Recovery rate measures the share of an exposure a lender gets back after default. LGD is simply 1 minus the recovery rate. They’re two sides of the same calculation — recovery rate measures what a bank regains, LGD measures what it loses.
Q2. Why can’t banks just use OLS regression to model LGD?
LGD sits between 0 and 1, often clusters near full recovery or full loss, and shows different variance across segments — all of which break OLS’s core assumptions. Beta regression and related bounded-response models handle these properties directly and keep predictions within the valid range.
Q3. What is downturn LGD, and why does it matter?
Downturn LGD reflects economic stress conditions — depressed collateral values and elevated workout costs — rather than long-run average experience. Basel’s Advanced IRB framework requires it because average-cycle LGD alone can understate capital needs exactly when losses are most likely to spike.
Q4. Does India use the IRB approach for LGD under Basel?
No. The RBI has mandated the Standardised Approach for credit risk capital for all commercial banks under its jurisdiction, effective April 1, 2027, rather than adopting Internal Ratings-Based approaches. LGD modeling still matters for Indian banks under RBI’s parallel Expected Credit Loss (ECL) provisioning framework, which also takes effect April 1, 2027.
Q5. How does collateral type affect LGD?
Collateral type drives recovery outcomes more than almost any other factor. Financial collateral — cash, government securities — typically produces the highest, most stable recoveries. Real estate stays generally strong but cyclical. Receivables, inventory, and specialized physical assets tend to show lower and more volatile recovery rates, and unsecured exposures show the lowest recoveries of all.
Conclusion
LGD estimation sits at the intersection of statistics, legal recovery process, and collateral economics, which is exactly why it resists simple modeling shortcuts. A defensible LGD framework needs clean segmentation by collateral and seniority, a disciplined collateral valuation methodology, a modeling technique that respects LGD’s bounded and skewed distribution, and calibration that accounts for downturn conditions rather than relying solely on benign-cycle averages. Whether you’re building models under Basel’s Advanced IRB approach or an IFRS 9/RBI-style ECL framework, the fundamentals covered here — from workout LGD construction to Beta regression — form the technical backbone of credible LGD estimation.
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
Every loan a bank writes carries a question that has to be answered with a number. How likely is this borrower to stop paying?
PD estimation is the discipline that produces that number. It underpins regulatory capital, loan loss provisions, pricing, approval cut-offs and portfolio strategy. Consequently, an error in either direction is expensive. Overstate risk and the bank prices itself out of good business. Understate it and losses arrive before provisions do.
The number is also more visible than it used to be. Under Basel, PD fed a capital calculation reviewed quarterly. Under IFRS 9 and CECL, it feeds provisions that flow into reported earnings every reporting period. As a result, PD estimation has moved from a specialist modelling concern to something auditors, boards and analysts examine directly.
This guide covers the main estimation methods, how retail and corporate portfolios differ, what forward-looking PD requires, and how to validate the result.
PD estimation is the process of quantifying the probability that a borrower will fail to meet contractual obligations over a defined time horizon. The result is expressed between 0% and 100%. It is one of three inputs to expected loss, alongside Loss Given Default (LGD) and Exposure at Default (EAD).
The expected loss identity is straightforward:
Expected Loss = PD × LGD × EAD
Three parameters must be fixed before a PD figure means anything.
The default definition
Basel and most supervisory regimes trigger default at 90 days past due. Alternatively, a lender may declare default earlier if it judges the obligor unlikely to pay without recourse to collateral. Importantly, a model built on a 90-DPD definition is not comparable to one built on 30-DPD. Mixing definitions across portfolios is a common source of inconsistency.
The time horizon
Regulatory capital under the Internal Ratings-Based approach uses a 12-month PD. IFRS 9 and CECL, by contrast, require a lifetime PD for exposures that have deteriorated since origination. These are different quantities. A 12-month PD of 2% does not imply a five-year PD of 10%, because default hazard varies with loan age.
Point-in-time or through-the-cycle
A point-in-time (PIT) PD reflects current conditions and moves with the economic cycle. A through-the-cycle (TTC) PD averages across a full cycle and stays deliberately stable. Basel capital wants TTC; accounting provisions want PIT. Most banks estimate one and transform to the other.
Global Credit Data’s benchmarking of large corporate portfolios illustrates how stable TTC estimates are meant to be: a TTC PD of 0.22% for investment-grade exposures and 2.82% for speculative-grade, held broadly constant over time. Observed annual default rates fluctuate around those anchors.
How to read a PD figure
A PD of 3% does not mean one borrower defaults 3% of the time. Rather, it means that within a pool of borrowers sharing that risk profile, roughly three in a hundred are expected to default over the horizon. PD describes a population and applies that description to an individual.
Regulators also set floors, because estimates near zero are not credible. Under the finalised Basel III standards, the PD input floor is 0.05% for corporate and institutional exposures and 0.1% for qualifying revolving retail revolvers.
How do you calculate probability of default?
There are four principal methods: the historical default rate, logistic regression, machine learning models such as random forest and gradient boosting, and survival analysis. The choice depends on portfolio size, available data, the horizon required, and whether the output must satisfy a regulator.
Method
Output
Ranks borrowers
Lifetime PD
Best suited to
Historical default rate
One rate per segment
No
No
Low-default and homogeneous pools
Logistic regression
PD at fixed horizon
Yes
Bolt-on only
Regulatory PD, scorecards
Machine learning
PD at fixed horizon
Strongest
Bolt-on only
Origination decisioning
Survival analysis
Full PD term structure
Yes
Native
Lifetime ECL, Stage 2 assets
Most institutions run two or three in combination. The sections below explain why.
The historical default rate method
The oldest approach is also the simplest. Segment the portfolio, count defaults over the period, then divide by the number of performing accounts at the start.
The formula
PD(segment) = Defaults during period ÷ Performing accounts at period start
A worked example
Suppose a bank holds 12,000 performing small-business loans in a single rating grade at the start of the year. During that year, 384 accounts reach 90 days past due.
PD = 384 ÷ 12,000 = 0.032 = 3.2%
Averaging that figure across several years — ideally a full economic cycle — produces a serviceable through-the-cycle PD for the grade.
When to use it
This method suits homogeneous, high-volume portfolios with stable underwriting. Credit cards, auto loans and standardised consumer finance all qualify. Furthermore, it is often the only option for low-default portfolios, where regression cannot be fitted at all. It also provides a benchmark against which any more sophisticated model should be checked.
Basel’s IRB requirements expect at least five years of retail data, and for corporate exposures a period ideally spanning a full cycle.
Limitations
It is backward-looking. The rate observed last year describes last year’s conditions. Should a rate shock or sectoral downturn follow, the historical figure forecasts badly.
It cannot rank within a segment. Every borrower in the bucket receives an identical PD, so all discriminatory information goes unused.
Segmentation is unstable. Cut too coarsely and the estimate is meaningless. Cut too finely and default counts fall to single digits, where sampling error dominates.
Zero-default periods break it. A segment with no defaults yields an estimate of 0%, which is both false and inadmissible under Basel floors.
A cautionary example
The 2007–09 financial crisis demonstrated the weakness plainly. US subprime mortgage pools showed low realised default rates through 2005 and 2006, and provisioning followed those observed rates. Underwriting quality had deteriorated well before defaults surfaced, but a backward-looking estimate had no mechanism to register it. Default rates then rose by an order of magnitude within eighteen months.
Similarly, European corporate speculative-grade default rates have swung from below 3% to forecasts approaching 8.5% during stressed periods, before settling back toward 4%. Any estimate anchored solely to the preceding year would have missed both turns.
Logistic regression for PD estimation
Logistic regression remains the industry standard for regulatory PD, for reasons that are as much supervisory as statistical.
The formula
The model expresses log-odds of default as a linear function of borrower characteristics:
ln( PD / (1 − PD) ) = β₀ + β₁x₁ + β₂x₂ + … + βₖxₖ
Rearranged, this gives PD directly:
PD = 1 / (1 + e^−(β₀ + β₁x₁ + … + βₖxₖ))
The logistic function maps any score onto the interval between 0 and 1. Usefully, that is exactly what a probability requires, so no clipping or rescaling is needed.
Why it dominates
Coefficients are interpretable. Each β represents a change in log-odds per unit of the predictor. Exponentiate it and a credit officer has an odds ratio to reason about.
It converts to a scorecard. Weight-of-evidence binning combined with logistic regression produces the points-based scorecards underwriting systems actually run.
It is stable on modest data. A few thousand observations with a reasonable default rate estimate coefficients reliably. Tree ensembles at that sample size tend to overfit.
Supervisors accept it. Model risk expectations under Basel and IFRS 9 weight explainability and documentation heavily.
Implementation steps
Define the target: 1 if the account defaulted within the horizon, 0 otherwise.
Assemble candidate drivers — financial ratios, behavioural variables, bureau data.
Treat outliers, then bin variables using weight of evidence.
Screen candidates on information value and check for multicollinearity.
Fit the model and retain variables with defensible signs and significance.
Calibrate the intercept to the portfolio’s central tendency.
Validate on an out-of-time sample, not merely a random hold-out.
That final point matters more than it appears. A random split shares the economic environment between training and test data, which flatters the model considerably.
Machine learning approaches to PD estimation
Tree ensembles have earned a place in PD estimation, particularly for retail and small-business portfolios with large data volumes and genuine non-linearity.
Random forest
Random forest grows many decision trees on bootstrapped samples with randomised feature subsets, then averages their predicted probabilities. It resists outliers, handles missing values gracefully and needs little tuning.
However, its outputs are frequently poorly calibrated. Averaged vote shares are not probabilities. Isotonic or Platt calibration is therefore close to mandatory before the figures enter a provisioning calculation.
XGBoost and gradient boosting
Gradient boosting builds trees sequentially, each correcting the residual errors of those before it. On structured credit data with rich behavioural features — utilisation trends, bureau enquiry velocity, payment patterns — it routinely delivers a 3 to 8 point Gini improvement over a well-built scorecard.
On thin-file applicants with a dozen variables, by contrast, that advantage largely disappears.
The constraints
Explainability. SHAP values provide local attributions and now appear as standard in documentation. Nevertheless, a SHAP plot explains a single prediction rather than a stable economic relationship. Supervisors will ask whether each driver relates to PD monotonically and sensibly. Monotonic constraints address this directly and should be used.
Overfitting on rare events. Default is an imbalanced outcome. Boosted trees will memorise the few defaulters in a training set unless time-series cross-validation and early stopping are applied.
Instability under distribution shift. Tree ensembles extrapolate poorly. When conditions move outside the training range, a logistic model degrades gradually while a boosted ensemble can fail abruptly.
A common architecture resolves the tension: gradient boosting powers the origination decision engine, where predictive accuracy converts into approval quality, while a logistic scorecard supplies the regulatory PD. The scorecard is then benchmarked against the ML model to confirm no material signal is being left unused.
Retail vs corporate PD estimation
Retail PD estimation is pooled and statistical, because portfolios contain thousands of homogeneous exposures. Corporate PD estimation is rating-based and judgement-assisted, because portfolios contain fewer, larger and more idiosyncratic exposures with too few defaults to model directly.
Retail
Corporate
Unit of estimation
Pool or segment
Individual obligor rating grade
Typical data volume
10,000s to millions
Hundreds to low thousands
Defaults observed
Sufficient for regression
Often single digits per grade
Dominant method
Scorecards, ML
Rating models, shadow ratings
Key drivers
Behavioural, bureau, utilisation
Financial statements, sector, structure
Expert judgement
Minimal
Formal override process
Basel treatment
No maturity adjustment
Maturity adjustment applies
What differs in practice
Retail portfolios support genuine statistical estimation. Behavioural data updates monthly, defaults are frequent enough to fit models, and segments are large enough that pooled estimates are stable. Consequently, the modelling challenge is technical: feature engineering, calibration and drift monitoring.
Corporate portfolios pose the opposite problem. A large corporate book may generate a handful of defaults across an entire cycle. Maximum likelihood estimation degrades badly at those counts, and complete separation is common.
How corporate PD gets estimated anyway
Banks generally use one of three routes. First, internal rating models score obligors on financial and qualitative factors, then map grades to a PD calibrated against external agency default studies. Second, shadow rating approaches replicate agency ratings on unrated obligors. Third, for very sparse portfolios, Bayesian methods and the “most prudent estimate” principle produce conservative PDs from limited or zero defaults.
Low-default portfolio methodology exists precisely because the standard toolkit fails here. This includes exposures to sovereigns, financial institutions, and specialised lending, where high exposure meets very low borrower counts.
Forward-looking PD under IFRS 9
IFRS 9 changed what a PD model must deliver. Provisions are no longer recognised when a loss is incurred; they are recognised when a loss is expected.
The three-stage model
Stage 1 covers exposures without significant deterioration since origination. These require a 12-month expected credit loss, and therefore a 12-month PD.
Stage 2 covers exposures that have shown a significant increase in credit risk (SICR). These require lifetime ECL, and therefore a lifetime PD across remaining maturity.
Stage 3 covers credit-impaired exposures, also on a lifetime basis.
What this demands of the model
Two requirements follow, and neither is satisfied by a standard 12-month scorecard.
A term structure is needed, not a point. Lifetime PD requires a curve across remaining maturity. Banks typically build one by chaining marginal PDs, applying rating transition matrices, or fitting survival models that produce the term structure natively. Survival methods also handle censoring correctly — loans that prepay or refinance are censored observations, not clean non-defaults, and treating them as the latter understates lifetime risk.
Conditioning must be point-in-time. The relationship between portfolio default rates and macroeconomic drivers is estimated, then applied to forecast scenarios. The Vasicek transformation is the standard vehicle:
PD(PIT) = Φ[ (Φ⁻¹(PD_TTC) − √ρ · Z) / √(1 − ρ) ]
Here Φ is the standard normal CDF, ρ is asset correlation, and Z is a systematic factor derived from macro variables.
Scenario weighting
IFRS 9 requires probability-weighted outcomes across multiple scenarios — typically a baseline, an upside and a downside, with documented weights. Because the relationship between macro conditions and PD is convex, probability-weighted ECL exceeds ECL calculated from the baseline alone. That result is a feature of the framework rather than a modelling artefact.
Macro drivers that carry signal include GDP growth, unemployment, policy rates and sector-specific indices. Notably, portfolio-relevant series consistently outperform generic aggregates: house price indices for mortgages, freight indices for transport lending.
How to validate and backtest a PD model
PD model validation tests three properties: discrimination (can the model separate defaulters from non-defaulters), calibration (are predicted levels accurate), and stability (has the population shifted). Each should be assessed at least annually.
Discrimination tests
Metric
Formula or definition
Acceptable range
AUC
Area under the ROC curve
0.70–0.80 application; 0.85+ behavioural
Gini
2 × AUC − 1
0.40–0.60 application scorecards
KS statistic
Max distance between cumulative distributions of defaulters and non-defaulters
Above 30 for retail applications
An AUC above 0.95 warrants investigation rather than celebration. In practice, a figure that high usually indicates target leakage — a variable that encodes the outcome, such as a collections flag populated only after default was known. It will look excellent in development and fail in production.
Calibration tests
Discrimination and calibration are independent. A model can rank borrowers perfectly and still predict 2% where the true rate is 6%. That is acceptable for approval decisions and unacceptable for provisioning.
To test calibration, bin the portfolio by predicted PD, then compare predicted against observed default rates within each bin. The Hosmer-Lemeshow test and a binomial test per rating grade are standard. Persistent one-directional deviation across grades signals a calibration problem rather than noise.
Stability tests
The Population Stability Index compares the score distribution at development against the current book:
Above 0.25 — material shift, model review warranted
Compute PSI on individual characteristics as well as the final score. A stable overall score can conceal offsetting drifts in two underlying variables, and only the characteristic-level view reveals them.
Choosing a PD estimation method
No single method is correct across all portfolios. The decision rests on four considerations.
Data depth. Fewer than roughly 50 defaults makes regression unreliable, which points toward historical rates or low-default portfolio methods. Above 500 defaults, machine learning becomes viable.
Horizon required. A 12-month regulatory PD can come from logistic regression. A lifetime PD needs survival analysis or a transition matrix.
Regulatory use. Where the output feeds capital or provisions, interpretability outweighs a marginal gain in predictive accuracy.
Portfolio type. Retail supports pooled statistical estimation. Corporate generally requires rating-based approaches with a formal judgement overlay.
For most institutions the practical answer combines several: a logistic core for rank ordering and regulatory reporting, a survival or transition layer for lifetime term structure, a macro overlay for point-in-time conditioning, and machine learning where origination accuracy justifies it.
Frequently Asked Questions
What is the difference between PD and default rate?
PD is a forward-looking estimate of default likelihood for a borrower or pool. The default rate is the backward-looking observed proportion of performing borrowers that actually defaulted within a period. PD is what the model predicts; the default rate is what happened.
How many years of data are needed for PD estimation?
Basel’s IRB minimum requirements expect at least five years of data for retail PD estimation. For corporate exposures, the observation period should ideally span a full economic cycle, since a shorter window will not capture a downturn.
Can logistic regression produce a lifetime PD?
Not directly. Logistic regression fitted on a 12-month default flag produces a single estimate at a single horizon. Lifetime PD requires chaining marginal PDs, applying a rating transition matrix, or fitting a survival model. Each route adds assumptions that must be separately documented and validated.
What is a low-default portfolio?
A low-default portfolio contains too few observed defaults for standard statistical estimation — typically sovereign, financial institution, and specialised lending exposures. Bayesian methods, external rating mapping and the most prudent estimate principle are used instead of regression.
S&P Global Ratings, Default, Transition and Recovery series
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
In practice, every rupee a bank lends carries an implicit question. But what is the chance this borrower stops paying?
PD estimation turns that question into a defensible number. Naturally, it sits at the center of modern credit risk management. Specifically, it drives regulatory capital, loan loss provisions, loan pricing, credit approval cut-offs, and portfolio strategy.
Notably, the consequences of error run in both directions. On one hand, overstate risk, and the bank prices itself out of good business. On the other hand, understate it, and losses arrive faster than banks built provisions.
Recently, the stakes in India have risen sharply. Nevertheless, the headline numbers look reassuring. Indeed, the RBI’s Financial Stability Report of June 2026 placed the gross NPA ratio of scheduled commercial banks at 1.8% as of March 2026. In other words, that is a multi-decadal low. Furthermore, the central bank’s baseline projects only a modest rise, to around 1.9% by March 2028. Additionally, capital ratios sit at multi-decade highs, with CRAR at 17.7% and CET1 at 15.3%.
In effect, this single change moves Indian banks off the incurred-loss model. Previously, provisions followed a default event. Now, under the Expected Credit Loss framework, banks must estimate a lifetime PD term for every performing exposure that shows a significant increase in credit risk.
Consequently, PD estimation stops being a capital input that a small modelling team owns. Instead, it becomes a line item flowing straight into the profit and loss account every quarter. As a result, boards, auditors, and supervisors will all read it.
To that end, this guide covers the principal PD estimation methods used in practice. For each, it sets out what the method does well, where it fails, and finally how to validate the result.
What Is Probability of Default?
Put simply, probability of default is the likelihood that a borrower fails to meet contractual obligations over a defined horizon. Specifically, it runs on a 0 to 1 scale, or equivalently 0% to 100%. At the extremes, zero means default is impossible while one means default is certain. In practice, of course, a well-specified model never produces either extreme.
Before any PD estimation exercise yields a meaningful number, however, three elements must be fixed.
The default definition
As a rule, Basel and RBI’s IRACP norms both trigger default at 90 days past due. Alternatively, lenders may declare default earlier if they judge the obligor unlikely to pay without realising collateral.
However, a model trained on a 90-DPD definition is not comparable to one trained on 30-DPD. Indeed, mixing the two remains one of the most common sources of inconsistency in Indian retail portfolios.
The horizon: 12-month versus lifetime
For instance, regulatory capital under the Internal Ratings-Based approach uses a 12-month PD. By comparison, IFRS 9 and RBI’s ECL directions ask for more. Specifically, Stage 1 assets need a 12-month PD, whereas Stage 2 and Stage 3 assets need a lifetime PD.
Importantly, a 12-month PD of 2% does not imply a five-year lifetime PD of 10%. In reality, default hazard rarely stays constant over time. This is precisely why survival methods matter.
Point-in-time versus through-the-cycle
A point-in-time (PIT) PD reflects current economic conditions and moves with the cycle. By contrast, a through-the-cycle (TTC) PD averages across a full cycle and stays deliberately stable.
Basel capital wants TTC. By contrast, ECL accounting wants PIT. Most banks therefore estimate one basis and transform to the other. Consequently, a great deal of model risk hides in that transformation.
How to read a PD number
Even so, interpretation deserves care. Notably, a PD of 3% does not mean a specific borrower will default 3% of the time. Rather, it means that within a homogeneous pool of borrowers sharing that risk profile, roughly three in a hundred will default over the horizon. In short, PD describes a population, then applies that description to an individual.
Similarly, regulators recognise that estimates near zero lack credibility. Under the finalised Basel III standards in BIS Basel Framework chapter CRE36, the PD input floor for corporate and institutional exposures rose from 0.03% to 0.05%. Meanwhile, qualifying revolving retail revolvers carry a 0.1% floor. In principle, these floors offset model risk, measurement error, and thin data. Usefully, they remind us that no PD estimate is exact.
Finally, PD is one of three parameters in the expected loss identity: EL = PD × LGD × EAD. For how the other two fit in, see our comprehensive guide to credit risk modeling, which covers Loss Given Default and Exposure at Default alongside PD.
PD Estimation Methods Compared
Before going into each method, here is how the four principal approaches differ. Ultimately, these dimensions determine which method you can actually use.
Historical Default Rate
Logistic Regression
Machine Learning
Survival Analysis
Output
One rate per segment
PD at a fixed horizon
PD at a fixed horizon
Full PD term structure
Ranks borrowers?
✗
✓
✓ strongest
✓
Lifetime PD?
✗
via bolt-on only
via bolt-on only
✓ native
Handles censoring
✗
✗
✗
✓
Minimum data
5+ yrs, ideally a full cycle
~1,000+ obs, 50+ defaults
10,000+ obs, 500+ defaults
Loan-level default timing
Interpretability
Complete
High
Low without SHAP
Moderate
Regulatory acceptance
High (benchmark use)
Highest
Conditional
Growing under IFRS 9
Best suited to
Low-default and homogeneous pools
Regulatory PD, scorecards
Origination decisioning
Stage 2 lifetime ECL
Most banks run two or three of these together rather than choosing one. Below, the reasons become clear.
Historical Default Rate Method for PD Estimation
Notably, the simplest approach to PD estimation is also the oldest. First, segment the portfolio. Then count defaults and divide by the number of accounts at the start of the period.
PD(segment) = Number of accounts defaulting in period / Number of performing accounts at period start
For example, consider a small portfolio. Suppose a bank holds 12,000 performing MSME loans in a given rating grade at the start of FY25. During the year, 384 of them hit 90-DPD. The observed one-year default rate is therefore 3.2%. Average that across several years, ideally a full cycle, and you then have a serviceable TTC PD for that grade.
Where the historical method works well
Generally, this method suits homogeneous, high-volume portfolios with stable underwriting. For instance, two-wheeler loans, gold loans, and standardised consumer durable finance all qualify.
Provided that the segment is genuinely homogeneous, the empirical rate is unbiased. Moreover, it needs no statistical assumptions at all. Additionally, it is the natural starting point for a low-default portfolio, where regression simply cannot be fitted. It also benchmarks any more sophisticated model that follows.
Basel’s IRB minimum requirements expect at least five years of data for retail PD estimation. For corporate exposures, meanwhile, they want a period spanning a full economic cycle. This is not bureaucratic conservatism. Rather, it responds directly to the method’s central weakness.
Where the historical method fails
It looks entirely backward. After all, the observed default rate for FY25 tells you only what happened under FY25 conditions. Consequently, if the next year brings a rate shock or a sectoral downturn, that rate forecasts badly. Regardless, Indian banking learned this expensively.
It cannot rank within a segment. Instead, every borrower in the bucket receives the same PD. For instance, a five-year-old MSME with declining coverage ratios gets the same number as a fifteen-year-old firm with improving margins. All the discriminatory information therefore sits unused.
Segment definition is arbitrary and unstable. Cut too coarsely and the estimate means nothing. Conversely, cut too finely and default counts collapse to single digits. At that point sampling error swamps signal. A segment with three observed defaults has a confidence interval so wide it barely constrains anything.
Low default rates break it. When a portfolio produces zero defaults in a year, the naive estimate is 0%. That figure is both false and, under Basel floors, inadmissible.
Case study: the 2015 Asset Quality Review
The corporate credit cycle of the mid-2010s illustrates historical-rate failure better than any hypothetical. Initially, through the boom years, observed default rates on large infrastructure and metals exposures stayed low. Provisioning duly followed those observed rates.
Underlying credit quality, however, had already deteriorated well before defaults surfaced. Neither the incurred-loss framework nor the historical-rate PD estimates feeding it could register that deterioration ahead of the event.
The RBI’s Asset Quality Review, launched in 2015, then forced consistent recognition across the system. Consequently, reported gross NPAs of scheduled commercial banks climbed to roughly 11.5% by 2018. Clean-up, recapitalisation, and IBC-driven resolution subsequently brought them down to 2.3% by March 2025 and 1.8% by March 2026.
Importantly, the lesson is not that banks acted dishonestly. Instead, it is that PD estimation drawn purely from recent realised defaults cannot anticipate a turning point. That structural gap is exactly what the ECL framework aims to close.
Logistic Regression for PD Estimation
Logistic regression remains the workhorse of PD estimation across the industry. Furthermore, it holds that position for reasons that are as much regulatory as statistical.
In essence, the model estimates the log-odds of default as a linear function of borrower and facility characteristics:
ln( PD / (1 − PD) ) = β₀ + β₁x₁ + β₂x₂ + … + βₖxₖ
Rearranging then gives the PD directly:
PD = 1 / (1 + e^−(β₀ + β₁x₁ + … + βₖxₖ))
Usefully, the logistic function maps any real-valued score onto the (0, 1) interval. That is exactly what a probability requires. Consequently, no transformation, clipping, or calibration hack is needed to keep estimates in range.
Why logistic regression dominates in practice
Coefficients are interpretable. Each β is a change in log-odds per unit of the predictor. Exponentiate it, therefore, and you get an odds ratio a credit officer can reason about. So when a supervisor asks why a borrower received a 4.8% PD, the model answers.
It converts cleanly to a scorecard. In addition, weight-of-evidence binning plus logistic regression produces the points-based scorecards that underwriting systems and branch staff actually use. Our walkthrough of logistic regression for PD modelling covers WOE and information value in full.
It stays stable on modest data. Typically, a few thousand observations and a reasonable default rate estimate the coefficients well. Conversely, tree ensembles at the same sample size tend to overfit.
It meets least regulatory resistance. Model risk expectations under the Basel III regulatory framework and RBI’s ECL directions weigh explainability and documented governance heavily.
How to implement logistic regression PD estimation in Python
import pandas as pd
import numpy as np
import statsmodels.api as sm
from sklearn.model_selection import train_test_split
Two practitioner notes follow from this. First, prefer statsmodels over scikit-learn for the development build. Model documentation needs p-values and standard errors to justify each variable. Unfortunately, scikit-learn does not surface them.
Second, always hold out an out-of-time sample rather than a random split. Otherwise, a random split shares the economic environment between train and test. As a result, it flatters the model considerably.
Limitations of Logistic Regression in IFRS 9 PD Modelling
Historically, logistic regression earned its dominance in a Basel world. There, the deliverable was a stable 12-month through-the-cycle PD, used once a quarter for capital.
IFRS 9 and RBI’s ECL directions ask for something structurally different. Under those new demands, several limitations become visible. None of them disqualifies the method; most banks will still build on a logistic core. Each one, however, requires a bolt-on that must itself be documented and validated.
Structural gaps: term structure and censoring
It produces a point, not a term structure. As noted, IFRS 9 requires a lifetime PD for Stage 2 and Stage 3 exposures. Meanwhile, a logistic model fitted on a 12-month default flag produces exactly one number, at one horizon.
In contrast, extending it to a 30-year mortgage means chaining marginal PDs, applying a rating transition matrix, or fitting separate models at multiple horizons. Every one of those routes introduces assumptions the original model never tested. Ultimately, the extrapolation rather than the regression drives most of the lifetime ECL.
It has no mechanism for censoring. Unfortunately, a binary setup treats loans that prepay, refinance, or remain performing at the data cut-off as clean non-defaults. In Indian retail books with high prepayment rates, this bites hard. In particular, housing and personal loans suffer most.
The reason is straightforward. Essentially, accounts that exited early get counted as successes rather than as observations that simply stopped being observed. Lifetime default risk therefore comes out understated. Survival methods handle this natively; logistic regression does not.
Conditioning and scenario problems
PIT conditioning sits outside the model. Fitted on pooled multi-year data, a logistic model delivers something closer to a hybrid or TTC estimate. Converting it to the PIT basis IFRS 9 requires means applying a separate macro scalar or Vasicek shift afterwards.
Because nobody estimates the macro relationship jointly with the borrower-level coefficients, the two components can drift apart. The model owner has no single likelihood to test.
Scenario sensitivity often runs too flat. In fact, auditors raise this criticism most consistently. During fitting, idiosyncratic borrower variables absorb the bulk of the variance. Little remains for macro drivers to explain.
As a result, the downside scenario moves ECL by only a few basis points. That implies the bank believes a severe recession barely affects its credit losses. Rarely is that credible, and defending it in an audit committee proves difficult.
SICR assessment needs an origination-date PD. Staging under IFRS 9 compares lifetime PD at the reporting date against lifetime PD at initial recognition. For loans booked before the model existed, which covers most of a legacy book, someone must reconstruct that origination PD retrospectively. The underlying data may never have been retained. Logistic regression offers no help here. The problem concerns data lineage, yet it lands squarely on the PD model owner.
Statistical constraints
Linearity in log-odds is a real constraint. Fundamentally, the method assumes each predictor moves log-odds linearly. In reality, credit drivers rarely oblige. Utilisation risk rises sharply past 80%. Vintage risk is hump-shaped. DSCR effects flatten at both tails.
Admittedly, weight-of-evidence binning is the standard fix. Nevertheless, coarse-classing discards within-bin information. The bin boundaries then become an unvalidated modelling choice that tends to destabilise on refresh.
Coefficients turn unstable in low-default portfolios. Large-corporate, NBFC, and sovereign exposures may generate a handful of defaults across an entire cycle. Consequently, maximum likelihood estimation degrades badly at those counts, and complete or quasi-complete separation is common. IFRS 9 still demands an ECL number for these exposures. Consequently, banks usually turn to external ratings-based PD mapping or a shadow-rating approach instead.
What this means for model architecture
Overall, the common architecture now runs three layers. First, a logistic model handles 12-month PD and rank ordering. Next, a survival or transition-matrix layer supplies the lifetime term structure. An explicit macro overlay delivers PIT conditioning.
That means three models, three validation exercises, and three sets of assumptions. In other words, the governance load is considerably heavier than the single scorecard that satisfied Basel. Plan for it well before the April 2027 deadline.
Machine Learning Methods for PD Estimation
Gradient-boosted trees and random forests have earned a genuine place in PD estimation. Specifically, they perform best in retail and MSME segments, where non-linearities and interactions run strong and data volumes are large.
Random Forest
Random Forest grows many decision trees on bootstrapped samples with randomised feature subsets, then averages their predicted class probabilities. Usefully, it resists outliers, handles missing values gracefully, and requires little tuning.
Its output probabilities, however, are frequently poorly calibrated. After all, averaged vote shares are not PDs. Isotonic or Platt calibration is therefore close to mandatory before the numbers touch an ECL calculation.
XGBoost and LightGBM
By contrast, these build trees sequentially. In turn, each new tree corrects the residual errors of the ensemble so far.
On structured credit data with rich behavioural features, gradient boosting routinely delivers a 3 to 8 point Gini improvement over a well-built logistic scorecard. For instance, utilisation trends, bureau enquiry velocity, EMI bounce patterns, and GST filing regularity all help here. On thin-file applicants with a dozen variables, though, the gain frequently disappears entirely.
import xgboost as xgb
from sklearn.calibration import CalibratedClassifierCV
from sklearn.metrics import roc_auc_score
base = xgb.XGBClassifier(
n_estimators=400,
max_depth=4, # keep shallow; deep trees overfit credit data
Explainability. Certainly, SHAP values give local attributions and now appear as standard in model documentation. Still, a SHAP plot explains a prediction rather than a stable economic relationship. Supervisors reviewing an IRB or ECL model will ask whether each driver relates to PD monotonically and sensibly. Monotonic constraints (monotone_constraints in XGBoost) address this directly, so use them.
Overfitting on rare events. By definition, default is an imbalanced outcome. Otherwise, boosted trees will happily memorise the handful of defaulters in a training set. Time-series cross-validation and aggressive early stopping are therefore essential.
Instability under distribution shift. Unfortunately, tree ensembles extrapolate poorly. When macroeconomic conditions move outside the training range, a logistic model degrades gracefully. A boosted ensemble, by contrast, can fail abruptly.
One architecture recurs across Indian banks, and it is defensible. Gradient boosting runs the origination decision engine, where predictive power converts directly into approval quality. A logistic scorecard supplies the regulatory PD for capital and ECL. Teams then benchmark the scorecard against the ML model, confirming that no material discriminatory power goes unused.
Survival Analysis for Lifetime PD Estimation
As discussed, logistic regression answers a binary question over a fixed window. Default in twelve months: yes or no?
By contrast, survival analysis answers a richer one. When is default likely to occur? What is the instantaneous risk at each point in the loan’s life?
Naturally, that distinction became commercially important the moment lifetime ECL arrived. Specifically, Stage 2 assets require a PD term structure across the remaining maturity. Survival models produce exactly that, natively.
Hazard rates and survival curves
The hazard h(t) is the instantaneous default rate at time t, conditional on having survived to t. Correspondingly, S(t) — the survival function — is the probability of surviving beyond t. Cumulative PD to time t is then simply 1 − S(t). A lifetime PD therefore reads directly off the survival curve at loan maturity.
Kaplan-Meier estimation produces a non-parametric survival curve from observed default times. Critically, it handles censoring correctly. Loans that prepay, refinance, or remain performing at the data cut-off are censored, not non-defaults.
By contrast, a naive logistic setup treats a loan booked three months ago as a “non-default” observation. That choice systematically biases PD estimation downwards.
Cox proportional hazards
Accordingly, Cox regression extends the survival curve to covariates:
h(t | x) = h₀(t) · exp(β₁x₁ + … + βₖxₖ)
Here, the baseline hazard h₀(t) captures the shape of default timing across the portfolio. Covariates then shift risk multiplicatively.
from lifelines import KaplanMeierFitter, CoxPHFitter
In practice, the payoff shows up in the seasoning curve. For example, Indian unsecured personal loan portfolios typically peak in hazard between months 9 and 18, then decline. A 12-month logistic PD cannot represent that shape at all. Over a five-year loan, the shape separates a defensible lifetime ECL from a guess.
How Macroeconomic Variables Affect PD Estimation
A PD model built on borrower characteristics alone assumes the economy of the training period repeats. In reality, it will not. Indeed, ECL under RBI’s directions and IFRS 9 explicitly requires forward-looking information. Macroeconomic conditioning is therefore no longer optional.
The Vasicek transformation
The standard approach links the portfolio’s observed default rate to macro drivers, then applies that relationship to forecast scenarios. Of the available vehicles, the Vasicek/Merton transformation is by far the most common:
Here Φ is the standard normal CDF, ρ is asset correlation, and Z_t is a systematic factor estimated from macro variables. A favourable macro environment, meaning positive Z, pulls PIT PD below the TTC level. Conversely, a downturn pushes it above.
Macro drivers that carry signal in Indian portfolios
Real GDP growth — the broadest indicator. It typically enters with a lag of two to four quarters, since credit stress follows activity rather than coinciding with it.
Policy rate and lending rate spreads — these matter most for floating-rate retail and MSME exposures. Here, a repo increase transmits directly to EMI burden.
CPI inflation — this compresses household surplus and drives unsecured retail delinquency.
Sector-specific series — IIP for manufacturing, residential price indices for mortgages, freight indices for commercial vehicle finance. Portfolio-relevant series consistently outperform generic aggregates.
Notably, the June 2026 FSR reinforces that last point. Despite a system average of just 1.8%, agriculture carried the highest sectoral GNPA ratio at 5.1%. An aggregate macro variable would miss exactly that kind of dispersion.
Two disciplines that matter more than variable choice
First, every coefficient must point in an economically defensible direction. For example, a model where higher GDP growth increases PD has found a spurious correlation. Therefore, drop the variable, regardless of its statistical significance.
Second, RBI’s ECL directions require probability-weighted scenarios: a baseline, an upside, and a downside, with documented weights. Because the relationship between macro factors and PD is convex, the probability-weighted ECL will exceed the ECL computed from the baseline alone. That convexity is a feature of the framework rather than an artefact.
How to Validate and Backtest a PD Model
A PD model is a regulated artefact. Accordingly, it needs evidence of discrimination, calibration, and stability, refreshed at least annually.
Discrimination: can the model separate defaulters from non-defaulters?
ROC and AUC. The ROC curve plots the true positive rate against the false positive rate across all cut-offs. In turn, AUC is the area beneath it.
At the floor, an AUC of 0.5 is random. Meanwhile, application scorecards typically land between 0.70 and 0.80. Meanwhile, behavioural models with repayment history commonly exceed 0.85. Anything above 0.95, however, should prompt a hunt for target leakage rather than celebration.
Gini coefficient. The industry’s preferred expression is simply Gini = 2 × AUC − 1.
Kolmogorov-Smirnov statistic. KS is the maximum vertical distance between the cumulative distributions of defaulters and non-defaulters. It answers a slightly different question from AUC: where in the score range does separation run strongest?
That question is operationally useful, since the answer often marks where the approval cut-off should sit. Generally, KS above 30 is acceptable for retail application models.
Importantly, discrimination and calibration are independent properties. For example, a model can rank perfectly and still predict 2% where the true rate is 6%. That is fine for approval decisions, yet disastrous for ECL.
To test this, bin the portfolio by predicted PD. Compare predicted against observed default rates in each bin, then test the difference. The Hosmer-Lemeshow test and the binomial test per rating grade are the standard tools. Ultimately, persistent one-directional deviation across grades indicates a calibration problem rather than noise.
Stability: has the population moved?
Population Stability Index (PSI) compares the score distribution at development against the current book:
Conventional thresholds run as follows. Below 0.10 is stable. Between 0.10 and 0.25 warrants monitoring. Above 0.25 signals a material shift requiring investigation.
Also compute PSI on individual drivers, not just the final score. After all, a stable overall score can conceal offsetting drifts in two underlying variables. Fortunately, the characteristic-level view catches them.
Choosing the Right PD Estimation Method
On the whole, no single approach to PD estimation is correct. Instead, the right choice depends on portfolio size, data depth, the horizon required, and the regulatory use the number will serve.
Matching method to portfolio
Broadly, historical default rates remain the sensible baseline for homogeneous portfolios. They are also the only viable route where defaults are too scarce to model.
Meanwhile, logistic regression is the default choice for regulatory PD. In a supervised environment, interpretability and stability are worth more than a marginal lift in Gini. Under IFRS 9, though, it is a starting point rather than a complete answer. Lifetime term structure, censoring, and PIT conditioning all sit outside what the regression itself estimates.
Similarly, machine learning earns its place where data is rich and non-linearity is real, provided outputs stay calibrated and monotonicity constrained. Survival analysis is less an alternative than a necessary complement, because lifetime PD term structures cannot be built from a 12-month binary model. Macroeconomic conditioning, meanwhile, is what turns any of them into a forward-looking estimate.
The road to April 2027
RBI’s ECL framework takes effect on 1 April 2027, with a glide path running to March 2031. Indian banks and NBFCs are therefore mid-way through a data and modelling build-out that will define credit risk practice for the next decade.
Unfortunately, institutions that treat this as a compliance exercise will produce models that pass validation and inform nothing. Conversely, those that treat it as a chance to understand their loan books properly will get both.
In short, PD estimation quantifies the probability that a borrower will default over a defined horizon. The result is expressed between 0% and 100%. It is one of three inputs to expected loss, alongside Loss Given Default and Exposure at Default. In practice, it drives regulatory capital, loan pricing, credit approval decisions, and provisioning under IFRS 9 and RBI’s ECL directions.
Which PD estimation method is most accurate?
In practice, no single method wins across all portfolios. Admittedly, gradient boosting delivers the strongest discrimination on large retail books with rich behavioural data, often 3 to 8 Gini points above a logistic scorecard.
However, rank ordering is not the only requirement. Equally, regulatory PD must stay interpretable, stable, and calibrated. On balance, logistic regression usually wins on that combined test. That is why it remains the industry standard for capital and ECL despite lower raw predictive power.
What is the difference between 12-month PD and lifetime PD?
A 12-month PD is the probability of default within one year from the reporting date. Notably, Basel capital and IFRS 9 Stage 1 assets both use it.
A lifetime PD is the cumulative probability of default over the remaining contractual life of the exposure. Conversely, Stage 2 and Stage 3 assets require it. Importantly, you cannot derive lifetime PD by scaling the 12-month figure, because default hazard varies with loan seasoning.
What is a good AUC or Gini for a PD model?
For application scorecards, an AUC between 0.70 and 0.80 is typical and acceptable. Equivalently, that corresponds to a Gini of 0.40 to 0.60. Behavioural models with repayment history routinely exceed 0.85.
However, an AUC above 0.95 usually signals target leakage, meaning a variable in the model encodes the outcome. Investigate before deployment rather than celebrating.
How often should a PD model be validated?
At minimum, annually. Specifically, validation should cover discrimination (AUC, Gini, KS), calibration (predicted versus observed default rates by grade), and stability (PSI on the score and on individual characteristics).
Additionally, more frequent monitoring is warranted after material changes to underwriting policy, product mix, or macroeconomic conditions. Under RBI’s ECL directions, model validation sits within a three-tier model risk management structure spanning business, risk, and audit functions.
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
Credit Risk in Indian Banking: What RBI’s Data Actually Shows
Every risk professional in Indian banking eventually asks the same question: is credit risk actually improving, or does it just look that way in aggregate numbers? Based on the Reserve Bank of India’s own published data, the answer is both. System-wide asset quality has genuinely strengthened over the past five years. But the composition of that risk is shifting in a direction that deserves closer attention.
This piece works through RBI’s Financial Stability Reports (FSR), sectoral credit data, and the newly finalized Expected Credit Loss (ECL) framework. Together they show what’s really happening with credit risk in Indian banking between 2020 and 2025. It does not rely on a proprietary survey or projected estimates dressed up as findings. In fact, every figure below is sourced directly to a named RBI report. That distinction matters: in a domain where regulators, auditors, and rating agencies check your numbers, credibility is the entire product.
Three questions structure the analysis. How has aggregate asset quality moved since the 2020 pandemic shock? Where is risk concentrating today, even as headline numbers improve? And what does the incoming ECL regime signal about how Indian banks will need to manage credit risk going forward?
Methodology and Data Sources
This analysis draws on primary RBI publications, cross-checked across multiple reporting periods for consistency.
RBI Financial Stability Reports (FSR): Published twice yearly. They consolidate Gross NPA (GNPA) and Net NPA (NNPA) ratios of Scheduled Commercial Banks (SCBs), capital adequacy (CRAR), bank-group-wise asset quality, and stress test results. Specifically, this piece uses FSR editions from January 2021 through December 2025.
RBI Sectoral Deployment of Bank Credit data: Monthly data on credit growth to industry, services, agriculture, and personal loans. It’s sourced from 41 banks representing roughly 95% of non-food credit.
RBI’s ECL framework releases: The draft ECL directions (October 2025) and final directions (April 2026). These describe the shift from incurred-loss to forward-looking PD/LGD/EAD-based provisioning, effective April 1, 2027.
One scope note: RBI does not publish a standardized “default rate by loan product” table. Where this article cites loan-category or bank-group figures, they are GNPA ratios: the share of gross advances classified as non-performing. It’s the metric RBI itself uses, and the one directly comparable across periods.
Finding 1: Asset Quality Has Improved for Five Consecutive Years
SCB GNPA: from 8% to 2.1% in five years
Period
GNPA Ratio
NNPA Ratio
Source
March 2020
8.4%
—
RBI FSR, Jan 2021
September 2020
7.5%
—
RBI FSR, Jan 2021
March 2024
2.8%
0.6%
RBI FSR, Jun 2024
March 2025
2.3%
0.5%
RBI FSR, Jun 2025
September 2025
2.1%–2.2%
—
RBI FSR, Dec 2025
March 2027 (projected, baseline)
1.9%
—
RBI FSR, Dec 2025
RBI’s January 2021 report recorded a September 2020 GNPA ratio of 7.5%, down from 8.4% in March 2020. That was a system still absorbing the pandemic shock. GNPA had fallen to 2.8% by June 2024, then to 2.3% by March 2025. It touched a multi-decade low of 2.1% by September 2025, and RBI projects further improvement to 1.9% by March 2027 under its baseline scenario.
In practice, this reflects five years of balance sheet cleanup: post-IBC resolution of legacy corporate stress, tighter underwriting after the 2018–2020 NBFC stress episode, and stronger capital buffers overall. Meanwhile, system-wide CRAR remains comfortably above regulatory minimums, with public sector banks at 16% and private banks at 18.1% as of September 2025.
In short, aggregate GNPA is a lagging confirmation of underwriting discipline, not a leading indicator. A PD model trained mainly on 2020–2022 stressed data will overstate current default risk. One trained only on 2023–2025 benign data risks understating tail risk in the next downturn.
Explore our Credit Risk Modeling Certification Training for a structured approach to PD estimation across credit cycles.
Finding 2: Improvement Isn’t Even Across Bank Groups
PSBs are catching up fast
For instance, PSB GNPA fell sharply from 3.7% in March 2024 to 2.8% in March 2025. Meanwhile, private bank GNPA held roughly stable at 2.8% over the same period, and foreign banks improved from 1.2% to 0.9%.
Even so, this convergence matters. For most of the post-2015 asset-quality-review era, PSB asset quality lagged private banks significantly, largely on corporate exposures. That gap has now nearly closed at the aggregate level. However, remaining risk differs by bank group, which leads to the more consequential finding below.
Finding 3: Unsecured Retail Is Where New Risk Concentrates
The retail risk hiding inside a good headline number
This is the most important finding for practitioners, because it sits underneath the reassuring headline number. According to RBI’s December 2025 FSR, roughly 53.1% of retail loan slippages now originate from unsecured products like personal loans and credit cards. At private banks, unsecured loans account for nearly 76% of fresh slippages. GNPA on unsecured retail loans stood at 1.8%, versus 1.1% for overall retail advances.
In other words, the 2.1% aggregate GNPA figure blends a very clean secured/corporate book with a smaller, faster-deteriorating unsecured retail book. RBI flagged this as a fintech-adjacent risk, tied to fast credit growth in small-ticket personal loans to borrowers under 35 through digital lending channels.
This pattern, in fact, tracks with operational experience. Unsecured lending has weaker recovery mechanics (no collateral to liquidate, higher LGD), shorter behavioral history on new-to-credit borrowers, and faster origination cycles that compress underwriting review. Moreover, it is the segment where forward-looking provisioning matters most, since unsecured risk builds up quietly between formal NPA recognition points.
As a result, portfolio-level GNPA alone is no longer sufficient. Overall, segment-level GNPA and vintage curves for unsecured retail belong alongside the aggregate number in any board-level risk dashboard.
Finding 4: ECL Will Formalize This Shift
Why the 2027 ECL shift matters here
RBI has issued directions introducing forward-looking ECL provisioning, replacing the incurred-loss model. It takes effect April 1, 2027, for scheduled commercial banks excluding RRBs, Small Finance Banks, and payments banks. ECL provisioning must be based on a bank’s own historical PD and LGD data spanning at least five years, subject to RBI-specified floors. Accounts 30–90 days past due move into Stage 2, a materially earlier trigger than the current framework.
Overall, the shift aligns India’s prudential norms with global IFRS 9 standards. In addition, it requires closer integration between finance and risk functions, as forward-looking macroeconomic scenarios become a formal input to provisioning.
Indeed, this is a direct regulatory response to Finding 3. An incurred-loss model recognizes impairment only after default has effectively occurred. ECL requires estimating expected loss, via PD, LGD, and EAD, well before that point, catching unsecured deterioration earlier in the cycle.
Even so, for banks building this capability, it isn’t a compliance task to fully outsource. RBI has explicitly made a bank’s board and senior management responsible for the adequacy of the ECL framework. Consequently, internal teams need working fluency in PD/LGD/EAD construction, not just the ability to read vendor output. However, it’s worth noting that the standard formula, Expected Loss = PD × LGD × EAD, assumes independence between the three components. In practice they’re correlated: LGD tends to rise in the same downturns that push PD higher. That’s why RBI’s stress tests apply adverse scenarios jointly rather than multiplying baseline figures in isolation.
What This Means for Banks and Risk Teams
First, aggregate GNPA improvement is real but incomplete. Segment-level monitoring, especially for unsecured retail, deserves as much attention as the headline ratio.
PD/LGD model recency matters. RBI’s own five-year minimum spans both a stressed period (2020–2021) and a benign one (2023–2025). Models need to represent both.
Collateral still matters, but isn’t the whole story. Unsecured products drive a disproportionate share of new slippages. In turn, this argues for tighter underwriting in that segment, not a wholesale retreat from unsecured lending.
Finally, the 2027 ECL deadline is closer than it looks. In practice, building five years of clean PD/LGD data and validation capability is a multi-year undertaking. Banks starting in 2026 are already behind institutions that began in 2024–2025.
Recovery rate discipline matters for LGD. LGD = 1 − Recovery Rate only holds up when ‘recovery rate’ is the economic, discounted, net-of-cost rate, not the nominal amount eventually collected.
Explore our Credit Risk Modeling Certification Training to build PD, LGD, and EAD modeling skills ahead of the 2027 ECL transition, or see Understanding Credit Risk: Definition and Types for foundational concepts referenced throughout.
FAQ
What is the current GNPA ratio of Indian banks?
As of September 2025, SCB GNPA stood at 2.1%, a multi-decade low, per RBI’s December 2025 Financial Stability Report.
Is unsecured lending riskier than secured lending right now?
Yes, and the gap is widening. Unsecured retail GNPA was 1.8% versus 1.1% for overall retail advances, and unsecured products drove over half of all retail slippages.
When does RBI’s ECL framework take effect?
RBI’s ECL Directions were issued 27 April 2026 and take effect April 1, 2027. They apply to commercial banks, excluding small finance banks, payments banks, and local area banks.
Does EL = PD × LGD × EAD fully capture expected loss?
It’s the standard starting formula, but it assumes PD, LGD, and EAD move independently. In stress, they’re correlated — which is why RBI applies adverse scenarios jointly rather than multiplying baseline values.
Conclusion
The data supports a measured conclusion, not a triumphant one. Indeed, Indian banking’s asset quality genuinely improved for five straight years, and RBI’s own numbers back that up without embellishment. However, the same data shows risk isn’t disappearing. Instead, it’s relocating toward unsecured retail lending, addressed through a regulatory shift that will demand more rigorous PD, LGD, and EAD modeling capability than most institutions currently have in-house. For risk analysts, credit officers, and model validators, that combination is telling: improving headline numbers alongside a harder compliance mandate. It’s exactly why 2025–2027 is a build-capability window, not a wait-and-see one.
This analysis is based on RBI’s Financial Stability Reports, Sectoral Deployment of Bank Credit data, and RBI’s ECL Directions (2025–2026). Figures are reported as published at the cited dates; readers should consult original RBI releases for the most current data.
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
In March 2021, a little-known family office called Archegos Capital Management defaulted on margin calls from its banks. Within days, Credit Suisse lost $5.5 billion. Nomura lost close to $2.5 billion. Morgan Stanley and UBS lost roughly $1 billion and $774 million. Combined, global banks lost more than $10 billion — not because of a market crash, but because of one counterparty’s concentrated, hidden leverage.
That’s credit risk in its purest form. It’s the possibility that someone you’ve extended money or exposure to won’t pay it back, and the cascading damage that follows when large exposures go bad at once.
Most people equate credit risk with loan defaults. That’s only part of the picture. Credit risk shows up in derivatives, trade settlements, corporate bonds, and interbank lending. It appears anywhere one party depends on another to deliver.
This guide breaks down what credit risk actually means. It covers the distinct types every risk professional needs to recognize, and how banks measure and manage it in practice, in India and globally.
What is Credit Risk?
Credit risk is the possibility that a borrower or counterparty fails to meet a financial obligation, causing a loss to the lender. That’s the formal definition. In plain terms: it’s the risk that you lend money, extend credit, or enter a contract with someone, and they don’t hold up their end. The Reserve Bank of India’s Guidance Note on Credit Risk Management frames it more precisely. Credit risk can be an individual transaction risk — the chance that one specific loan goes bad. Or it can be a portfolio risk, which looks at how credit losses behave in aggregate across a bank’s book. A single bad loan is a manageable, expected cost of doing business. Thousands of correlated bad loans, going bad at once because they share a common vulnerability, is a solvency event.
Credit risk isn’t limited to banks lending to individuals or businesses. It appears in:
Loans and advances — the most familiar form, where a borrower fails to repay principal or interest
Bonds and fixed-income securities — where an issuer defaults on coupon payments or principal at maturity
Derivatives contracts — where a counterparty can’t meet its obligations under a swap, option, or forward
Trade finance and settlement — where one party in a transaction fails to deliver cash or securities as agreed
Guarantees and letters of credit — where a bank stands behind another party’s obligation and gets called on to pay
Under the Basel framework, credit risk-weighted assets typically make up the largest share of the capital a bank must hold. That’s larger than market risk or operational risk combined, for most commercial banks. This is why understanding credit risk isn’t a niche specialty. It’s the foundation most of banking risk management sits on.
Types of Credit Risk
Credit risk isn’t one uniform threat. RBI’s own framework splits transaction-level credit risk into default risk and rating migration risk. It splits portfolio-level risk into intrinsic risk and concentration risk. Layered on top of that, banking practice recognizes several distinct sub-types worth understanding individually. Three matter most for anyone building a working knowledge of the field: counterparty risk, concentration risk, and settlement risk.
The distinction isn’t academic. Each type demands a different measurement approach, a different mitigation strategy, and often a different team within a bank’s risk function. A credit officer underwriting a retail loan thinks primarily about default risk on that single borrower. A treasury desk trading derivatives thinks primarily about counterparty risk and daily mark-to-market exposure. A chief risk officer reviewing the whole institution thinks about concentration across the entire book. Does the bank have too much riding on one sector, one region, or one large group of related borrowers? Confusing these categories, or managing them with a single generic framework, is exactly how risks that look small individually compound into something systemic.
Counterparty Risk
Counterparty risk is the risk that the other party in a financial contract fails to fulfill their side of the deal. This isn’t a traditional borrower — it’s a trading or derivatives counterparty. It’s especially relevant in derivatives, securities lending, and prime brokerage relationships. There, exposure isn’t a fixed loan amount. It’s a fluctuating mark-to-market value.
The Archegos collapse is the clearest recent illustration. Archegos used total return swaps to build enormous, concentrated positions in a small number of stocks. It never owned the shares directly, and never disclosed the size of its bets. Its prime brokers — Credit Suisse, Nomura, Morgan Stanley, Goldman Sachs, and UBS — each saw only their own slice of Archegos’s exposure. None had visibility into the full picture. When the fund’s portfolio value fell roughly 30% in four days in March 2021, it couldn’t meet margin calls. Its brokers had to liquidate billions of dollars in positions simultaneously. Those unable to exit fast enough absorbed massive losses.
Analysts who studied the collapse point to a specific, technical failure mode: wrong-way risk. This occurs when a bank’s exposure to a counterparty grows precisely as that counterparty’s ability to pay deteriorates. Archegos’s swap exposure ballooned in tandem with its portfolio’s decline. The worse things got, the more the banks were owed, and the less able Archegos was to cover it. The European Central Bank later reviewed 23 major banks’ derivatives exposures. It found “material shortcomings” in how counterparty credit risk was governed across the industry.
Concentration Risk
Concentration risk arises when a bank’s credit exposure clusters too heavily around a single borrower, sector, geography, or risk factor. Diversification is supposed to protect a lender. If one borrower fails, the loss should be a small fraction of the total book. Concentration undermines that protection entirely.
Archegos illustrates this too, in a second way. The fund’s own portfolio was concentrated in a handful of technology and media stocks. When those specific names fell, there was no offsetting position to cushion the blow. The losses hit every position at once. Banks face the mirror image of this problem in lending. Heavy exposure to one industry, real estate for instance, or one large corporate group, means a single sector downturn can impair a disproportionate share of the loan book. RBI’s prudential framework directly addresses this. It sets single-borrower and group-borrower exposure limits specifically to prevent Indian banks from building the kind of concentrated exposure that made Archegos so dangerous to its lenders.
Settlement Risk
Settlement risk is sometimes called Herstatt risk, after the 1974 collapse of Bankhaus Herstatt. It’s the risk that one party in a transaction delivers its side — cash or securities — while the counterparty fails to deliver theirs. This typically happens because of timing gaps, or the counterparty’s failure between trade execution and final settlement. Herstatt was shut down by German regulators mid-day. By then it had already received Deutsche Mark payments from counterparties, but it hadn’t yet sent back the US dollars it owed in return. Banks on the other side of those foreign exchange trades lost their payments outright.
That single event reshaped global payments infrastructure. It led directly to the creation of CLS Bank, a settlement system designed specifically to eliminate this timing gap in foreign exchange transactions. CLS does this by settling both legs of a trade simultaneously. Settlement risk remains a live concern anywhere payment and delivery aren’t simultaneous. It’s a reminder that credit risk isn’t only about long-term loans — it can materialize in a transaction that’s supposed to finish within hours.
Default Risk Explained
Default risk sits at the center of every credit risk framework. It’s the risk that a borrower simply stops paying. Understanding what causes defaults, and how professionals frame the probability of one occurring, is foundational to everything else in credit risk modeling.
It’s also the oldest form of credit risk banks have grappled with. Long before derivatives, prime brokerage, or cross-border settlement systems existed, lenders were already trying to predict which borrowers would repay and which wouldn’t. Every other risk type covered in this guide is, in some sense, a more specialized variant of the same underlying question: will the party on the other side of this transaction deliver what they owe? Default risk asks that question in its most direct form, applied to a straightforward loan or credit exposure.
What Causes Defaults
Defaults rarely happen for one isolated reason. They typically result from a mix of factors:
Cash flow deterioration. A borrower’s income or business revenue falls below what’s needed to service debt. For individuals, this might mean job loss. For corporates, it might mean a demand shock or margin compression.
Over-leverage. A borrower takes on more debt than their income can realistically support, leaving no buffer when conditions turn even mildly unfavorable.
Macroeconomic shocks. Recessions, interest rate spikes, or sector-wide downturns push otherwise healthy borrowers into distress simultaneously — the mechanism behind concentration risk turning into realized losses.
Governance and fraud. Misrepresented financials, diverted funds, or outright fraud can turn an apparently strong borrower into a default risk overnight. Archegos itself is a governance case as much as a market one — Bill Hwang was later convicted on fraud and racketeering charges tied to how the fund misled its own counterparties about position size and concentration.
Willful default. Occasionally, a borrower who can pay simply chooses not to, usually when the cost of default (reputational or legal) is judged lower than the cost of repayment.
The Probability Framework
Rather than treating default as a binary surprise, credit risk professionals model it as a probability. A Probability of Default (PD) is expressed as a percentage likelihood that a borrower defaults within a defined time horizon, typically 12 months for regulatory purposes. A PD of 2% doesn’t mean a specific borrower is “slightly risky.” It means that, across a large pool of similar borrowers, roughly 2 in 100 are expected to default within that window.
This probabilistic framing matters. It converts an unpredictable individual event into something a bank can price, provision for, and manage at portfolio scale. Building a reliable PD estimate is a discipline in its own right — deciding which borrower characteristics matter, which statistical technique to use, and how to validate the result. We cover the full methodology, including logistic regression and survival analysis, in our detailed guide to PD estimation.
What matters at this stage is the underlying logic: default isn’t modeled as a yes/no outcome for an individual borrower. It’s modeled as a rate across a population, calibrated against historical data and adjusted for current conditions.
Credit Risk in International Banking
Credit risk doesn’t stop at national borders, and neither does its regulation. Two frameworks matter most for anyone working in or around Indian banking. One is the International Accounting Standards Board’s approach to credit loss recognition. The other is RBI’s domestic prudential framework.
Why does a domestic Indian bank need to care about an international accounting standard? Because capital markets are global, even when a bank’s loan book isn’t. Foreign investors, rating agencies, and cross-border lenders all benchmark a bank’s provisioning and capital adequacy against international norms, whether or not that bank operates outside India. A framework that looks conservative by domestic standards can still look under-provisioned by global standards. That gap shows up directly in borrowing costs, credit ratings, and investor confidence.
The IASB and IFRS 9
The International Accounting Standards Board (IASB) issued IFRS 9 specifically to fix a weakness exposed by the 2008 financial crisis. The old “incurred loss” model only recognized credit losses after a default had already happened — by which point it was too late for provisions to cushion the blow. IFRS 9 replaced that with an Expected Credit Loss (ECL) model. It requires banks to recognize losses based on forward-looking estimates, before default occurs. IFRS 9 has been adopted across more than 140 jurisdictions worldwide, making it one of the most widely applied accounting standards in global banking.
RBI’s Framework and India’s ECL Transition
India has historically run on a different system: the Income Recognition and Asset Classification (IRAC) norms. These classify loans as standard, sub-standard, doubtful, or loss, based on how many days they’re overdue, and provision accordingly. That’s closer to the old “incurred loss” approach IFRS 9 was designed to replace.
That’s now changing. RBI has confirmed that an ECL-based provisioning framework, with prudential floors, will apply to all Scheduled Commercial Banks from April 1, 2027. Under the new norms, banks will classify financial assets into Stage 1, Stage 2, or Stage 3. That classification depends on assessed credit losses at initial recognition, and at each subsequent reporting date — directly mirroring the IFRS 9 structure used globally.
RBI has also been actively updating its credit risk rules to align with Basel Committee on Banking Supervision (BCBS) standards. In 2026, it revised its counterparty credit risk framework specifically to bring India’s derivatives exposure measurement closer to global norms. That’s a change regulators internationally have prioritized in the years following the Archegos collapse. India’s Capital-to-Risk-Weighted-Assets Ratio (CRAR) requirement of 9% for scheduled commercial banks also exceeds the 8% global Basel III minimum. That reflects RBI’s consistently more conservative capital stance relative to international baselines.
Measuring Credit Risk
Defining and categorizing credit risk only gets a risk team so far. Managing it requires quantifying it — turning a qualitative concern into numbers that inform pricing, provisioning, and capital decisions. Four metrics form the core toolkit.
Probability of Default (PD)
As covered above, PD is the percentage likelihood that a borrower defaults within a set time horizon. It’s typically estimated through logistic regression or survival analysis, using historical repayment data, bureau scores, and financial ratios as inputs. PD is the starting point for nearly every downstream credit risk calculation.
Loss Given Default (LGD)
Default doesn’t automatically mean total loss. LGD measures the proportion of exposure a lender expects to actually lose after accounting for recoveries — collateral liquidation, guarantees, or restructuring. Getting LGD right requires more care than it first appears. It has to reflect the economic recovery rate, discounted for the time value of money and net of recovery costs — not just the raw cash eventually collected. A recovery that takes three years to materialize is worth meaningfully less than the same amount recovered immediately.
Credit Ratings
Credit ratings translate a borrower’s creditworthiness into a standardized letter grade, from investment-grade (AAA down to BBB-) to speculative or junk grades below that. In India, these come from agencies like CRISIL, ICRA, and CARE; internationally, from S&P, Moody’s, and Fitch. Ratings aren’t a substitute for internal PD modeling. But they serve two practical purposes. They give banks an independent, externally validated view of risk. And they directly feed into regulatory risk-weighting under the Basel Standardised Approach, where a lower rating translates into a higher capital charge against that exposure.
Expected Loss and Exposure at Default (EAD)
Exposure at Default (EAD) measures how much a lender is actually on the hook for at the moment a default happens. For revolving credit like credit cards, this can run higher than the current outstanding balance, since distressed borrowers often draw down more of their available limit before defaulting. Combining EAD with PD and LGD produces Expected Loss (EL) — the amount a bank should provision for a given exposure on average. It’s worth noting this combined figure is a standard industry simplification. It assumes PD, LGD, and EAD move independently, which understates risk in downturns, when all three tend to move against the lender simultaneously. Regulators address this specific gap by requiring “downturn” LGD and EAD estimates, rather than accepting benign-cycle averages alone.
Used together, these four metrics let a bank move from “this borrower feels risky” to a specific number. That number can be priced into a loan, provisioned for on the balance sheet, and reported to regulators with a defensible methodology behind it.
Conclusion
Credit risk is broader and more structured than “will this loan get repaid.” It spans counterparty exposure in derivatives, concentration across a portfolio, settlement timing in payments, and default at the level of an individual borrower. Each has distinct causes, distinct historical failures, and distinct measurement approaches. Understanding what is credit risk at this level of detail is the foundation every subsequent risk modeling technique builds on, from PD estimation through to full expected-loss calculation.
The Archegos collapse, the 2027 shift to ECL-based provisioning in India, and RBI’s ongoing alignment with Basel counterparty risk standards all point to the same conclusion. Credit risk management is not static. It evolves in direct response to the failures that expose its gaps.
Explore our Credit Risk Modeling Certification Training to master Risk Analytics. Build the PD, LGD, and EAD models covered here from scratch. Work through India’s transition to ECL provisioning, and learn the model validation techniques practicing risk teams use every day.
A practical framework for PD, LGD, EAD and expected loss modeling in banking
Introduction
Between March 2018 and March 2021, gross non-performing assets at India’s public sector banks fell from a peak of 14.58% to 9.11% of advances. By September 2025, the system-wide gross NPA ratio for all scheduled commercial banks had dropped further, to a multi-decade low of 2.15%. That turnaround wasn’t accidental. Tighter underwriting, RBI’s Asset Quality Review, the Insolvency and Bankruptcy Code, and — underlying all of it — better credit risk models drove it.
The lesson for anyone building a career in banking analytics is straightforward. Credit risk is the single largest risk category banks carry on their balance sheets. How well a bank measures it directly determines capital adequacy, profitability, and long-term survival. A bank that underestimates default risk lends into losses it didn’t provision for. A bank that overestimates it prices good borrowers out of the market and loses share to competitors.
This guide walks through the complete framework professionals use to model credit risk. It covers the regulatory logic that makes credit risk modeling non-negotiable. It covers the three parameters that quantify it — PD, LGD, EAD. And it covers the statistical and machine learning techniques used to estimate each one. Maybe you’re a risk analyst preparing for an FRM exam. Maybe you’re a data scientist moving into banking analytics. Maybe you’re evaluating a credit risk modeling certification. Either way, this article builds the structured foundation technical interviews and real project work both expect.
By the end, you’ll understand the formulas. You’ll understand why each one exists, how professionals estimate it in practice, and where Indian banks apply it today.
What is Credit Risk and Why It Matters
Credit risk is the possibility that a borrower fails to meet a contractual debt obligation. The borrower could be an individual, a company, or a counterparty. Either way, it results in a financial loss for the lender. It is the risk that repayment doesn’t happen as agreed: a missed EMI, a defaulted corporate bond, a counterparty that can’t settle a derivative contract.
For a bank, credit risk isn’t one risk among many — it’s usually the dominant one. Loans and advances typically make up the largest share of a bank’s assets. Under the Basel framework, credit risk-weighted assets (RWA) form the biggest component of the capital a bank must hold. When credit risk is mismeasured, everything built on top of that measurement goes wrong too: capital adequacy ratios, loan pricing, provisioning, dividend capacity.
Why Credit Risk Modeling Matters to a Bank’s Survival
Three consequences flow directly from how well a bank models credit risk:
Capital adequacy. Under Basel III, banks must hold capital proportional to the risk-weighted value of their assets. The Capital-to-Risk-Weighted-Assets Ratio (CRAR) for India’s scheduled commercial banks stood at a strong 17.2% as of September 2025 — well above the regulatory minimum. Risk measurement and provisioning discipline improved sharply after 2015, and it shows.
Provisioning under IFRS 9. IFRS 9 replaced the incurred-loss model with an expected credit loss (ECL) model. Banks must now recognize losses before a default happens, based on modeled probabilities. That makes credit risk models a direct input into the profit and loss statement, not just a back-office risk tool.
Pricing and portfolio strategy. A bank that can accurately rank borrowers by risk can price loans correctly, set appropriate limits, and choose which segments to grow or shrink. A bank that can’t ends up either losing good customers to competitors with sharper pricing, or accumulating bad loans that eventually show up as NPAs.
The RBI Regulatory Backdrop
The Reserve Bank of India has steadily tightened the credit risk framework banks operate under. Its 2015 Asset Quality Review forced transparent recognition of stressed assets that restructuring had previously hidden. More recently, in November 2023, the RBI raised risk weights on unsecured consumer credit — from 100% to 125% for banks. NBFC exposures saw a further increase. The goal: curb underpriced risk-taking in retail lending. The RBI has also begun articulating principle-based guidance for AI use in credit decisioning through its evolving regulatory framework. Model governance, not just model accuracy, is now squarely on the regulator’s radar.
This regulatory environment is precisely why credit risk modeling has become a core competency banks and NBFCs actively hire for — not an optional analytics add-on.
Core Components of Credit Risk Modeling
Every credit risk model, regardless of the statistical technique behind it, answers one question: how much money could the bank lose on this exposure, and how likely is that loss?
The industry-standard approach breaks “expected loss” into three independently modeled components:
1. Probability of Default (PD)
PD is the likelihood that a borrower will fail to meet their obligations within a defined time horizon. That’s typically 12 months for regulatory capital purposes, or lifetime PD under IFRS 9 for certain asset stages. PD is expressed as a percentage. A PD of 3% means a 3% chance of default within the horizon, based on the borrower’s characteristics.
2. Loss Given Default (LGD)
Default doesn’t always mean total loss. If a borrower defaults, the bank usually recovers something — through collateral liquidation, guarantees, or restructuring. LGD is the proportion of the exposure the bank expects to actually lose after recoveries. It’s expressed as a percentage of the exposure. An LGD of 40% means that, on average, the bank recovers 60% of what it’s owed after a default.
3. Exposure at Default (EAD)
EAD is the total value the bank carries at the moment of default — not the sanctioned limit, but the amount actually outstanding. For a term loan, this is close to the outstanding balance. Revolving facilities like credit cards or cash credit limits work differently. There, EAD must account for the possibility that the borrower draws down more of the limit before defaulting.
Putting It Together: Expected Loss
These three parameters combine into the foundational credit risk formula:
Expected Loss (EL) = PD × LGD × EAD
Consider a simple example: a bank has an outstanding exposure (EAD) of ₹1 crore to a borrower with a PD of 4% and an LGD of 45%.
Expected Loss = 0.04 × 0.45 × ₹1,00,00,000 = ₹1,80,000
This ₹1.8 lakh isn’t a one-off loss estimate for a single account. It’s the amount the bank should provision for, on average, across a portfolio of similar loans. Multiply this calculation across thousands of accounts, segmented by product, geography, and borrower type. The result is a portfolio-level expected loss figure that feeds directly into provisioning, pricing, and capital planning.
Why This Formula Is a Simplification
EL = PD × LGD × EAD is the standard, industry-wide formula — it’s not wrong. But it rests on an assumption worth naming explicitly: that PD, LGD, and EAD are independent of one another. In reality they aren’t. Recoveries tend to get worse exactly when default rates spike, because both track the same macroeconomic cycle. A recession depresses collateral values and pushes more borrowers into default at the same time. Multiplying three independent point estimates understates loss in exactly the scenarios that matter most. That’s why regulators require downturn LGD and downturn EAD/CCF add-ons rather than accepting benign-cycle averages.
From Single-Period EL to Lifetime ECL
The formula above is also a single-period number, typically 12 months. Under IFRS 9, lifetime ECL for Stage 2 and Stage 3 assets isn’t one multiplication. It’s a discounted sum of marginal expected losses across every remaining period of the loan’s life:
That distinction matters in practice. A 12-month EL figure and a lifetime ECL figure for the same loan can differ substantially. Conflating the two is a common — and consequential — modeling error.
Each of the three components — probability of default, loss given default, and EAD — demands its own data, statistical techniques, and validation standards. Getting the overall expected loss number right depends entirely on getting each of these three right individually. It also depends on staying honest about where the simplifying assumptions break down.
Probability of Default (PD) Estimation
PD estimation is where most credit risk modeling careers begin, because it has the richest data history and the most mature statistical toolkit.
The Data Foundation
PD models draw on historical loan performance data: borrower demographics, financial ratios, repayment history, bureau scores (like CIBIL in India), and macroeconomic variables. The target variable is binary. Did the borrower default within the observation window, typically defined as 90 days past due (DPD), or not?
Method 1: Logistic Regression
Logistic regression remains the industry workhorse for PD modeling, and for good reason. It’s interpretable, regulator-friendly, and produces a probability output directly — exactly what PD requires.
The logistic regression model estimates:
PD = 1 / (1 + e^-(β₀ + β₁X₁ + β₂X₂ + … + βₙXₙ))
X₁ through Xₙ are borrower characteristics: debt-to-income ratio, bureau score, vintage, loan-to-value ratio, and so on. Analysts estimate the β coefficients from historical default data.
Worked example: Suppose a simplified model uses two variables — bureau score and debt-to-income (DTI) ratio — and produces this equation:
PD = 1 / (1 + e^17.165) ≈ 0.00003, or effectively near-zero risk — consistent with a strong bureau score and moderate leverage.
In practice, banks convert these raw PD outputs into a credit scorecard. It’s a points-based system: each variable band — score range, income bracket, tenure — contributes points that sum to a final score. That score maps back to a PD and a risk grade, AAA to D for instance. Most retail lending decision engines run on exactly this system.
Method 2: Survival Analysis
Logistic regression answers “will this borrower default within 12 months?” It doesn’t naturally answer “when.” Survival analysis models the time to default instead, using techniques like the Cox Proportional Hazards model or Kaplan-Meier estimation. It treats default as an “event.” It treats non-defaulted, still-active loans as “censored” observations.
This matters for two practical reasons. First, it lets banks estimate lifetime PD curves, which IFRS 9 requires for Stage 2 and Stage 3 assets. Second, it naturally handles loans that are still performing at the end of the observation period. It doesn’t treat them as “non-defaults” the way a simple logistic model would. That subtlety matters a great deal in mortgage and long-tenure corporate lending portfolios.
Machine Learning Extensions
Random forests, gradient boosting (XGBoost, LightGBM), and neural networks increasingly supplement — not always replace — logistic regression. This works particularly well where alternative data is available: transaction behavior, digital footprint, utility payment history. Industry research on large Indian banks backs this up. ML models can improve default prediction accuracy over ratio-based or bureau-only assessments, particularly for thin-file borrowers without a long credit history. The trade-off is interpretability. Regulators and internal model validation teams expect PD models to be explainable. That’s why techniques like SHAP (SHapley Additive exPlanations) have become standard for justifying ML-based PD outputs to auditors and regulators.
Common Mistakes in PD Modeling
Treating bureau score as a sufficient standalone predictor without controlling for portfolio-specific behavior
Ignoring population stability — a PD model trained on pre-pandemic data can be badly miscalibrated for a post-pandemic portfolio
Failing to account for right-censoring when using a fixed observation window
Overfitting on a small default sample, which is common in low-default portfolios like corporate or sovereign lending
Loss Given Default (LGD) and Recovery
PD tells you whether a loss will happen. LGD tells you how big it will be once it does. Most practitioners consider it the harder of the two to model well. Default events are relatively rare, and recovery processes can take years to conclude.
What Drives LGD
LGD is shaped primarily by three factors:
Collateral coverage and quality. A secured home loan with a well-documented, liquid property as collateral typically carries a far lower LGD than an unsecured personal loan. The reason: the bank has a tangible asset to recover value from.
Seniority of claim. In corporate lending, senior secured lenders recover more than subordinated or unsecured creditors in a resolution or liquidation.
Recovery mechanism and timeline. In India, recovery routes include SARFAESI Act enforcement, Debt Recovery Tribunals, and the Insolvency and Bankruptcy Code (IBC). These matter enormously for LGD estimation. IBC-driven recoveries have averaged around 94% of the fair value of resolved businesses, though considerably less against the original claim amount. System-wide NPA recovery rates for scheduled commercial banks have roughly doubled, from 13.2% in FY18 to 26.2% in FY25, reflecting stronger legal recovery infrastructure.
The Recovery Rate Relationship
LGD and recovery rate are two sides of the same coin:
LGD = 1 − Recovery Rate
That relationship is correct. But it’s easy to misapply, because “recovery rate” is doing a lot of work in that equation. It has to mean the economic recovery rate: the present value of net recoveries, after two adjustments. First, discount every recovered cash flow back to the default date, to account for the time value of money. Second, subtract the direct and indirect costs of recovery — legal fees, collateral liquidation costs, workout team overhead. A workout can take two to three years in India, even under IBC timelines. A rupee recovered in year three is worth meaningfully less than a rupee recovered on day one.
The common mistake: using the nominal recovery rate instead. That’s raw cash eventually recovered, divided by exposure, with no discounting and no cost deduction. It overstates the recovery rate and, by direct consequence, understates LGD. Say a bank nominally recovers 65% of exposure post-default, but that recovery arrives over three years and costs 8% of exposure in legal and liquidation expenses. The economic recovery rate falls well below 65%. The resulting LGD lands correspondingly higher than the naive “35%” the nominal figure would suggest.
Collateral Valuation
For secured lending, LGD modeling starts with realistic collateral valuation — not the value at origination, but the expected value at the time of liquidation, discounted for:
Market depreciation of the asset class (property, equipment, vehicles)
Haircut for forced-sale conditions versus fair market value
Time-to-recovery, since a three-year legal process erodes present value even if the nominal recovery is high
Direct recovery costs (legal fees, auctioneer fees, administrative costs)
Statistical Methods for LGD
Workout LGD (the standard approach): Banks track every actual default in their historical data. They record all cash flows recovered post-default: collateral sale proceeds, settlement payments, guarantee invocations. Then they discount those cash flows back to the default date, using an appropriate discount rate, and net off recovery costs. This produces an empirical, account-level LGD. Analysts then average and segment it by product, collateral type, and vintage.
Regression-based LGD models: Raw workout LGD is bounded between 0 and 1. It can occasionally exceed 1, when recovery costs surpass recoveries. Banks often use techniques suited to bounded outcomes instead. One option is beta regression. Another is a two-stage model: first predict whether any recovery occurs at all, then model the recovery amount conditional on recovery happening.
Downturn LGD: Basel requires banks to estimate LGD under economic downturn conditions, not just average conditions, because collateral values and recovery rates both tend to fall exactly when default rates rise. This is a critical, frequently underestimated requirement. An LGD model calibrated only on benign-cycle data will understate loss severity in a stress scenario.
A Practical Note on LGD in Indian Retail Lending
For Indian home loans, LGD modeling relies heavily on loan-to-value (LTV) ratio at origination and current LTV, adjusted for property price movements. Property is the dominant recovery source here. Unsecured personal loans and credit cards work differently — they carry substantially higher and more volatile LGD, since recovery depends almost entirely on borrower cooperation, collection agency effectiveness, or write-off. That’s part of why the RBI’s 2023 risk-weight increase on unsecured credit targeted loss severity concerns explicitly. <h2id=”exposure-at-default”>Exposure at Default (EAD) Calculation
EAD often gets treated as the “simple” component of the PD × LGD × EAD formula, but any product with a revolving or undrawn component needs its own careful modeling.
Why EAD Isn’t Just the Current Balance
For a fully drawn term loan, EAD is straightforward. It stays close to the outstanding principal balance at any point in time, adjusted for scheduled amortization. Products like credit cards, overdrafts, and cash credit facilities work differently. A borrower can draw down additional funds between the assessment date and the moment of default. A borrower approaching financial distress often draws their credit line closer to the limit right before defaulting. That means EAD tends to run higher than the current outstanding balance, for exactly the accounts where it matters most.
The Credit Conversion Factor (CCF)
To capture this, banks use the Credit Conversion Factor (CCF) — the proportion of the currently undrawn commitment they expect to see drawn down before default.
EAD = Current Outstanding Balance + (CCF × Undrawn Commitment)
Worked example: A borrower has a credit card with a sanctioned limit of ₹5,00,000. The current outstanding balance is ₹2,00,000, leaving an undrawn commitment of ₹3,00,000. Historical data shows that, on average, borrowers who eventually default draw down 60% of their remaining undrawn limit before the default event (CCF = 0.60).
This runs meaningfully higher than the ₹2,00,000 current balance. The expected loss calculation should use the ₹3,80,000 figure, not the current balance.
How CCF Is Estimated
Banks typically estimate CCF empirically, using a cohort approach. They identify accounts that defaulted. They look back at each account’s utilization level 12 months before default. Then they measure how much of the then-undrawn limit got drawn down by the time of default. Averaging this ratio across the portfolio — segmented by product type and current utilization band — produces the CCF estimates that feed EAD models.
CCF runs highest for revolving retail products like credit cards and overdrafts. It runs near-zero for term loans with no undrawn component. Corporate revolving credit facilities and working capital limits also require dedicated CCF modeling. Distressed corporate borrowers frequently draw down committed but unused credit lines as a liquidity buffer before default becomes evident.
EAD Under Basel and IFRS 9
Under the Basel Internal Ratings-Based (IRB) approach, EAD estimation follows the same downturn-conditioning logic as LGD. CCFs should reflect what happens under stressed conditions, since utilization tends to spike precisely when the broader environment deteriorates. Under IFRS 9, EAD projections also need to extend across the lifetime of the facility for Stage 2 and Stage 3 exposures, not just a fixed 12-month window.
Real-World Implementation in Banking
The theory behind PD, LGD, and EAD only matters if it translates into disciplined implementation. India’s banking sector over the past decade offers a clear illustration of what that looks like at scale.
The Turnaround in Numbers
Following the RBI’s 2015 Asset Quality Review, public sector banks’ gross NPA ratio rose sharply, as transparent recognition brought hidden stress to light. It peaked at 14.58% in March 2018. What followed was a sustained, model-driven cleanup. Recapitalization, tighter underwriting standards, and the Insolvency and Bankruptcy Code combined to bring the PSB gross NPA ratio down to 9.11% by March 2021, and further to 2.58% by March 2025. System-wide, across all scheduled commercial banks, the gross NPA ratio reached a multi-decade low of around 1.8%–2.15% through late 2025 and into 2026, per the RBI’s Financial Stability Report.
At the institution level, this shows up clearly in individual bank results. For the quarter ended March 2026, State Bank of India — India’s largest lender — reported a gross NPA ratio of 1.49% and a net NPA ratio of 0.39%. HDFC Bank reported a gross NPA ratio of 1.15% and net NPA ratio of 0.38%. These aren’t accidents of a benign credit cycle alone. They reflect years of investment: early warning systems, scorecard-based underwriting, stressed-asset monitoring infrastructure.
What Changed Operationally
Three implementation shifts are consistently cited across the sector:
Early Warning Systems (EWS). Public sector banks have rolled out automated EWS frameworks with roughly 80 distinct triggers. These pull in third-party data to flag stress in borrowing accounts before they slip into NPA status. This shifts credit risk management from reactive classification to proactive monitoring.
Machine learning-augmented scorecards. Industry research examining major Indian banks — SBI, HDFC, ICICI, Kotak Mahindra — found that machine learning techniques consistently outperform pure ratio-based or bureau-score-only assessments. The winning combination: logistic regression alongside random forests and neural networks, incorporating alternative data like transaction behavior and digital usage patterns. It works especially well for borrowers with thin credit files.
Faster, more granular underwriting. The same research found loan approval times at many Indian banks have compressed, from several days to a matter of minutes for eligible segments. Automated, model-based decisioning drives this, rather than manual file review. That shift only became possible once PD and exposure models earned enough validated trust to run with minimal manual override.
The Governance Layer
None of this works without governance. The RBI’s ongoing regulatory review — including its recently articulated principle-based framework for AI use in banking — signals something important. As banks lean further into ML-driven credit risk models, model validation, explainability, and monitoring will matter as much as raw predictive accuracy. A model that can’t be explained to an auditor or regulator can’t be deployed at scale in a regulated balance sheet, no matter how well it performs statistically.
Interview Questions to Test Your Understanding
Write out the expected loss formula and explain what each component represents.
Why is logistic regression still preferred over more complex ML models for regulatory PD models?
What’s the difference between workout LGD and downturn LGD, and why does Basel require the latter?
Explain why EAD for a credit card is typically higher than the current outstanding balance.
How would you validate a PD model’s performance? (Hint: think Gini coefficient, KS statistic, and calibration testing.)
Summary
Credit risk modeling breaks down a complex question: how much could a bank lose, and how likely is it? It splits that question into three independently estimated, rigorously validated components — Probability of Default, Loss Given Default, and Exposure at Default — combined through EL = PD × LGD × EAD. Analysts typically estimate PD through logistic regression and survival analysis, increasingly supplemented by explainable machine learning. LGD depends on collateral quality, recovery mechanisms, and downturn conditions. EAD requires modeling credit conversion factors for any revolving exposure. Together, these three parameters drive capital adequacy, IFRS 9 provisioning, and loan pricing. India’s banking sector’s asset quality turnaround over the past decade proves the point: disciplined credit risk modeling delivers at scale.
Frequently Asked Questions
Q1. What is the difference between credit risk and credit risk modeling?
Credit risk is the underlying possibility of borrower default. Credit risk modeling is the quantitative discipline of measuring that risk. It uses statistical and machine learning techniques to estimate PD, LGD, and EAD, so a bank can price, provision for, and manage the risk.
Q2. Is credit risk modeling the same as credit scoring?
Credit scoring is a subset of credit risk modeling, focused specifically on PD estimation for underwriting decisions. Full credit risk modeling also covers LGD, EAD, portfolio-level expected loss, stress testing, and regulatory capital calculation.
Q3. Which technique is better for PD modeling: logistic regression or machine learning?
Neither wins universally. Logistic regression remains preferred for regulatory capital models, thanks to its interpretability and regulator familiarity. Machine learning models can improve accuracy, particularly with alternative data. But they require additional explainability tooling — like SHAP — to meet governance and audit requirements.
Q4. What skills do I need to build a career in credit risk modeling?
You need statistics (logistic regression, survival analysis), proficiency in Python or SAS, familiarity with Basel and IFRS 9, and hands-on exposure to real credit datasets. Structured training that pairs regulatory context with practical model-building is generally the fastest path in.
Q5. How is credit risk modeling connected to IFRS 9?
IFRS 9 requires banks to recognize expected credit losses using forward-looking PD, LGD, and EAD estimates, rather than waiting for an actual default. Stage 1 assets use a 12-month figure, close to the standard EL = PD × LGD × EAD calculation. Stage 2 and 3 assets use lifetime ECL instead — a discounted sum of marginal expected losses across the loan’s remaining life, not a single multiplication.
Ready to Build These Skills Hands-On?
Understanding the theory behind PD, LGD, and EAD is the first step. Building bankable, interview-ready models — in Python or SAS, on real credit datasets, aligned to Basel and IFRS 9 — is what actually moves a career forward.
Explore Dexlab Analytics’ Credit Risk Modeling certification program to build PD, LGD, and EAD models from scratch, work through IFRS 9 ECL frameworks, and learn model validation techniques used by practicing risk teams.
In our previous blog we discussed about few of the basic functions of MQL like .find() , .count() , .pretty() etc. and in this blog we will continue to do the same. At the end of the blog there is a quiz for you to solve, feel free to test your knowledge and wisdom you have gained so far.
Given below is the list of functions that can be used for data wrangling:-
updateOne() :- This function is used to change the current value of a field in a single document.
After changing the database to “sample_geospatial” we want to see what the document looks like? So for that we will use .findOne() function.
Now lets update the field value of “recrd” from ‘ ’ to “abc” where the “feature_type” is ‘Wrecks-Visible’.
Now within the .updateOne() funtion any thing in the first part of { } is the condition on the basis of which we want to update the given document and the second part is the changes which we want to make. Here we are saying that set the value as “abc” in the “recrd” field . In case you wanted to increase the value by a certain number ( assuming that the value is integer or float) you can use “$inc” instead.
2. updateMany() :- This function updates many documents at once based on the condition provided.
3. deleteOne() & deleteMany() :- These functions are used to delete one or many documents based on the given condition or field.
4. Logical Operators :-
“$and” : It is used to match all the conditions.
“$or” : It is used to match any of the conditions.
The first code matches both the conditions i.e. name should be “Wetpaint” and “category_code” should be “web”, whereas the second code matches any one of the conditions i.e. either name should be “Wetpaint” or “Facebook”. Try these codes and see the difference by yourself.
So, with that we come to the end of the discussion on the MongoDB Basics. Hopefully it helped you understand the topic, for more information you can also watch the video tutorial attached down this blog. The blog is designed and prepared by Niharika Rai, Analytics Consultant, DexLab AnalyticsDexLab Analytics offers machine learning courses in Gurgaon. To keep on learning more, follow DexLab Analytics blog.
MongoDB is a document based database program which was developed by MongoDB Inc. and is licensed under server side public license (SSPL). It can be used across platforms and is a non-relational database also known as NoSQL, where NoSQL means that the data is not stored in the conventional tabular format and is used for unstructured data as compared to SQL and that is the major difference between NoSQL and SQL. MongoDB stores document in JSON or BSON format. JSON also known as JavaScript Object notation is a format where data is stored in a key value pair or array format which is readable for a normal human being whereas BSON is nothing but the JSON file encoded in the binary format which is quite hard for a human being to understand. Structure of MongoDB which uses a query language MQL(Mongodb query language):- Databases:- Databases is a group of collections. Collections:- Collection is a group fields. Fields:- Fields are nothing but key value pairs Just for an example look at the image given below:-
Here I am using MongoDB Compass a tool to connect to Atlas which is a cloud based platform which can help us write our queries and start performing all sort of data extraction and deployment techniques. You can download MongoDB Compass via the given link https://www.mongodb.com/try/download/compass
In the above image in the red box we have our databases and if we click on the “sample_training” database we will see a list of collections similar to the tables in sql.
Now lets write our first query and see what data in “companies” collection looks like but before that select the “companies” collection.
Now in our filter cell we can write the following query:-
In the above query “name” and “category_code” are the key values also known as fields and “Wetpaint” and “web” are the pair values on the basis of which we want to filter the data. What is cluster and how to create it on Atlas? MongoDB cluster also know as sharded cluster is created where each collection is divided into shards (small portions of the original data) which is a replica set of the original collection. In case you want to use Atlas there is an unpaid version available with approximately 512 mb space which is free to use. There is a pre-existing cluster in MongoDB named Sandbox , which currently I am using and you can use it too by following the given steps:- 1. Create a free account or sign in using your Google account on https://www.mongodb.com/cloud/atlas/lp/try2-in?utm_source=google&utm_campaign=gs_apac_india_search_brand_atlas_desktop&utm_term=mongodb%20atlas&utm_medium=cpc_paid_search&utm_ad=e&utm_ad_campaign_id=6501677905&gclid=CjwKCAiAr6-ABhAfEiwADO4sfaMDS6YRyBKaciG97RoCgBimOEq9jU2E5N4Jc4ErkuJXYcVpPd47-xoCkL8QAvD_BwE 2. Click on “Create an Organization”. 3. Write the organization name “MDBU”. 4. Click on “Create Organization”. 5. Click on “New Project”. 6. Name your project M001 and click “Next”. 7. Click on “Build a Cluster”. 8. Click on “Create a Cluster” an option under which free is written. 9. Click on the region closest to you and at the bottom change the name of the cluster to “Sandbox”. 10. Now click on connect and click on “Allow access from anywhere”. 11. Create a Database User and then click on “Create Database User”. username: m001-student password: m001-mongodb-basics 12. Click on “Close” and now load your sample as given below :
Loading may take a while…. 13. Click on collections once the sample is loaded and now you can start using the filter option in a similar way as in MongoDB Compass In my next blog I’ll be sharing with you how to connect Atlas with MongoDB Compass and we will also learn few ways in which we can write query using MQL.
So, with that we come to the end of the discussion on the MongoDB. Hopefully it helped you understand the topic, for more information you can also watch the video tutorial attached down this blog. The blog is designed and prepared by Niharika Rai, Analytics Consultant, DexLab AnalyticsDexLab Analytics offers machine learning courses in Gurgaon. To keep on learning more, follow DexLab Analytics blog.