FUTURECENTRAL PRESS · BOOK SAMPLE
Python Fundamentals for Finance
From Your First Line of Code to a Complete Analytical Project
Chapter 9: A Complete Worked Project
From data download to written interpretation: a notebook another analyst can re-run and a committee can review
Learning Outcomes
By the end of this chapter, you will be able to:
Specify a finance-domain analytical project against a structured project brief and a deliverable definition.
Choose between credit-, investments-, and risk-domain project universes based on your career direction and the data available to you.
Execute a complete end-to-end notebook from data download, through cleaning and EDA, through a small modeling step, into a written interpretation.
Produce a professional-quality Jupyter or Colab notebook that another analyst can re-run on the same inputs and reach the same conclusions.
Write a five-paragraph analytical summary that translates the notebook’s charts and tables into a memo a committee can act on.
Evaluate your own analysis against a self-directed rubric covering data, method, output, and reproducibility.
Opening Vignette
Picture a New Analyst at a Chennai-based corporate-finance advisory firm who was given his first standalone project assignment.²² The brief was open-ended in the way most first projects are. Pick an area of finance you are interested in. Pull together public data. Run the analyses you have learned in the company’s induction program. Produce a deliverable that another team member can read and a partner can sign off on. Two weeks. No vendor data. No ground-truth answer.
He spent the first morning choosing the universe. The firm’s practice areas spanned mid-corporate credit advisory, family-office investment counsel, and an operational-risk consulting service for mid-sized banks. Each had its own data-and-analysis pattern. He chose credit because his undergraduate background was in commerce and the credit-domain analytical artefacts (vintage default curves, segmentation tables, transition matrices) were the ones he had practised most in the firm’s induction. The choice halved the second-day workload by aligning the project with the analytical patterns he could already produce fluently.
The first week went into downloading and cleaning. The LendingClub historical loan-level archive, the one Chapter 7 introduced as the closest publicly-available analogue to a real loan portfolio, gave him a million rows of US peer-to-peer-lending data spanning 2007 through 2018, with origination, performance, and default fields included. He filtered to a manageable subset of consumer loans originated between 2014 and 2017, applied the defensive read_csv template from Chapter 5, normalized the column names, and ran the five-layer EDA pass from Chapter 8 against the cleaned table. By the end of week one the data was loaded, validated, and described.
The second week went into the analysis and the write-up. The five charts in his notebook (loan-amount distribution, default-rate-by-loan-purpose bar chart, default-rate-by-credit-grade bar chart, vintage-default trajectory, and a correlation heatmap of borrower-level features) answered the questions the brief implicitly asked: who defaults, when, and what features differentiate defaulters from non-defaulters. A five-paragraph written summary translated the charts into actionable observations for the firm’s credit-advisory practice, and the deliverable was reviewed without major revisions; the New Analyst was added to the firm’s credit-advisory client team the following month.
By the end of this chapter, the project pipeline that took the New Analyst two weeks will be a working template you can deploy in a single weekend on data of your own choosing.
9.1 The Project Brief
Every analytical project (at a firm, in a course, on a job application) has the same five-element structure. Producing the brief is the first step before the project touches any data, and is the discipline that prevents two-week projects from becoming six-week projects.
The five elements
Question. The single substantive question the project answers, written as a sentence ending in a question mark. "Which loan-portfolio segments carry the highest default risk?" "How has the cross-asset correlation structure between Indian and US equities evolved over five years?" "What is the operational-loss tail behavior in the bank’s payments-business line?"
Universe. The data the analysis covers: the time window, the asset list or borrower list, the loss series. "LendingClub consumer loans originated 2014-2017, US." "Daily closes for ten Indian large-caps and four US large-caps, 2019-2024." "Operational-loss incidents above ₹1 lakh in the bank’s payments business line, FY 2022-24."
Method. The analytical approach: at this stage of the reader’s training, this is the EDA workflow of Chapter 8 plus an optional simple modeling step. "Five-layer EDA on the loan portfolio plus a logistic-regression default-prediction model." "Tri-asset EDA tear sheet plus a rolling-correlation regime analysis."
Deliverables. The artefacts the project produces: typically a notebook, a memo, and a small set of charts. "A Jupyter notebook (.ipynb), a five-page memo (.docx), and three publication-quality charts (.png at 150 dpi)."
Audience. The reader who will consume the deliverable: a partner, a board, an interview panel, a course evaluator. The audience determines the tone of the memo, the level of technical detail, and the length.
Write the brief on a single page before opening any tool, and well before the data download begins. The discipline is the same one Chapter 6 applied to chart selection: identify the question first, then everything else.
9.2 Three Universes to Choose From
Three project universes cover the three finance domains this book has worked with. Choose the one that fits your career direction, the data you have access to, and the analytical patterns you are most fluent in. The next three sections work through one example in each universe.
Credit universe
The credit universe is built around a loan-portfolio dataset. Public sources include the LendingClub historical archive (Chapter 7 endnote 17), the Kaggle credit-risk competitions, and the academic mirrors that several universities maintain. The LendingClub archive is the open dataset closest to a real loan portfolio that an analyst can practise on. The analytical artefacts the project produces include the loan-amount distribution, the default-rate-by-segment table, the vintage-default trajectory, and an optional logistic-regression default-prediction model. The audience is typically a credit-advisory practice, a banking-strategy team, or an analytics-for-credit interview panel.
Investments universe
The investments universe is built around a multi-asset return panel. Public sources include yfinance, NSEpy, and FRED, the four-source pipeline of Chapter 7. The analytical artefacts include the return-distribution overlay, the rolling-volatility chart, the rolling-correlation chart, and the cross-asset correlation matrix. The audience is typically an asset-management research team, a wealth-advisory practice, or a sell-side equity-research interview panel.
Risk universe
The risk universe is built around an operational-loss series or a market-loss series. Public sources are sparser (see Chapter 7 Section 7.8) and many projects in this universe use simulated data calibrated against published industry aggregates. The analytical artefacts include the loss-distribution histogram, the frequency-severity time series, the peaks-over-threshold tail analysis, and a Monte Carlo VaR estimate. The audience is typically a market-risk team, an operational-risk team, or a regulatory-stress-testing role.
How to choose
Three considerations point to the right universe. Career direction is the most important: pick the universe that matches the role you want next, because the deliverable doubles as a portfolio piece. Data access is the second: the universe with the cleanest available data will produce the best deliverable in the time available. Pattern fluency is the third: the universe whose analytical artefacts you can produce most quickly leaves more time for the interpretation step that distinguishes a good project from a competent one.
9.3 A Worked Example: Credit Portfolio Analysis
This worked example uses the LendingClub historical loan-level archive, the closest open analogue to a real loan portfolio that is freely available. The pattern reproduces what the New Analyst built in the Opening Vignette.
Step 1: Download and load
import pandas as pd
import numpy as np
URL = ("https://futurecentral.in/datasets/"
"credit/lendingclub_2014_2017.csv.gz")
loans = pd.read_csv(
URL,
parse_dates=["issue_d"],
na_values=["NA", "NULL", "n/a", ""],
dtype={"id": str, "member_id": str, "grade": str, "sub_grade": str},
compression="gzip",
)
loans.columns = loans.columns.str.strip().str.lower()
print(f"Loaded {len(loans):,} loans across {loans.shape[1]} columns.")Step 2: Inspect and define the default flag
print(loans["loan_status"].value_counts())
# Standard LendingClub default convention. Charged off, default, or 31-120 day late
default_statuses = ["Charged Off", "Default",
"Late (31-120 days)", "Does not meet the credit policy. Status:Charged Off"]
loans["default_flag"] = loans["loan_status"].isin(default_statuses).astype(int)
print(f"Portfolio default rate: {loans['default_flag'].mean():.2%}")The default-flag definition is itself a methodological choice that should be documented in the notebook. LendingClub publishes several status categories; the analyst groups them into binary defaulted/not-defaulted classifications according to the project’s working definition. The literature has converged on the four-category grouping above as the standard convention.¹
Step 3: Run the EDA workflow
# Layer 1: profile
print(loans[["loan_amnt", "int_rate", "annual_inc", "dti"]].describe().round(2))
print(loans["purpose"].value_counts(normalize=True).round(3))
print(loans["grade"].value_counts(normalize=True).round(3))
# Layer 2: distribution
import matplotlib.pyplot as plt
import seaborn as sns
sns.set_theme(style="whitegrid")
fig, axes = plt.subplots(1, 3, figsize=(14, 4))
axes[0].hist(loans["loan_amnt"], bins=50, edgecolor="black")
axes[0].set_title("Loan amount ($)")
axes[1].hist(loans["int_rate"], bins=40, edgecolor="black")
axes[1].set_title("Interest rate (%)")
axes[2].hist(loans["dti"].clip(0, 50), bins=50, edgecolor="black")
axes[2].set_title("Debt-to-income ratio (%)")
plt.tight_layout()
plt.show()
# Layer 3: Time series - vintage default
loans["issue_quarter"] = loans["issue_d"].dt.to_period("Q")
vintage = (loans
.groupby("issue_quarter", observed=True)["default_flag"]
.agg(n_loans="count", default_rate="mean")
.round({"default_rate": 4}))
print(vintage)
# Layer 4: dependence -- default rate by purpose and grade
by_purpose = (loans.groupby("purpose")["default_flag"]
.mean().sort_values(ascending=False).round(4))
by_grade = (loans.groupby("grade")["default_flag"]
.mean().sort_values().round(4))
print(by_purpose)
print(by_grade)Step 4: A simple default-prediction model
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
# Minimal feature set. Full version belongs to Chapter 10
features = ["loan_amnt", "int_rate", "annual_inc", "dti"]
X = loans[features].dropna()
y = loans.loc[X.index, "default_flag"]
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=42)
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1])
print(f"Test AUC: {auc:.3f}")A four-feature logistic-regression model fitted to predict default on this dataset typically lands at an AUC between 0.65 and 0.70. Modest but non-trivial.² Higher AUCs in published academic work usually involve more features (FICO score, employment history, geography), longer training samples, or richer model classes. Chapter 10 covers the next steps;¹¹ the more demanding model classes belong to follow-on courses. The interpretation paragraph in the memo should explain what the AUC means and what its level implies for the practical use of the model.¹² For credit projects with class imbalance, the analyst typically reports the precision-recall curve and the confusion matrix alongside the AUC.¹⁸
Step 5: Memo and chart artefacts
The credit-portfolio project ships with a five-paragraph memo and three named charts: the default-rate-by-purpose horizontal bar chart, the vintage-default-rate line chart, and the credit-grade-default-rate horizontal bar chart. Section 9.6 covers the memo structure.
9.4 A Worked Example: Equity Portfolio EDA
This worked example uses the multi-source data pipeline of Chapter 7 (yfinance for global equities, plus an optional FRED macro overlay) to build a tri-asset cross-comparison artefact. The pattern reproduces and extends the tear sheet from the Chapter 8 Python Lab.
Step 1: Download and prepare
import yfinance as yf
import pandas as pd
import numpy as np
tickers = ["RELIANCE.NS", "TCS.NS", "HDFCBANK.NS", "INFY.NS", "ICICIBANK.NS",
"AAPL", "MSFT", "JPM", "GOOGL"]
panel = {}
for t in tickers:
df = yf.download(t, start="2019-01-01", end="2024-12-31",
auto_adjust=True, progress=False)
panel[t] = df["Close"]
prices = pd.DataFrame(panel).dropna()
returns = prices.pct_change().dropna()Step 2: Run the EDA workflow
# Layer 1-2: descriptive statistics
summary = pd.DataFrame({
"ann_return": returns.mean() * 252,
"ann_vol": returns.std() * np.sqrt(252),
"skew": returns.skew(),
"kurtosis": returns.kurtosis(),
"sharpe": (returns.mean() / returns.std()) * np.sqrt(252),
}).round(3)
print(summary)
# Layer 3: Rolling volatility
rolling_vol = returns.rolling(60).std() * np.sqrt(252)
# Layer 4: correlation matrix
corr = returns.corr()
# Layer 5: drawdown
def drawdown(s):
cum = (1 + s).cumprod()
return cum / cum.cummax() - 1
dd = returns.apply(drawdown)
max_dd = dd.min().round(3)
print(f"\nMax drawdown by ticker:\n{max_dd}")Step 3: Cross-asset comparison panel
Build the six-panel tear sheet from the Chapter 8 Python Lab (cumulative return, distribution overlay, rolling volatility, rolling correlation, drawdown, and full-sample correlation heatmap) using the matplotlib subplot pattern of Chapter 6. The output is a single PNG that summarizes the equity panel’s risk-and-return profile in one screen.
Step 4: Optional macro overlay
import pandas_datareader.data as pdr
macro = pdr.DataReader(
["CPIAUCSL", "FEDFUNDS", "T10Y2Y"],
"fred",
start="2019-01-01",
end="2024-12-31",
).ffill()
# Joined panel. Returns and macro on the same date index
joined = returns.join(macro, how="left").ffill()
# Correlation between portfolio return and federal funds rate
portfolio_ret = returns.mean(axis=1) # equal-weighted portfolio
print(f"Portfolio return vs. Fed funds rate: "
f"{portfolio_ret.corr(joined['FEDFUNDS']):.3f}")
print(f"Portfolio return vs. CPI: "
f"{portfolio_ret.corr(joined['CPIAUCSL'].pct_change()):.3f}")The macro overlay¹⁹ extends the basic equity-EDA project into a cross-asset macro context. Useful when the project audience is interested in the relationship between portfolio behavior and the broader interest-rate or inflation cycle. The overlay is optional; for a self-contained equity project it can be skipped.
Step 5: Memo and chart artefacts
The equity-portfolio project ships with a memo and three named charts: the cumulative-return time series with rebased base date, the rolling-volatility comparison line chart, and the cross-asset correlation heatmap. The risk-and-return summary table sits in the memo as a second-page reference.
9.5 A Worked Example: Portfolio Risk Analysis
This worked example is the Monte Carlo VaR exercise of Chapter 4, scaled into a complete project with cleaning, calibration, simulation, validation, and a written interpretation. The pattern is the production-grade version of the Chapter 4 Python Lab.
Step 1: Calibrate from history
import yfinance as yf
import pandas as pd
import numpy as np
tickers = ["RELIANCE.NS", "TCS.NS", "HDFCBANK.NS", "INFY.NS", "ICICIBANK.NS"]
prices = yf.download(tickers, start="2019-01-01", end="2024-12-31",
auto_adjust=True, progress=False)["Close"]
returns = prices.pct_change().dropna()
# Annualized mean and volatility per asset
ann_means = returns.mean() * 252
ann_vols = returns.std() * np.sqrt(252)
# Correlation matrix from the historical sample
corr = returns.corr()
print(ann_means.round(3))
print(ann_vols.round(3))
print(corr.round(2))Step 2: Build the Cholesky factor
daily_means = (ann_means / 252).values
daily_vols = (ann_vols / np.sqrt(252)).values
# Covariance = D @ R @ D
D = np.diag(daily_vols)
covariance = D @ corr.values @ D
# Cholesky factor. See Chapter 4 Python Lab
L = np.linalg.cholesky(covariance)Step 3: Run the Monte Carlo
rng = np.random.default_rng(seed=42)
n_scenarios = 100_000
weights = np.array([0.25, 0.20, 0.20, 0.15, 0.20])
portfolio_value = 1e8 # Rs. 10 crore
shocks = rng.normal(size=(n_scenarios, len(tickers))) @ L.T
asset_returns = daily_means + shocks
portfolio_returns = asset_returns @ weights
pnl = portfolio_value * portfolio_returns
var_95 = -np.percentile(pnl, 5)
var_99 = -np.percentile(pnl, 1)
es_99 = -pnl[pnl <= -var_99].mean()
print(f"95% one-day VaR: Rs. {var_95:,.0f}")
print(f"99% one-day VaR: Rs. {var_99:,.0f}")
print(f"99% one-day ES: Rs. {es_99:,.0f}")Step 4: Validation against historical
# Historical-simulation cross-check
hist_pnl = portfolio_value * (returns @ weights)
hist_var_99 = -np.percentile(hist_pnl, 1)
print(f"Historical 99% VaR: Rs. {hist_var_99:,.0f}")Reconciling the Monte Carlo and the historical-simulation VaR estimates is the validation step the project must include. A wide divergence between the two (say, more than twenty percent) is the signal to inspect the calibration assumptions: the lookback window, the normality assumption that the parametric simulation rests on, and the stationarity of the correlation matrix.
Step 5: Sensitivity analysis
A risk-project memo without a sensitivity analysis is incomplete. The two sensitivities the audience will ask about: how does the VaR change when the correlation matrix is shocked toward one (the stress-correlation case), and how does the VaR change when the volatility inputs are scaled by one-and-a-half or two (the stress-volatility case). Each sensitivity is a five-line modification to the simulation code; the memo reports the headline number from each scenario.
Step 6: Memo and chart artefacts
The risk-portfolio project ships with a five-paragraph memo and three named charts. The simulated portfolio P&L distribution with VaR and ES annotation lines, the historical-versus-simulated VaR comparison bar chart, and the sensitivity-analysis table rendered as a heatmap. The memo concludes with a recommendation on the VaR number to adopt for the position-limit framework.
9.6 Interpretation and Written Summary
The notebook is the analytical artefact.¹⁵ The memo is the deliverable. A senior reader, an interview panel, or a partner reads the memo first; the notebook is consulted only if the memo raises a question. The discipline of writing the memo well is what separates a competent project from a useful one.
The write-up structure
Paragraph 1: The question and the headline answer. State the question the project answered and the one-line headline finding. "Default rates across the LendingClub 2014-2017 vintage vary from 6.2 percent for the lowest-risk loans to 22.4 percent for the highest-risk loans, with loan purpose and credit grade explaining most of the variation."
Paragraph 2: The data and the method. Identify the dataset, the time window, the universe size, and the analytical method. Cite the source, declare the cleaning steps, and acknowledge any limitations of the data. "The analysis uses the LendingClub public archive, restricted to consumer loans originated 2014-2017. The five-layer EDA pass of Chapter 8 was applied; an optional logistic-regression default-prediction model was fitted as a supplementary step."
Paragraph 3: The most consequential finding, in detail. The single most analytically important number or pattern, with the chart that supports it. "The default-rate-by-credit-grade chart (Figure 1) shows monotonic deterioration from grade A through grade G,¹³ with the gradient steepest between grades C and D. The pattern is consistent across vintages and is the strongest single predictor in the analysis."
Paragraph 4: The supporting findings. Two or three secondary results, each one or two sentences. The supporting findings should reinforce, qualify, or contextualise the primary one. "The vintage-default trajectory (Figure 2) shows that 2016 and 2017 cohorts deteriorated relative to 2014 and 2015, suggesting underwriting-quality drift over the sample. The default-rate-by-purpose breakdown (Figure 3) shows debt-consolidation and small-business loans concentrated in the higher-risk grades."
Paragraph 5: The implication and the limitation. What the audience should take away, and what the analysis cannot say. The implication is the call-to-action; the limitation is the honest qualification that marks a mature analyst’s output. "The findings support a tightening of underwriting in the C-and-below grades for the 2016-2017 vintages. The analysis cannot speak to the post-Covid period (the dataset ends in 2018) and the conclusions should be revalidated against current loan performance before any policy change is made."
The structure is portable across credit, investments, and risk projects. The headline-finding sentence in Paragraph 1 is the one the audience will remember; spend disproportionate time on it.
9.7 Self-Evaluation Rubric
Before submitting the deliverable, run the notebook against the rubric below. Each row is a yes-or-no check; any "no" is a candidate for further work before the deliverable leaves your hands.
Data and provenance
The notebook’s opening cell records the source URL, the date of pull, and the version of the wrapper library used.
Every column used in the analysis has had its dtype, range, and missing-value pattern inspected.
Outliers and anomalies have been inspected manually before any decision was made about them.
The data file or the cached parquet is included in the project folder and the path is relative, not absolute.
Method and reproducibility
The notebook runs end-to-end on a fresh kernel with no errors and no manual intervention.
All random seeds are set explicitly with np.random.default_rng(seed=…) or scikit-learn’s random_state argument.
Every methodological choice (the default-flag definition, the lookback window, the train-test split) is documented in a markdown cell next to the code that uses it.
The notebook is committed to a version-controlled folder with a meaningful commit message, even if the version control is local rather than remote.
Output and audience
Every chart has axis labels with units, a title that states the chart’s conclusion, a date range, and a source attribution.
The five-paragraph memo answers the question, identifies the data and method, presents the headline and supporting findings, and acknowledges a limitation.
The PNG charts have been saved at 150 dpi or higher and the PDF charts at vector quality.
Another analyst could take the project folder, open the README, run the notebook, and produce the same numbers and charts.
A project that passes every row of the rubric is one a senior reader will sign off on. A project that fails one or two rows is one the senior reader will return for revision. A project that fails five or more is one the senior reader will not read at all.
Framework: The Six-Phase Project Pipeline
This six-phase pipeline is the spine of the whole chapter. It is the working sequence the New Analyst followed in the Opening Vignette and the sequence the chapter applied across the three worked examples. The phases are sequential; the discipline is to complete the deliverables of each phase before starting the next.
Phase 1: Brief. Write the question, universe, method, deliverables, and audience on a single page. Cost: thirty minutes. Saves: between two and ten hours of mid-project rework. Skipping the brief is the single most common cause of two-week projects becoming six-week projects.
Phase 2: Data. Download, cache, and validate the source data using the defensive pipeline of Chapter 7. Document the source URL, the pull date, and any cleaning decisions in the notebook’s opening cells. Cost: a few hours. Saves: every downstream debugging session that would have traced back to a data-loading shortcut.
Phase 3: EDA. Apply the five-layer EDA pass of Chapter 8 to the cleaned data. Generate the profile, distribution, time-series, dependence, and outlier views. The output is the analytical understanding the rest of the project rests on.
Phase 4: Analysis. Execute the analytical method named in the brief: segmentation tables, rolling statistics, a small predictive model, a Monte Carlo simulation. The phase produces the headline numbers and the supporting charts that the memo will reference.
Phase 5: Memo. Write the brief of Section 9.6. Translate the notebook’s charts and tables into prose a non-technical reader can act on. The memo is the deliverable; the notebook is the supporting evidence.
Phase 6: Self-evaluation. Run the notebook against the rubric of Section 9.7. Fix every "no" before the deliverable leaves your hands. The phase costs an hour and is the difference between a deliverable a senior reader signs off on and one returned for revision.
Use this pipeline²¹ whenever you sit down to a finance-domain analytical project. The temptation under deadline pressure is to skip Phase 1, Phase 5, or Phase 6 and go straight from Data to Analysis. Running all six in order, even under deadline, is what keeps a two-week project from becoming a six-week one.
India Cases
Indian MBA finance capstone projects and the credit-portfolio pattern. MBA finance programs at the Indian Institutes of Management, the Indian School of Business, and other leading business schools routinely set capstone or term-paper projects on credit-portfolio analysis using public datasets.³ The analytical pattern the students apply is the credit-universe pattern of Section 9.3: load, EDA, segmentation, a small predictive model, written interpretation. The vocabulary varies by faculty; the structure does not. The deliverable on these capstones is typically a Jupyter notebook, a short memo, and a defense presentation; the rubric the faculty grade against is recognizably the rubric of Section 9.7.
Asset-management research-internship deliverables in India. The summer-internship programs at major Indian asset-management firms typically end with the intern presenting a research project to the firm’s investment committee.⁴ HDFC AMC, ICICI Prudential, SBI Mutual Fund, Axis AMC, and Nippon India Mutual Fund all run programs of this shape. The deliverables follow the investments-universe pattern of Section 9.4: a multi-asset return panel, a tear sheet, a written summary, and a recommendation. The patterns are stable across firms; an intern at one firm produces output a colleague at another firm would recognize immediately.
Risk-consulting client engagements at the major Indian advisories and the Monte Carlo VaR pattern. Risk-advisory engagements at the Indian arms of the major consulting firms (KPMG, PwC, EY, Deloitte) and at boutique Indian risk-consulting firms deliver Monte Carlo VaR analyses to mid-sized banks, NBFCs, and treasury teams as a routine offering.⁵ The Section 9.5 worked example reproduces the pattern of these engagements at smaller scale. Calibrate from history, build the Cholesky factor,¹⁶ simulate, validate against historical,¹⁷ run sensitivities, and write up. The senior consultant on the engagement reviews the deliverable against an internal checklist that closely resembles the Section 9.7 rubric.
Public-policy-research outputs and reproducibility standards. The Reserve Bank of India Occasional Papers, the SEBI Working Paper series, and the various NCAER, NIPFP, and ICRIER research outputs are all structurally the same kind of deliverable as the chapter’s worked examples.⁶ Each follows the data-EDA-analysis-interpretation arc. The reproducibility standards the publication processes apply (committed code, sourced data, documented methodology) match the standards the self-evaluation rubric of Section 9.7 prescribes. Reading a recent Occasional Paper alongside the rubric is a useful exercise in seeing the standards applied at a publication-grade level.
Global Cases
Kaggle competitions and the credit-modeling project canon. Kaggle, the data-science competition platform now owned by Google, has hosted dozens of credit-risk competitions since its founding.⁷ The Home Credit Default Risk competition, the LendingClub historical analyses, and the various credit-fraud-detection competitions form a large public archive of credit-modeling project deliverables. The winning solutions are published with code, data, and a written write-up, the closest open analogue to the credit-universe project this chapter teaches, run at the scale of a thousand-team competition.
CFA Institute Research Challenge and the equity-research project pattern. The CFA Institute Research Challenge is an annual global competition in which student teams produce an equity-research report on a publicly-listed company.⁸ The deliverable is a written report, an equity-research valuation model, and an oral defense. The investments-universe pattern of Section 9.4 (multi-asset EDA, return analysis, written interpretation) overlaps substantially with the CFA Research Challenge methodology, scaled from a single-asset valuation to a multi-asset cross-comparison.
The Federal Reserve Bank of New York and the Liberty Street Economics blog. The Federal Reserve Bank of New York’s Liberty Street Economics blog publishes worked analytical pieces on US and international macroeconomic-and-financial questions, with code and methodology notes alongside.⁹ The articles use the same five-element project structure this chapter teaches: question, universe, method, deliverable, audience. The scale is published-research rather than analyst-deliverable, but the structure is recognizably the same. They are a useful reference for any analyst building a macro-overlay project on top of the equity-EDA pattern of Section 9.4.
Practitioner's Lens
Imagine a Project Lead at a mid-sized Indian financial-analytics consulting firm in her fifth year of project-based delivery work. She runs a four-person team that ships two-to-three week analytical projects for a rotating client base: banks, asset-management firms, family offices, and the occasional fintech.²² The projects span all three finance domains; the structure is portable; the discipline is what makes the throughput possible.
Her project kickoff is the brief. The first deliverable on every engagement is a one-page project brief signed off by the client before any download begins. The brief specifies the question, the universe, the method, the deliverables, and the audience, the five elements of Section 9.1 of this chapter. The discipline has, in her experience, prevented every scope-creep argument she has had with a client over five years; when the conversation drifts mid-project, she returns to the brief.
Her data-and-cleaning week is non-negotiable. On any project longer than five working days, the first week is data: download, clean, validate, cache, document. The temptation to start the analysis on day two is the temptation she trains junior team members out of. The deliverable that emerges from a project where the data layer was rushed is one that fails the client’s reproducibility check. The deliverable that emerges from a project where the data layer was patient is one the client signs off on without revisions.
Her notebook discipline is markdown first, then code. Every section of every notebook her team ships opens with a markdown cell that states the question being answered, identifies the inputs, and previews the output. The code that follows is the implementation of what the markdown cell already described. The reader of the notebook (the client, the audit team, the colleague six months later) can navigate by reading the markdown alone and dropping into the code only where the implementation matters.
Her memo-writing time runs to a third of the project. She allocates roughly thirty percent of total project time to the memo and the chart preparation that supports it. The proportion is higher than most clients expect and is the conversation she has at the brief stage about the deliverable’s purpose. The memo is the artefact the client will read; the notebook is the evidence behind it; the proportion of time that goes into each reflects the proportion of attention each will receive on the receiving end.
Python Lab: A Reusable Project Notebook Template
This lab provides a reusable notebook template that students and analysts can clone for any of the three project universes. The template is the spine of every project in the chapter.
Cell 1: Project header (markdown)
# Project Brief
#
# Question: [one-sentence question]
# Universe: [data and time window]
# Method: [analytical approach]
# Deliverable: [notebook + memo + charts]
# Audience: [who will read the deliverable]
#
# Author: [name]
# Pull date: [YYYY-MM-DD]
# Library versions: pandas, numpy, matplotlib, seaborn, yfinance
Cell 2: Imports and configuration
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
sns.set_theme(style="whitegrid")
# Reproducibility
RNG = np.random.default_rng(seed=42)
# Path conventions
DATA_DIR = "data_cache"
OUTPUT_DIR = "outputs"
import os
os.makedirs(DATA_DIR, exist_ok=True)
os.makedirs(OUTPUT_DIR, exist_ok=True)Cell 3: Data load with caching
def cached_yf(ticker, start, end):
"""Defensive cached yfinance download (Chapter 7 pattern)."""
fname = f"{DATA_DIR}/{ticker.replace('=','_')}_{start}_{end}.parquet"
if os.path.exists(fname):
return pd.read_parquet(fname)
import yfinance as yf
df = yf.download(ticker, start=start, end=end,
auto_adjust=True, progress=False)
if not df.empty:
df.to_parquet(fname)
return dfCell 4: Five-layer EDA pass
def eda_pass(df, label="dataset"):
"""Run the Chapter 8 five-layer EDA pass on a numeric DataFrame or Series."""
print(f"=== {label} ===")
print(f"Shape: {df.shape}")
print(f"Date range: {df.index.min()} to {df.index.max()}" if hasattr(df.index, "min") else "")
print(f"Missing values:\n{df.isna().sum()}")
print(f"\nDescribe:\n{df.describe().round(4)}")
if isinstance(df, pd.DataFrame) and len(df.columns) > 1:
print(f"\nCorrelation:\n{df.corr().round(2)}")Cell 5: Analysis (project-specific)
# Project-specific analysis goes here.
# Credit: groupby segmentation, vintage default trajectory, optional logistic
# Investments: rolling vol, rolling correlation, drawdown, tear sheet
# Risk: Cholesky simulation, VaR/ES, sensitivity analysis
Cell 6: Save artefacts
def save_artefacts(figures, tables, name):
"""Save a dict of figures (PNG + PDF) and tables (CSV)."""
for label, fig in figures.items():
fig.savefig(f"{OUTPUT_DIR}/{name}_{label}.png", dpi=150, bbox_inches="tight")
fig.savefig(f"{OUTPUT_DIR}/{name}_{label}.pdf", bbox_inches="tight")
for label, df in tables.items():
df.to_csv(f"{OUTPUT_DIR}/{name}_{label}.csv")Save the template as project_template.ipynb in your working directory¹⁰ and clone it for every new project. The template is the difference between a project that takes a weekend and one that takes a fortnight. The discipline is in the structure,¹⁴ cell by cell.
Summary and Bridge
The project pipeline is now assembled. The five-element project brief (question, universe, method, deliverable, audience) sets the project up, and the three universes (credit, investments, risk) cover the analytical patterns most readers will encounter in their first three years of professional practice. Three worked examples demonstrated the pattern in each domain. Your memo structure translates the notebook’s output into a deliverable a non-technical reader can act on. The four-section self-evaluation rubric is the discipline that separates a deliverable a senior reader signs off on from one returned for revision.
The Six-Phase Project Pipeline framework (Brief, Data, EDA, Analysis, Memo, Self-Evaluation) is the sequence to carry into every project. Make it routine. In project work, consistency beats sophistication: run the six phases in order,²⁰ giving each the time and attention it deserves.
The reusable notebook template in the Python Lab is the spine of every project in the three universes. Clone it for each new project, fill in the project-specific cells, and the structure is in place from the first cell.
Chapter 10 turns to the next layer in the analyst’s toolkit: machine learning for finance. Where this chapter taught you how to assemble a complete project around an EDA-and-segmentation backbone, Chapter 10 teaches you how to add a small predictive model to that backbone: a regression for return prediction, a classification for default prediction. The scikit-learn API pattern, the supervised-learning setup, the train-test split, and the realistic expectations for ML performance on financial data are the topics. The project pipeline you have built in this chapter becomes the container into which the modeling step of Chapter 10 fits.
Key Terms
Project brief. A one-page document specifying the five elements of an analytical project: question, universe, method, deliverable, and audience. Written before any data is downloaded.
Universe. The data and time window the project covers. The universe statement is one sentence and identifies what is in the project and what is not.
Deliverable. The set of artefacts that the project produces and ships to the audience. Defined in the brief and refined as the project develops.
Audience. The eventual consumer of the deliverable: a partner, an investment committee, an interview panel, or a course evaluator. The audience profile drives the writing register and the level of technical detail.
Vintage analysis. The credit-domain technique of grouping loans by origination period and tracing the default rate within each cohort over time. Used to detect underwriting-quality drift across cohorts.
Tear sheet. A single-page or two-page summary of an asset’s or portfolio’s key risk-and-return statistics, charts, and metadata. Used in fund factsheets, fund-comparison databases, and analyst-prepared internal references.
Cholesky decomposition. The factorisation of a positive-definite matrix into the product of a lower-triangular matrix and its transpose. Used in finance to generate correlated random shocks from independent normal draws; see Chapter 4.
Historical simulation. A risk-measurement approach that uses the empirical distribution of historical returns or losses, without parametric assumptions, to compute VaR and ES. The standard cross-check against parametric Monte Carlo simulation.
Reproducibility. The property of an analysis that produces the same numbers and charts when re-run by another analyst on the same inputs. Achieved through committed code, sourced data, documented methodology, and explicit random seeds.
AUC. Area under the receiver-operating-characteristic curve. The standard performance metric for a binary classification problem; ranges from 0.5 (no skill) to 1.0 (perfect classifier).
Self-evaluation rubric. The checklist the analyst applies to the deliverable before submission, covering data, method, output, and reproducibility. Each row is a yes-or-no check; any "no" is a candidate for further work.
Markdown cell. A Jupyter Notebook cell containing formatted text rather than executable code. Used to introduce a section, document a methodological choice, or interpret an output.
Discussion Questions
1. The chapter argues that the one-page project brief (question, universe, method, deliverable, audience) should be written before any data is downloaded. Identify a project you started without a brief and walk through what specifically would have been different if a brief had been written upfront. What part of the eventual rework would the brief have prevented?
2. The three project universes (credit, investments, risk) cover most analyst work in the first three years of professional practice. Identify a project area you are interested in that does not fit cleanly into one of the three. What does the misfit reveal about the area, and how would you adapt the project pipeline to accommodate it?
3. The Practitioner’s Lens describes a Project Lead who allocates thirty percent of total project time to the memo and the chart preparation that supports it. What does this allocation cost in analytical depth, and under what conditions would a different project legitimately allocate more or less?
4. The Section 9.6 summary structure prescribes the question and headline answer first, the data and method second, the consequential finding third, the supporting findings fourth, and the implication-with-limitation fifth. Identify a memo you have written or read that deviates from this structure. What does the deviation gain, and what does it lose?
5. The Section 9.7 self-evaluation rubric covers data and provenance, method and reproducibility, output and audience. Add a fourth dimension to the rubric (one that you believe a working senior reviewer would care about) and justify the addition. What does adding the dimension exclude, and is the exclusion acceptable?
6. The chapter argues that the data-and-cleaning week is non-negotiable on any project longer than five working days. Identify a project where the data layer was rushed and walk through the consequences. What would the disciplined version of the data layer have looked like, and how much time would it have added or saved?
7. The reusable notebook template in the Python Lab is a six-cell scaffold. Identify a seventh cell you would add to the template based on your own project experience. What does the cell deliver that the existing six do not?
Self-Check Questions
Test your understanding of the worked project. Choose the best answer.
1. (Credit) Which dataset do you load to begin the credit portfolio worked example in Section 9.3?
a) S&P 500 daily returns
b) The lending-club style loan portfolio CSV referenced in Chapter 4
c) A synthetic options chain
d) FX tick data
Answer: b
2. (Credit) The default rate for a segment is computed as:
a) Sum of loan amounts divided by count
b) Count of defaulted loans divided by count of loans in the segment
c) Mean of interest rates
d) Median of credit scores
Answer: b
3. (Credit) When segmenting the credit portfolio by purpose and credit-score band, the resulting object is best described as:
a) A scalar
b) A pivot table or grouped DataFrame
c) A NumPy 1-D array
d) A Python set
Answer: b
4. (Credit) Which classification metric is most appropriate when default events are rare and false negatives are costly?
a) Plain accuracy
b) Recall (or AUROC) on the positive class
c) Mean squared error
d) R-squared
Answer: b
5. (Investments) The recommended workflow for the equity EDA in Section 9.4 begins with:
a) Fitting a deep neural network
b) Loading prices, computing returns, then summary statistics and visualization
c) Estimating VaR
d) Running a Monte Carlo simulation
Answer: b
6. (Investments) Daily simple returns for an equity series are computed as:
a) prices.diff()
b) prices.pct_change()
c) np.log(prices)
d) prices.cumsum()
Answer: b
7. (Investments) A correlation matrix close to the identity matrix across a basket of equities suggests:
a) Strong diversification benefit
b) Perfect co-movement
c) The data is corrupt
d) Returns are non-stationary by definition
Answer: a
8. (Risk) Historical Value-at-Risk at the 95% level for a P&L series is estimated by:
a) The sample mean
b) The 5th percentile of the loss distribution
c) The maximum drawdown
d) The standard deviation
Answer: b
9. (Risk) A loss distribution with negative skew indicates:
a) More frequent small losses than gains
b) A longer left tail (rare large losses)
c) Symmetric tails
d) No tail risk
Answer: b
10. (Risk) The threshold-exceedance count for a daily P&L series is:
a) The number of days the loss breached a stated threshold
b) The average loss
c) The volatility
d) The Sharpe ratio
Answer: a
Exercises
Apply the worked-example workflows to fresh data. Solutions are provided in Appendix E.
Exercise 9.1 (Credit). Reproduce the Section 9.3 credit portfolio analysis on the subprime segment of the loan dataset (credit-score band below 660). Compute default frequency by purpose, segment counts, and a baseline classifier; compare the subprime cut to the full-portfolio results from §9.3.
Exercise 9.2 (Investments). Repeat the Section 9.4 equity EDA on a basket of five NSE largecaps of your choice. Produce daily returns, summary statistics, the correlation heatmap, and a one-paragraph interpretation of diversification within the basket.
Exercise 9.3 (Risk). Apply the Section 9.5 risk workflow to a different daily P&L series (for instance, an FX position you construct from public spot data). Recompute historical VaR at 95% and 99%, the rolling 21-day volatility, and the threshold-exceedance count for a loss threshold of your choice.
Solutions are provided in Appendix E.
Further Reading
For project-pipeline methodology in data-and-analytics work, the CRISP-DM methodology at crisp-dm.eu is the most-referenced practitioner framework. The methodology predates modern Python tooling but its six-phase structure (business understanding, data understanding, data preparation, modeling, evaluation, deployment) maps directly onto the project pipeline this chapter teaches. Reading the CRISP-DM specification alongside the chapter is a useful exercise in seeing the same discipline named differently.
For analyst-memo writing specifically, Barbara Minto’s The Pyramid Principle (Pearson, third edition 2009) is the canonical practitioner reference. The book teaches a hierarchical structure that maps the McKinsey-style consulting memo onto a tree of supporting arguments. The five-paragraph structure of Section 9.6 is a simplified version of the Minto principle adapted to analyst deliverables. The full Minto framework is the right next step for any reader who finds the simpler version constraining.
For credit-portfolio project depth, Tony Van Gestel and Bart Baesens, Credit Risk Management: Basic Concepts (Oxford University Press, 2009), pairs naturally with this chapter’s Section 9.3. Van Gestel and Baesens defines the analytical artefacts (vintage curves, transition matrices, expected-credit-loss inputs) that the credit-portfolio project ultimately produces.
For investments-project depth, Yves Hilpisch, Python for Finance, second edition (O’Reilly, 2018), Chapters 11 through 13, covers the rolling-statistics, return-distribution, and tear-sheet patterns this chapter’s Section 9.4 introduces, at a level beyond the introductory.
For risk-project depth, the JPMorgan RiskMetrics Technical Document (1996) and Philippe Jorion, Value at Risk: The New Benchmark for Managing Financial Risk, third edition (McGraw-Hill, 2007), are the canonical references. Both pair naturally with this chapter’s Section 9.5 and with the Chapter 4 Python Lab on correlated Monte Carlo simulation. For deeper quantitative coverage, Alexander J. McNeil, Rüdiger Frey, and Paul Embrechts, Quantitative Risk Management (Princeton University Press, second edition 2015), is the comprehensive practitioner reference.
Endnotes
1. The standard convention for deriving a binary default flag from LendingClub’s loan_status field is documented in the platform’s historical data documentation and in the academic credit-risk literature that uses the dataset. See, for example, Galindo and Tamayo, “Credit Risk Assessment Using Statistical and Machine Learning: Basic Methodology and Risk Modeling Applications,” Computational Economics 15, no. 1–2 (2000), and the various Kaggle Discussion archives on the LendingClub competition.
2. AUC (area under the receiver-operating-characteristic curve) is the standard performance metric for binary classification problems including credit default. AUCs in the 0.65-0.70 range with a small feature set on LendingClub-type data are typical of the academic literature; AUCs above 0.75 generally require either richer feature sets, ensemble models, or specific data-engineering steps. The metric is documented at scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html.
3. MBA finance capstone-project conventions at the Indian Institutes of Management and the Indian School of Business are documented in the programs’ public curriculum descriptions at the institutional websites. The credit-portfolio analysis pattern is the most-assigned project in the credit-and-banking elective tracks.
4. Indian asset-management firms’ summer-internship programs are documented in the firms’ campus-recruitment communications and on the firms’ careers pages. The investment-committee presentation as the closing deliverable is the standard convention across HDFC AMC, ICICI Prudential, SBI Mutual Fund, Axis AMC, and Nippon India Mutual Fund.
5. Risk-advisory engagement structures at the Big Four consulting firms’ Indian practices and at boutique risk-consulting firms are documented in the firms’ service-line descriptions and case-study collections at kpmg.com, pwc.in, ey.com/en_in, and deloitte.com/in.
6. The Reserve Bank of India Occasional Papers are at rbi.org.in under Publications. The SEBI Working Paper series is at sebi.gov.in. The National Council of Applied Economic Research, the National Institute of Public Finance and Policy, and the Indian Council for Research on International Economic Relations publish their working papers at ncaer.org, nipfp.org.in, and icrier.org respectively.
7. Kaggle is at kaggle.com. The platform was founded in 2010 and acquired by Google in 2017. The Home Credit Default Risk competition, the LendingClub-related competitions, and the various credit-fraud-detection competitions are searchable at kaggle.com/competitions.
8. The CFA Institute Research Challenge lives at cfainstitute.org/en/programs/research-challenge. The competition runs annually with regional, sub-regional, and global rounds, and the deliverable structure has been stable since the early 2010s.
9. The Federal Reserve Bank of New York’s Liberty Street Economics blog is at libertystreeteconomics.newyorkfed.org. The blog publishes short analytical pieces by Federal Reserve research staff with methodology notes and supporting code where applicable. The Federal Reserve Bank of San Francisco maintains a comparable Economic Letter series at frbsf.org/economic-research/publications/economic-letter/.
10. The pandas to_parquet method, used in the Python Lab’s caching pattern, is documented at pandas.pydata.org/docs/reference/api/pandas.DataFrame.to_parquet.html. Parquet is the recommended format for cached analytical data because it preserves dtypes through round-trips and is fast to read.
11. The scikit-learn train_test_split utility you'll find at scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html. The random_state argument is the canonical reproducibility lever for the split.
12. Logistic regression in scikit-learn is documented at scikit-learn.org/stable/modules/linear_model.html#logistic-regression. The library is the standard introductory scikit-learn classifier and the right starting point for a binary classification project at the level Chapter 10 covers.
13. The pandas groupby user guide at pandas.pydata.org/docs/user_guide/groupby.html is the reference for the segmentation patterns the credit-portfolio worked example uses.
14. The matplotlib savefig API, including the dpi and bbox_inches parameters used in the Python Lab template, is described in matplotlib.org/stable/api/_as_gen/matplotlib.figure.Figure.savefig.html.
15. The Jupyter Notebook documentation at jupyter.org/documentation describes the markdown-and-code cell convention the Python Lab template depends on. The same structure works in JupyterLab and in Google Colab without modification.
16. The Cholesky decomposition pattern for correlated Monte Carlo simulation is introduced in Chapter 4 of this book and documented in Glasserman, Monte Carlo Methods in Financial Engineering (Springer, 2003), Chapter 2.
17. Historical-simulation VaR as a cross-check on parametric Monte Carlo VaR is the standard model-validation pattern in market-risk practice. See JPMorgan’s RiskMetrics Technical Document (1996) and Philippe Jorion, Value at Risk: The New Benchmark for Managing Financial Risk, third edition (McGraw-Hill, 2007).
18. The scikit-learn classification-metrics module (covering precision, recall, F1, the confusion matrix, and the precision-recall curve in addition to the ROC AUC) is documented at scikit-learn.org/stable/modules/model_evaluation.html. For credit-risk projects with class imbalance, the precision-recall curve is often more informative than the ROC AUC alone.
19. The Federal Reserve Economic Data repository at fred.stlouisfed.org and the pandas-datareader package at pydata.github.io/pandas-datareader together provide the macro-overlay path used in the Section 9.4 worked example. See Chapter 7 endnotes 11 through 13 for further documentation.
20. The Awesome Quant repository at github.com/wilsonfreitas/awesome-quant maintains a curated list of finance-and-quant Python resources, including dataset sources, modeling libraries, and example projects across the three domains this chapter covers.
21. Project-management literature on six-phase analytical workflows draws on, among other sources, the CRISP-DM methodology for data-mining projects and on the project-management practitioner literature published by the Project Management Institute. CRISP-DM is at crisp-dm.eu.
22. The Opening Vignette and the Practitioner’s Lens narratives are illustrative composites drawn from publicly described project-delivery experience at multiple Indian financial-analytics and consulting firms; no single individual or firm is depicted.