Research note
Quantamental Equity Research Model
A transparent systematic stock selection research framework combining fundamental quality, valuation discipline, market recognition, and reinvestment behaviour for long-term U.S. equity investing.
Research Question
How can fundamental quality, valuation discipline, market recognition, and reinvestment behaviour be combined into a transparent systematic stock selection model for long-term U.S. equity investing?
Project Overview
This project develops a point-in-time quantamental equity research model for long-term U.S. stock selection. The model is designed to combine fundamental research logic with systematic cross-sectional factor analysis, using approximately 30 years of historical data across around 2,600 U.S.-listed equities.
The central objective is not to build a black-box prediction model, but to create an interpretable research framework that evaluates companies through transparent metrics, industry-specific valuation logic, and portfolio-level risk controls.
The framework follows a full research pipeline:
- Historical equity data cleaning
- Point-in-time-aware data controls
- Tradable universe construction
- GICS industry group classification
- Industry-specific valuation rule design
- Seven signal family construction
- Cross-sectional factor ranking
- Risk, liquidity, and valuation overlay
- Portfolio construction
- Backtest evaluation and robustness analysis
Data and Universe Construction
The model uses approximately 30 years of historical U.S. equity data across around 2,600 listed stocks. The dataset includes price history, liquidity measures, corporate-action-related data, and company-level fundamentals.
The data pipeline focuses heavily on reducing common backtest distortions. Point-in-time-aware controls are applied to reduce look-ahead bias, while monthly completeness checks are used to avoid incomplete-period effects. End-of-month price validation and rolling liquidity filters are used to reduce stale-price and illiquidity-related distortions.
The tradable universe is constructed through a combination of:
- Common stock filtering
- Historical price availability checks
- End-of-month price validation
- Complete-month controls
- Rolling liquidity filters
- Corporate-action awareness
- Removal or flagging of stale, illiquid, or unreliable observations
The goal is to ensure that the model is tested on a realistic universe rather than a clean but biased dataset.
Industry-aware Valuation Logic
A core part of the project is the classification of stocks by GICS industry groups. Instead of applying one universal valuation metric across all companies, the model evaluates companies using industry-specific valuation logic.
This matters because valuation metrics are not equally meaningful across all business models. For example, earnings-based valuation may be more useful for mature profitable companies, while revenue or gross-profit-based measures may be more relevant for earlier-stage growth companies. Asset-heavy businesses may require different valuation anchors from software, financials, industrials, or cyclical companies.
The valuation framework therefore focuses on metric applicability and valuation rules:
| Valuation Logic Family | Main GICS Industry Groups | Primary Valuation Anchor | Main Risk | Initial Backtest Priority |
|---|---|---|---|---|
| Financial Balance Sheet | Banks, Financial Services, Insurance | P/TBV, P/B, business-line normalized P/E | Regulatory capital, credit losses, accounting heterogeneity | High |
| Asset Yield / Regulated | Utilities, Equity REITs, Real Estate Management & Development, selected Mortgage REITs | P/FFO, AFFO yield, cap-rate or rate-base spread | Non-GAAP comparability, rate sensitivity, external financing reliance | Medium |
| Mature Compounder | Consumer Staples, Health Care Equipment & Services, mature Software & Services, branded Industrials and Discretionary | FCF yield, EV/EBIT, reverse DCF gap | Overpaying for quality, buybacks masking slower growth | Highest |
| Cyclical / Capital Intensive | Energy, Materials, Capital Goods, Transportation, Autos, Semiconductors, Tech Hardware | EV/normalized EBIT, mid-cycle EBITDA, trough EV/Sales | Low P/E illusion near cycle peaks, inventory and commodity-price volatility | Highest |
| High Growth / Intangible | Software & Services, Media & Entertainment, platform Consumer Services, selected Health Care Tech | EV/Gross Profit, quality-matched EV/Sales, reverse DCF implied margin | Duration risk, SBC dilution, overheated narratives | High |
| Turnaround / Distressed | Cross-GICS state-based overlay | EV/Sales, post-turnaround EV/EBIT, balance-sheet floor valuation | Refinancing risk, delisting risk, false recovery | Low to Medium |
| Pre-profit Optionality | Pharma, Biotech & Life Sciences, selected early medtech and hard-tech | EV/Net Cash, EV/NTM Revenue | Binary events, dilution, limited comparable samples | Lowest |
Signal Family Design
The model uses seven interpretable signal families. Each family is designed to capture a different economic dimension of stock selection.
| Signal Family | Purpose |
|---|---|
| Valuation / Mispricing Anchor | Identifies whether a stock is attractively valued relative to its industry, fundamentals, and growth profile. |
| Quality / Profitability / Durability | Measures business strength through profitability, margin quality, cash generation, and durability of returns. |
| Price Momentum | Captures market price confirmation and trend persistence. |
| Recognition / Revision | Measures whether the market is beginning to recognize improving fundamentals through estimate revisions, earnings surprises, or improving sentiment. |
| Growth / Runway / Operating Leverage | Evaluates revenue growth, margin expansion potential, scalability, and long-term runway. |
| Capital Allocation / Financing / Shareholder Return | Assesses management behaviour through buybacks, dividends, leverage, dilution, reinvestment, and financing quality. |
| Investment / Reinvestment Quality | Evaluates whether reinvestment is creating value through ROIC, incremental returns, capex efficiency, and long-term compounding quality. |
The model is designed to avoid relying on a single factor. A stock is more attractive when multiple signals point in the same direction, such as reasonable valuation, strong quality, improving recognition, sustainable growth, and disciplined capital allocation.
Cross-sectional Factor Ranking
The model ranks companies cross-sectionally within relevant peer groups. This avoids comparing structurally different companies in inappropriate ways.
For example, a software company, a bank, and an industrial manufacturer should not be judged using the same valuation metric with the same threshold. Instead, each company is evaluated relative to the metrics that are most applicable to its industry group and business model.
The cross-sectional ranking process is designed to answer:
- Which companies are attractively valued relative to comparable peers?
- Which companies have stronger quality and profitability?
- Which companies show improving recognition or revision trends?
- Which companies combine growth potential with reasonable valuation?
- Which companies reinvest capital efficiently?
- Which companies should be avoided due to poor liquidity, excessive valuation, weak quality, or unstable fundamentals?
Portfolio Construction
After signal construction, the model applies portfolio-level overlays before forming the final long-only portfolio.
| Overlay | Purpose |
|---|---|
| Risk Overlay | Reduces unintended concentration, excessive beta exposure, and unstable return behaviour. |
| Liquidity Overlay | Avoids illiquid stocks and reduces exposure to names that may not be realistically tradable. |
| Valuation Overlay | Avoids overextended stocks where price momentum or growth signals are not supported by valuation discipline. |
| Industry Exposure Review | Checks whether performance is driven by stock selection or unintended industry bets. |
| Turnover Control | Reduces excessive rebalancing and improves transaction-cost realism. |
The final portfolio is designed to capture stock-selection alpha while remaining explainable, diversified, and implementable.
Evaluation Framework
The model is evaluated across five areas: portfolio performance, benchmark-relative performance, factor validation, implementation, and robustness.
1. Portfolio Performance
| Metric | Purpose |
|---|---|
| CAGR | Measures long-term annualized return. |
| Cumulative Return | Measures total return over the full backtest period. |
| Annualized Volatility | Measures return stability. |
| Max Drawdown | Measures the largest peak-to-trough portfolio loss. |
| Drawdown Duration | Measures how long the portfolio takes to recover from losses. |
| Sharpe Ratio | Measures risk-adjusted return relative to total volatility. |
| Sortino Ratio | Measures return quality relative to downside volatility. |
| Calmar Ratio | Measures annualized return relative to maximum drawdown. |
2. Benchmark-relative Performance
| Metric | Purpose |
|---|---|
| Alpha | Measures excess return not explained by market beta. |
| Beta | Measures sensitivity to broad market movement. |
| Active Return | Measures annualized outperformance versus benchmark. |
| Tracking Error | Measures volatility of active return. |
| Information Ratio | Measures active return per unit of active risk. |
| Up Capture | Measures how much upside the model captures in rising markets. |
| Down Capture | Measures how much downside the model experiences in falling markets. |
3. Factor Validation
| Metric | Purpose |
|---|---|
| Rank IC | Measures whether factor rankings are positively related to future stock returns. |
| ICIR | Measures the stability of factor predictive power. |
| Positive IC Months | Measures how often the factor direction is positive. |
| Top-minus-Bottom Quantile Spread | Measures return difference between highest-ranked and lowest-ranked stocks. |
| Quantile Monotonicity | Tests whether higher-ranked factor buckets deliver progressively better returns. |
| Factor Decay | Measures whether factor effectiveness persists across 1-month, 3-month, 6-month, and 12-month horizons. |
4. Implementation Controls
| Metric | Purpose |
|---|---|
| Turnover | Measures how frequently the portfolio changes holdings. |
| Transaction Cost-adjusted Return | Tests whether performance remains attractive after trading costs. |
| Liquidity Exposure | Checks whether returns rely on illiquid or hard-to-trade stocks. |
| Average Holdings | Measures diversification. |
| Max Single-name Weight | Controls individual stock concentration. |
| Sector / Industry Exposure | Checks whether returns are driven by stock selection or industry allocation. |
5. Robustness Tests
| Test | Purpose |
|---|---|
| Subperiod Performance | Tests whether the model works across different market regimes. |
| Out-of-sample Performance | Tests whether the model remains effective outside the research period. |
| Ablation Study | Measures the contribution of industry valuation rules and portfolio overlays. |
| Equal-weight Benchmark Comparison | Tests whether performance is driven by broad market structure. |
| Signal-family Attribution | Identifies which signal families contribute most to return and risk control. |
Illustrative 10-year Backtest Output
The following table represents an illustrative output structure for a 10-year backtest. Actual values should be updated after final backtest validation.
| Area | Metrics | Illustrative Output |
|---|---|---|
| Performance | CAGR, Volatility, Max Drawdown, Sharpe, Sortino | 14.2% net CAGR, 16.9% volatility, -27.4% max drawdown, 0.74 Sharpe |
| Benchmark-relative | Alpha, Beta, Information Ratio, Tracking Error | 2.4% annualized alpha, 0.94 beta, 0.43 information ratio, 6.5% tracking error |
| Factor Validation | Rank IC, ICIR, Quantile Spread, Factor Decay | 0.028 Rank IC, 0.35 ICIR, 0.58% monthly top-bottom spread |
| Implementation | Turnover, Liquidity Exposure, Transaction Cost-adjusted Return | 18% monthly turnover, >$50m median ADV, 15.0% gross CAGR, 14.2% net CAGR |
| Robustness | Subperiod Test, Out-of-sample Test, Ablation Study | Positive alpha in 3/4 market regimes, 1.5-2.0% early OOS alpha, overlay reduced peak drawdown by approximately 5 percentage points |
Expected Result Interpretation
The ideal result of this framework is not an unrealistic high-Sharpe strategy, but a credible long-only quantamental equity model that generates moderate benchmark-relative outperformance with stronger risk-adjusted returns and better downside control.
A successful result would show:
- CAGR higher than the benchmark
- Sharpe and Sortino ratios higher than the benchmark
- Max drawdown lower than the benchmark or recovery time shorter
- Positive alpha with beta close to or below 1
- Positive information ratio
- Positive Rank IC across signal families
- Clear top-minus-bottom quantile spread
- Reasonable turnover and transaction-cost-adjusted return
- No excessive dependence on illiquid stocks, one sector, or one market regime
The model is therefore best understood as a systematic active equity research framework rather than a black-box trading strategy. Its value comes from combining data discipline, industry-aware valuation, interpretable factor design, and portfolio-level risk control.
Limitations and Further Work
Results are based on a point-in-time monthly backtest with estimated transaction costs. Out-of-sample evidence is preliminary due to limited sample length, and the current research design still requires careful validation.
Survivorship Bias / Delisting Bias
Although the backtest uses historical data rather than only current index constituents, survivorship and delisting bias may remain due to incomplete coverage of delisted securities and identifier changes.
Point-in-Time Data / Look-ahead Bias
Financial statement and valuation signals rely on point-in-time alignment. Reported results may still be affected by residual look-ahead bias if filing dates, restatements, or vendor update timestamps are imperfectly captured.
Data Vendor Coverage Bias
The project uses publicly accessible market and fundamental data rather than institutional CRSP/Compustat-style datasets, which may limit the precision of historical coverage, delisting returns, corporate actions, and identifier continuity.
Backtest Period Limitation
The main backtest focuses on approximately the past ten years, which may not fully capture long-cycle market environments such as prolonged inflation, deep credit stress, or multi-year value/growth regime reversals.
Factor Data Mining / Multiple Testing Risk
Since multiple signals and portfolio rules are evaluated, the results may be exposed to data-mining and multiple-testing risk. Factor validation therefore emphasizes Rank IC, ICIR, top-bottom spreads, regime tests, and out-of-sample checks rather than relying only on headline backtest performance.
Further work includes improving point-in-time fundamental data validation, testing delisting and survivorship effects, adding more realistic transaction cost assumptions, conducting signal-family attribution, and expanding out-of-sample validation across different market regimes.