Mingyang Liu

On the 2026–27 academic job market.

Portrait of Mingyang Liu

I am a Visiting Researcher at the Imperial Centre of Excellence in Quantitative Finance, Imperial College London, hosted by Robert Kosowski, and a part-time AI and machine-learning adviser to the Cost of Capital team at KPMG LLP (UK). My primary research fields are empirical asset pricing, machine learning and big data in finance, and financial econometrics. Working with large panels of stock characteristics and regularised machine-learning estimators, I ask which parts of a factor model’s answer (its factors, its alphas, its portfolios) reflect information in the data, and which reflect only how the model measures that information. I follow these measurement choices from identification and inference to the portfolios an investor ends up holding.

I received my PhD in Finance from Imperial College London in 2026, supervised by Paolo Zaffaroni and Pasquale Della Corte, after an MRes in Finance with Distinction (2020). Before Imperial, I earned a B.S. in Economics and Mathematics (Mathematics of Finance and Risk Management track) from the University of Michigan and an M.A. in Economics (Financial Econometrics) from Columbia University.

Static version of Figure 1: the 8-by-8 signed-overlap table for the window ending March 2009. Beside it, the timeline of average overlap size and participation dimension from 1983 to 2025.
Characteristic12345678

Scale of the signed overlap: +1 means the same stocks score high on both, −1 the opposite. A value of 0 means the scores do not line up. Each characteristic is signed so that the named trait scores high. Diagonal cells are 1 by construction. The average in the readout ignores the sign.

Timeline, 1983 to 2025. Left axis, fixed from 0.20 to 0.35: the average size of the eight characteristics' overlaps, sign ignored. Right axis, fixed from 20 to 24: the participation dimension of the whole library. The cursor marks the window shown.

Use the slider or the arrow keys to move the window; the table updates.

Across the windows the average size of the overlaps ranges from 0.24 to 0.29. The overlap between small size and illiquidity ranges from 0.61 to 0.93. The participation dimension of the whole library stays between 21.7 and 22.9.

Figure 1: How eight related characteristics overlap, 1983–2025. Each cell shows the signed overlap of two characteristics' scores across stocks (small, illiquid, thinly traded, volatile, distressed, levered, unprofitable or young firms), measured with the characteristic ruler (the Gram metric) on the 20 years ending in the month shown. These eight were picked in advance as a group of related traits, and most of their pairs overlap more than is typical: across the whole library most pairs barely overlap (median 0.06). Under the plain ruler (the Euclidean or identity metric) every off-diagonal cell is 0 by assumption; both rulers are conventions, not discoveries. A descriptive exhibit from the author's research materials, not from a paper on this site: the timeline's participation dimension (how many equally large directions of variation all 153 characteristics behave like) stays close to 22 and counts directions in the data, not priced risks; the footnote gives the construction.1

Figure 1 shows what the characteristic ruler measures and the plain ruler assumes away: how much eight related characteristics overlap across stocks, re-estimated each month from 1983 to 2025.1 The next section explains what a ruler is, and why switching rulers changes how returns are split between factors and the part no factor explains, and what an investor holds.

Research agenda

Same information, different answer: before trusting what a factor model says, check its ruler.

A factor model explains many stocks’ returns with a few common drivers. Built on stock characteristics (size, value, momentum and many more), it gives three answers: which risks are priced, how much expected return its factors leave unexplained (the alpha), and what to hold. I ask how much of each answer comes from information in the data, and how much from choices that add no information. The central choice is the ruler. To credit a stock’s expected return to its drivers, a model must first decide when two combinations of characteristics are really different and how large each one is. The rule it uses for that is the ruler. I estimate the ruler from the data. I show that changing it changes how returns are split between the factors and the part no factor explains (the intercept; the alpha is a close relative, measured against a stated benchmark model), but not how well they are fitted, and I carry the ruler’s own estimation error into the confidence intervals. I then ask what a smaller alpha really reflects, and how the ruler, and even the way a library of characteristics is written down, changes what an investor holds. Researchers and portfolio managers change such conventions routinely, and a change of convention is easy to mistake for a discovery.

Give each characteristic a weight and add them up: every stock gets a score. Such a weighted combination is called a direction (it is not a portfolio). The plain ruler (the Euclidean or identity metric) judges a direction by its weights alone. The characteristic ruler (the Gram metric) judges it by the scores it gives real stocks: how large they typically are, and whether two directions’ scores line up across stocks, high on the same stocks and low on the same stocks.

The ruler is the convention a model uses to measure how large a direction is and how much two directions overlap. Two value signals that pick out the same cheap stocks count as overlapping under the characteristic ruler but as unrelated under the plain ruler.

what goes in Library of characteristics how it is measured Ruler what the model says Forecast factors + intercept what the investor holds Portfolio 1 Same data, different factors 2 Same fit, different split 3 Same ruler, different target 4 Same forecast, different portfolio 5 Same information, different holdings Job market paper
Figure 2: Where each paper enters the pipeline. Numbers give the reading order; the job market paper follows a library change all the way to holdings and returns after trading costs.

Read in this order

Each paper holds everything else fixed, changes one choice, and asks what moves.

  1. Which ruler? → Geometric Framework (earlier, broader version)
  2. What does the ruler decide, and how reliable is the answer? → Characteristic-Space Metrics
  3. What does a smaller alpha mean? → Interpreting Pricing Errors
  4. Same forecast, different portfolio? → Characteristic Geometry
  5. Same information, different holdings? → Characteristic Libraries Job market paper

Short on time? Start with the job market paper.

Research

Job market paper Working paper · September 2026

Characteristic Libraries and Portfolio Decisions

Same information, different holdings: copying or reweighting characteristics a library already has adds no information, yet it changes what a tuned investment rule holds, because the model’s penalty on large bets charges a copied theme less. Balancing the themes alone moves four such rules’ target holdings by 19–28% of their average total position size (long plus short). Whether returns after trading costs rise or fall depends on the method, the cost level and how missing returns are valued.

Paper page · PDF

SSRN preprint · March 2026 (preliminary version)

A Geometric Framework for Identification in Characteristic-Based Factor Models

Same data, different factors: which combinations of characteristics count as distinct depends on the ruler. The paper takes the ruler from how characteristics spread and overlap across stocks, and builds identification, no-arbitrage restrictions and inference on it.

Earlier, broader version of Characteristic-Space Metrics.

Paper page · PDF · SSRN

SSRN preprint · September 2026

Characteristic-Space Metrics in Factor Models: Identification and Inference

Same fitted returns, different split: the ruler decides how returns are credited to hidden factors, observed factors such as the market, and an intercept. When the ruler is estimated from the data, its error belongs in the confidence intervals.

Paper page · PDF · SSRN

Working paper · September 2026

Interpreting Estimated Pricing Errors: Evidence from Characteristic-Based Return Forecasts

Same ruler, different target: an alpha can shrink because the factors explain more, or because the expected-return estimate they must explain (the target) has changed. Pairing each model’s estimate with each model’s factors on the same stocks, the paper finds that the large alpha gap between two training rules mainly reflects different targets.

Paper page · PDF

Working paper · September 2026

Characteristic Geometry and Portfolio Choice

Same forecast, different portfolio: a small fixed stability constant in the portfolio rule is written in the ruler’s units, so changing the ruler changes how strongly each stock position is held back. Remove only that constant and the weight gap nearly disappears.

Paper page · Slides (PDF)

Path

The Gothic stone facade of the University of Michigan Law Library, its tall leaded windows lit at dusk.

Ann Arbor · 2014–2017

University of Michigan

B.S. in Economics and Mathematics

Photo: University of Michigan

Low Memorial Library at Columbia University, with the Alma Mater statue on the steps in front.

New York · 2017–2019

Columbia University

M.A. in Economics (Financial Econometrics)

Photo: Columbia Alumni Association, Columbia University

The Sherfield Building and Library at Imperial College London, seen across Queen’s Lawn.

London · 2019–2026

Imperial College London

MRes in Finance (2020) and PhD in Finance (2026); Visiting Researcher from 2026

Photo: Imperial College London

Contact


  1. Characteristics: the Jensen–Kelly–Pedersen library of 153 U.S. stock characteristics (JKP153), the same library the job market paper uses, each characteristic rank-centred within the month and signed so that the trait named in the row scores high. Each window averages, with equal weights, the monthly 153×153 second-moment matrices of these characteristics over the 240 months ending in the month shown. Window ends run from June 1983 to November 2025. There are 510 windows, and no future months enter any window. The 240-month equal-weighted window is this exhibit’s own choice. Characteristic Geometry reports the rolling-60, equal-120 and EWMA (exponentially weighted moving average) versions of the characteristic ruler, and the path depends on the window. The table takes the 8×8 block for market equity (sign reversed), Amihud illiquidity, turnover (reversed), idiosyncratic volatility, a distress score (Ohlson O-score), debt to market equity, operating profitability (reversed) and firm age (reversed), and rescales it to unit diagonals, so the cells are approximately correlations. The block is a chosen high-overlap corner of the library, not a representative sample; summaries of the overlaps among all 153 characteristics are in Characteristic-Space Metrics, Table I (p. 31) and Figure 1 (p. 32). The most overlapping pair is small size and illiquidity in every window; Amihud illiquidity is built from dollar volume, so a large overlap with size is expected. Accounting inputs change only when new financial statements are filed, and coverage, the share of missing values and the stock universe change over time; those changes move these averages as well as any change in how firms line up. The participation dimension in the timeline is computed from the full 153×153 window matrix, not from the 8×8 block, and runs from 21.7 to 22.9. This is a descriptive visualisation, not a new estimator or causal test.↩︎