Universe data and appraisal unsmoothing¶
Author: Artur Sepp
The universe container and its unsmoothing transform are implemented in OptimalPortfolios. Software citation: CITATION.cff.
Overview¶
UniverseData holds one investment universe: a price panel, one metadata row per asset, up to
two levels of group loadings, and the labels that identify cash, equity, bond and private-equity
holdings. It is a frozen dataclass that checks on construction that these parts describe the same
assets. Its parts, not the container, are then passed to the covariance estimators, to group
constraints and to group risk budgets, as the worked example shows.
Private-equity and real-estate valuations are appraisals, which move slowly and carry over part
of the previous valuation. Their reported returns are smoothed: positively autocorrelated, with a
volatility well below that of the asset itself. copy_universe_data_with_unsmoothed_prices
returns a copy of the universe in which the flagged price columns are unsmoothed by
qis and everything else is unchanged.
Estimation then samples returns from the copy with compute_returns_from_prices. The
ROSAA case study also unsmooths its
private-asset returns before estimating the covariance.
This page derives what smoothing does to volatility and autocorrelation, states what the package’s copy changes, and shows on a simulated four-asset universe that unsmoothing lowers the autocorrelation of the private-equity series, restores its volatility and changes a risk-budgeting allocation.
Inputs, notation, and assumptions¶
Convention |
This article |
|---|---|
Return basis |
Log returns: the unsmoothing regression ( |
Estimation grid |
Quarter ends: the unsmoothing default |
Rebalancing grid |
None for the container; the example solves once, at the last quarter end of the sample |
Covariance units |
Annual: |
Expected returns |
None |
Weight state |
None for the container; the example’s risk-budgeting weights are target weights |
Solver |
None for the container and the unsmoothing, which are closed-form rolling regressions in qis; the example’s risk budgeting uses the package’s in-house solver |
The notation follows the conventions page. In addition:
Symbol |
Meaning |
|---|---|
\(x_t\) |
Reported quarterly log return of an appraisal-valued asset |
\(r_t\) |
Its true quarterly log return, unobserved in practice and simulated in the example |
\(\phi\) |
Smoothing weight of the previous reported return, \(0 \leq \phi \lt 1\) |
\(\rho_1\) |
Lag-1 autocorrelation |
\(L^{(1)}\), \(L^{(2)}\) |
Group loadings of levels 1 and 2: one row per asset, one column per group |
\(G\), \(n_g\) |
Number of groups; number of assets in group \(g\) |
\(\ell_g\), \(u_g\) |
Lower and upper limits on the allocation to group \(g\) |
\(y_t\), \(m_t\) |
Any sampled log return and its EWMA mean, with the decay \(\lambda\) of a span \(s\) |
Prices are positive levels on a DatetimeIndex, one column per asset. The smoothing results
assume that the true returns \(r_t\) are independent over time, with mean \(\mu\) and variance
\(\sigma^2\). The example’s panel is synthetic and drawn from a fixed seed; it is a teaching sample,
not an estimate for any private-equity index.
Methodology¶
The container and what construction checks¶
A UniverseData has ten fields:
Field |
Default |
Content |
|---|---|---|
|
required |
Price panel: a |
|
required |
One row per asset, indexed by the asset names |
|
|
Enum whose values name the metadata columns that must be present and complete |
|
|
Loadings \(L^{(1)}\): one row per asset, one column per group |
|
|
Loadings \(L^{(2)}\) of a second grouping |
|
|
Value of the asset-class column that marks cash-like assets |
|
|
Value of the asset-class column that marks equities |
|
|
Value of the asset-class column that marks bonds |
|
|
Name of the appraisal-valued private-equity asset |
|
|
Whether construction calls |
With validate_on_init=True, construction runs validate(), which raises ValueError at the
first failed check, in this order:
Alignment. The set of price columns equals the set of metadata rows; the message starts
Asset mismatchand names the assets found on one side only.Required fields. Every value of
metadata_fieldsis a metadata column.Duplicates. No asset name appears twice among the price columns or the metadata rows.
Nulls. No required metadata column holds a missing value.
Group loadings. When a level is given, the set of its row labels equals the set of price columns.
With validate_on_init=False the instance is built unchecked, and validate() can be called
later. The checks compare labels, not values or order: the metadata rows may come in another order
than the price columns, and missing or negative prices, fractional loadings, and identifiers that
name no asset or class all pass.
Metadata fields¶
MetadataField is a string enum with the members NAME = 'name',
ASSET_CLASS = 'asset_class' and CURRENCY = 'currency'; a member compares equal to its value.
With the default metadata_fields=MetadataField, the metadata needs the three columns name,
asset_class and currency, without missing values; further columns are allowed and unchecked.
Another enum sets another schema: its values become the required columns. The properties name,
asset_class and currency of the container read the default column names whatever the enum.
Group loadings, constraints and budgets¶
A loading matrix \(L\) has one row per asset and one column per group, and \(L^{\top} w\) is the
group exposure of the weights \(w\). With one-hot loadings, which qis.set_group_loadings builds
from a label column, each asset belongs to one group and a group’s exposure is the sum of its
members’ weights. The container holds two such matrices. Which grouping is level 1 and which is
level 2 is the caller’s convention; the example uses asset classes for level 1 and liquidity for
level 2. The container checks only their row labels.
The loadings reach the optimisers through the caller:
Group constraints.
GroupLowerUpperConstraintstakes a loading matrix asgroup_loadings, withgroup_min_allocationandgroup_max_allocation, and imposes \(\ell_g \leq \sum_i L_{ig} w_i \leq u_g\) on every group with a limit.GroupTrackingErrorConstraintandGroupTurnoverConstrainttake the samegroup_loadings. The constraints page describes each.Group risk budgets.
compute_group_risk_budgetstakes group labels, which one-hot loadings give back asidxmax(axis=1). With its defaultgroup_size_exponent=0, every nonempty group receives the same share of risk and splits it equally among its members, as HRP and cluster risk budgets derives:
Asset-class identifiers¶
liquidity_ac_id, equity_ac_id and bond_ac_id are values of the asset_class column, with
the defaults 'Liquidity', 'Equities' and 'Bonds'. They let code select the cash, equity or
bond assets of a universe by class, as get_hedge_ratio(hedged_acs) does for any list of
classes. pe_asset_id is different: it names one asset, not a class, and marks the series whose
prices are appraisals. The container stores the four labels and checks none of them.
What appraisal smoothing does¶
In the first-order smoothing model behind the reverse filter of Geltner (1993), each reported return carries a fixed share \(\phi\) of the previous reported return:
Proposition 1 (smoothing). If the true returns are independent with mean \(\mu\) and variance \(\sigma^2\), and \(0 \leq \phi \lt 1\), the stationary reported returns have
Proof. Unrolling the recursion gives \(x_t = (1 - \phi) \sum_{k \geq 0} \phi^k r_{t-k}\). The weights \((1 - \phi) \phi^k\) sum to one, so the mean is \(\mu\). The variance is \((1 - \phi)^2 \sigma^2 \sum_k \phi^{2k} = (1 - \phi)^2 \sigma^2 / (1 - \phi^2)\), which simplifies as stated. Since \(r_t\) is independent of \(x_{t-1}\), the recursion gives \(\mathrm{Cov}(x_t, x_{t-1}) = \phi \mathrm{Var}(x_{t-1})\), and the two variances are equal in the stationary state. \(\square\)
Proposition 2 (reverse filter). With \(\phi\) known, the true return is recovered exactly from two consecutive reported returns:
Proof. Solve the recursion for \(r_t\). \(\square\)
In practice \(\phi\) is unknown and can change over time. qis estimates the smoothing from the reported series itself, with rolling EWMA regressions of each return on its previous returns, which also allows smoothing over several lags as in Getmansky, Lo and Makarov (2004), and inverts it in the manner of Proposition 2. The estimator, its bounds and its warmup are derived on the qis page Private-asset unsmoothing and de-levering; the package does not change them.
What the unsmoothing copy changes¶
copy_universe_data_with_unsmoothed_prices(universe_data, assets_for_unsmoothing, ...) takes a
boolean Series assets_for_unsmoothing whose index equals universe_data.metadata.index, in the
same order, and applies qis.compute_ar_unsmoothed_prices to the columns flagged True. It
passes six of its arguments to qis:
Argument |
Default |
Passed to qis as |
|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Every other qis argument keeps its qis default, among them ar_order=2 and
min_value_for_beta=-0.25.
Flagged columns. Each is replaced by the NAV that qis returns: missing until the first identified coefficient, 1.0 on the last date before the first unsmoothed return, then compounded from the unsmoothed returns, and forward-filled onto the universe’s dates. The original price level is not kept.
Everything else. The other price columns are unchanged bit for bit, and the copy is built with the same
metadata,metadata_fields, both group loadings,equity_ac_id,bond_ac_id,pe_asset_idandvalidate_on_init.liquidity_ac_idis not passed and returns to its default.Nothing flagged. The input universe itself is returned.
Errors. A
ValueErroris raised when the flag’s index, or the index of afreqSeries, differs from the metadata index.
The function does not read pe_asset_id; the caller builds the flag, typically from it.
Returns for estimation¶
compute_returns_from_prices takes prices and has the defaults returns_freq='ME',
demean=True, drop_first=True, is_first_zero=False, is_log_returns=True and span=52. It
samples log returns at returns_freq through qis.to_returns, without the first, empty row. With
demean=True it subtracts the EWMA mean of span span, which includes the current return.
estimate_current_ewma_covar and EwmaCovarEstimator build their returns with it, so this is
where an unsmoothed copy enters estimation.
Proposition 3 (demeaned returns). With \(m_1 = y_1\) and \(m_t = \lambda m_{t-1} + (1 - \lambda) y_t\), the demeaned return is
and zero at \(t = 1\), which is why, with drop_first=True, the function drops that row as well.
Proof. Substituting \(m_t\) gives \(y_t - \lambda m_{t-1} - (1 - \lambda) y_t\), which is \(\lambda (y_t - m_{t-1})\). \(\square\)
Each demeaned return is therefore the surprise against the previous mean, scaled down by \(\lambda\): by \(39/41 \approx 0.951\) at a span of 40 quarters.
Worked example¶
The canonical script of this page,
examples/docs/universe_data_and_unsmoothing.py,
runs offline and asserts every number quoted here, including the three propositions:
python -m examples.docs.universe_data_and_unsmoothing
The universe has four assets with quarterly prices from 31 December 1989 to 31 December 2024: government bonds, credit, listed equity and private equity. The true log returns are drawn from a fixed seed, with annual volatilities of 5%, 8%, 16% and 20%, and the private-equity price is built from returns smoothed with \(\phi = 0.6\). The level-1 loadings group the assets by asset class, the level-2 loadings by liquidity:
prices, true_returns = simulated_panel(seed=SEED)
metadata = pd.DataFrame({'name': NAMES, 'asset_class': ASSET_CLASSES, 'currency': 'USD'},
index=TICKERS)
level1 = qis.set_group_loadings(group_data=metadata['asset_class'])
level2 = qis.set_group_loadings(group_data=pd.Series(LIQUIDITY, index=TICKERS))
universe = op.UniverseData(prices=prices, metadata=metadata, group_loadings_level1=level1,
group_loadings_level2=level2, pe_asset_id='PE')
Each of the five inputs below fails one check on construction:
invalid = {
'metadata row missing': dict(prices=prices, metadata=metadata.drop(index='PE')),
'required column missing': dict(prices=prices, metadata=metadata.drop(columns='name')),
'duplicate price column': dict(prices=prices[TICKERS + ['PE']], metadata=metadata),
'null in required column': dict(
prices=prices, metadata=metadata.assign(currency=['USD', 'USD', 'USD', None])),
'loadings on other assets': dict(prices=prices, metadata=metadata,
group_loadings_level2=level2.drop(index='PE')),
}
errors = {case: construction_error(**arguments) for case, arguments in invalid.items()}
Input |
The |
|---|---|
Metadata without the |
|
Metadata without the |
|
|
|
A missing currency |
|
Level-2 loadings without the |
|
The flag for unsmoothing marks the asset that pe_asset_id names:
flag = pd.Series(universe.metadata.index == universe.pe_asset_id,
index=universe.metadata.index)
unsmoothed = op.copy_universe_data_with_unsmoothed_prices(universe_data=universe,
assets_for_unsmoothing=flag)
The government-bond, credit and equity prices of the copy equal the originals bit for bit, and the copy holds the same metadata and loading objects. The private-equity column is missing for the first 17 quarters, equals 1.0 on 31 March 1994 and compounds the unsmoothed returns from 30 June 1994. Over the sample, the coefficient sum that qis estimates stays between 0.45 and 0.67, inside its bounds.
Over the 123 quarters from June 1994, the reported series has a lag-1 autocorrelation of 0.61 and a volatility of 10.2%. The unsmoothed series has an autocorrelation of −0.04 and a volatility of 20.5%, close to the −0.02 and 20.4% of the simulated true returns:
before = op.compute_returns_from_prices(universe.prices, returns_freq='QE', demean=False)
after = op.compute_returns_from_prices(unsmoothed.prices, returns_freq='QE', demean=False)
window = after['PE'].dropna().index
stats = pd.DataFrame({
'lag-1 autocorrelation': [before.loc[window, 'PE'].autocorr(lag=1),
after.loc[window, 'PE'].autocorr(lag=1)],
'annualised volatility': [before.loc[window, 'PE'].std() * 2.0,
after.loc[window, 'PE'].std() * 2.0]},
index=['smoothed', 'unsmoothed'])

Figure: the reported, unsmoothed and simulated true returns of the private-equity column of the
example, as cumulative paths and as lag-1 autocorrelation and annualised volatility. Drawn by the
exhibit function of the canonical script; the analytics gallery lists
its provenance.
Insight
With \(\phi = 0.6\), the volatility factor \(\sqrt{(1 - \phi)/(1 + \phi)}\) of Proposition 1 is exactly one half, and the sample agrees: 10.2% reported against 20.4% true. A covariance estimated on the reported returns understates the variance of the private asset about fourfold, although smoothing leaves its expected return unchanged.
The loadings then feed an allocation. The level-1 loadings give two asset classes of two assets each, so equal group budgets give every asset a quarter of the risk. The level-2 loadings cap the illiquid group at 15% of capital. The same risk-budgeting problem is solved on the EWMA covariance of each universe, estimated over the common sample from March 1994:
budgets = op.compute_group_risk_budgets(groups=universe.group_loadings_level1.idxmax(axis=1))
illiquid_cap = op.GroupLowerUpperConstraints(
group_loadings=universe.group_loadings_level2, group_min_allocation=None,
group_max_allocation=pd.Series({'Liquid': 1.0, 'Illiquid': ILLIQUID_CAP}))
constraints = op.Constraints(is_long_only=True, group_lower_upper_constraints=illiquid_cap)
start = unsmoothed.prices['PE'].first_valid_index()
covars, weights = {}, {}
for label, data in {'smoothed': universe, 'unsmoothed': unsmoothed}.items():
covars[label] = op.estimate_current_ewma_covar(prices=data.prices.loc[start:],
returns_freq='QE', span=SPAN)
weights[label] = op.wrapper_risk_budgeting(pd_covar=covars[label],
constraints=constraints, risk_budget=budgets)
On the reported prices, the estimated volatility of private equity is less than half its unsmoothed value. Without the cap, risk budgeting would put 22% of capital in it for its quarter of the risk; with the cap, it holds exactly 15%. On the unsmoothed prices, private equity takes 10% of capital, the cap is slack, and every asset carries a quarter of the risk.
Implementation in optimalportfolios¶
The four objects are exported at the package root:
UniverseData, inuniverse/universe_data.py: the frozen container andvalidate().from_selection(prices, metadata, assets, ...)subsets the prices, metadata and loadings to the listed assets, in that order.save(file_name, local_path)writes the parts to CSV files through qis, andload(file_name, local_path, ...)reads them back; withoutmetadata_fields,loadmakes every metadata column it reads a required field.rename_index()relabels the assets by theirnamecolumn. The propertiesname,asset_class,currency,assets,n_assetsanddate_rangeread the parts, andget_hedge_ratio(hedged_acs)returns 1.0 for the assets of the listed classes and 0.0 for the others.get_asset_returns_dictis described with mixed-frequency data.MetadataField, in the same module: the default schema of the metadata.copy_universe_data_with_unsmoothed_prices, inuniverse/universe_transforms.py: the unsmoothing copy described above.compute_returns_from_prices, incovar_estimation/utils.py: the return construction of the EWMA covariance estimators.
Pitfall
The docstring of copy_universe_data_with_unsmoothed_prices promises AR(1)
unsmoothing, but the function does not pass ar_order, and qis.compute_ar_unsmoothed_prices
defaults to ar_order=2: the copy is unsmoothed with a rolling AR(2) filter. The example checks
that its column equals the qis AR(2) result and differs from the AR(1) result. For AR(1), call
qis.compute_ar_unsmoothed_prices with ar_order=1 and build the universe from its NAVs.
Interpretation and limitations¶
What Geltner’s model does not bring. The package takes from Geltner (1993) the first-order smoothing model that the example simulates and the reverse filter of Proposition 2. It does not inherit the paper’s derivation of the smoothing from appraisal behaviour, which avoids assuming that market returns are uncorrelated over time, nor its corrections for temporal aggregation and seasonal reappraisal. qis estimates the smoothing from the autocorrelation of the reported series, so all serial correlation is treated as smoothing, including any in the true returns.
Estimation noise. The coefficients come from a rolling EWMA regression with a span of 40 quarters. The example’s unsmoothed volatility is close to the true one, but other samples differ: with seed 3 instead of 19, the unsmoothed volatility is 30.7% against a true 19.5%.
Shorter history. The unsmoothed column is missing for the first 17 quarters of the example, and the example’s covariance starts after them. Its level starts at 1.0, so only its returns are comparable with the original.
Finer grids. On a monthly panel with the default
freq='QE', the unsmoothed column changes only at quarter ends. A monthly estimator then sees two zero returns in every quarter; keep such an asset in a quarterly bucket, as described with mixed-frequency data.Labels, not values. Construction accepts missing or negative prices, fractional loadings and identifiers that name nothing; incomplete histories describes how the solvers treat missing data.
Identifiers lost in transforms.
from_selectionandrename_indexreturn a universe with the default identifiers andpe_asset_id=None, and the unsmoothing copy resetsliquidity_ac_id. Pass the identifiers again where they are needed.
See also¶
References¶
Geltner, D. (1993). Estimating Market Values from Appraised Values without Assuming an Efficient Market. Journal of Real Estate Research, 8(3), 325–345. DOI 10.1080/10835547.1993.12090713.
Getmansky, M., Lo, A. W. and Makarov, I. (2004). An Econometric Model of Serial Correlation and Illiquidity in Hedge Fund Returns. Journal of Financial Economics, 74(3), 529–609. DOI 10.1016/j.jfineco.2004.04.001. Smoothing over several lags, the model that qis cites for its filters.
Sepp, A., Ossa, I. and Kastenholz, M. (2026). Robust Optimization of Strategic and Tactical Asset Allocation for Multi-Asset Portfolios. The Journal of Portfolio Management, 52(4), 86–120. DOI 10.3905/jpm.2025.1.806. Unsmooths its private-asset returns before estimation.