Methodology
Abstract
We constructed a new quality-adjusted commercial real estate rent index for U.S. office, retail, and industrial markets using more than one million CompStak lease transactions from 2010-2025. A hierarchical hedonic framework with building-, block-, and ZIP-level fixed effects allows us to control for both observable and unobserved quality, producing quality-adjusted rent indices. Nationally, these indices show that raw and simple hedonic rent series often overstate post-COVID recoveries, especially in office and retail, and at times understate pandemic-related declines, reflecting shifts in the quality mix of transacting properties rather than true market movements. Industrial rents, by contrast, exhibit only modest composition bias. Local indices further highlight stark geographic heterogeneity, including pronounced pandemic-related declines in San Francisco and only a modest post-pandemic recovery in Manhattan. Overall, our frame-work provides a more accurate measure of underlying commercial rent fundamentals by accounting for quality composition effects.
Data Foundation
Our main data source is commercial real estate lease data compiled by CompStak between 2010 and 2025. CompStak compiles a national database of commercial real estate leases by crowd-sourcing informa- tion from a network of verified brokers, appraisers, and researchers. The accuracy of every rental lease is verified by expert analysts and machine learning algorithms. For each lease, the data records, for example, the square footage being leased, the starting rent and rent schedule over the lease term, concessions includ- ing tenant improvements and free rent, the type of space being leased (e.g., office, industrial, retail), the lease type, the building address and its physical characteristics, and the dates on which the lease is signed and when it commences. Our main measure of rent is net effective rent per square foot. We also construct indices for starting rents, tenant improvements, and free rents. To assess the coverage of CompStak, we compare the occupied office inventory in our data to the occupied office inventory reported by Cushman & Wakefield. In major markets like Manhattan and San Francisco, CompStak covers around 90% of the office market, while in smaller markets the coverage is lower.
Market Definition
We estimate indices for markets across the office, retail, and industrial sectors. Our starting universe is the set of all U.S. metropolitan and micropolitan statistical areas (MSAs) that appear in our raw data, totaling 920 distinct MSAs as defined by the U.S. Office of Management and Budget.
To include a market in our estimation, we impose a minimum data-density requirement. A market must have at least five observations in a given quarter, with no more than five quarters falling below this threshold over the entire estimation window. The estimation period for each market is defined as the longest continuous span that (1) begins in 2010 or later, (2) extends through the most recent quarter, and (3) has its first quarter starting before 2018. These criteria ensure sufficient data coverage and minimize discontinuities while retaining the maximum feasible time series for each market.
For all three sectors, we construct indices at the MSA level, using all leases that fall within each MSA. For the office and retail sectors, we additionally estimate indices at the city level for all cities that satisfy the same data requirements. City-level markets are defined using CompStak's city designations, and city-level indices are estimated alongside (rather than as a substitute for) their corresponding MSA-level indices: a city that meets the sample thresholds appears both as its own city-level market and as part of the broader MSA market.
Market inclusion is determined automatically by the estimation pipeline based on data volume and consistency requirements, not manual selection, which is why coverage may vary by sector. Indices are generated only for market–sector combinations with sufficient data volume and consistency to support stable, quality-adjusted estimation.
Hierarchical Geographic Fixed Effects (HGFE) Methodology
To construct a quality-adjusted commercial rent index, we estimate a weighted least squares (WLS) hedonic regression using CompStak commercial real estate lease-level transaction data. Our approach employs a hierarchical geographic fixed effects framework that controls for both observable and unobserved quality differences across leases.
Main Regression Equation
For the national sample, we estimate:
where Rijst is the net effective rent per square foot in nominal dollars for lease i in building j of space type s (Office, Retail, Industrial) executed at time t. HGFEj is our hierarchical geographic fixed effect, Xijst is a vector of lease and building characteristics, and each lease is weighted by its transaction square footage.
Hierarchical Geographic Fixed Effects
A central contribution of our methodology is the construction of hierarchical geographic fixed effects. Given that the CompStak data provides a geocode (latitude and longitude) of the building, we map each lease to increasingly aggregate geographical areas: Census blocks, ZIP codes, Cities, and Metropolitan Statistical Areas (MSAs). In the HGFE framework, each lease is assigned to the most granular geography with at least five observations:
- Building FE: when there are at least five leases present in the sample for that building
- Block FE: if there are not five leases in the building but there are at least five leases in the Census block
- ZIP code FE: otherwise
- City FE, MSA FE, State FE: if all else fails
This structure maximizes the granularity of geographic controls while maintaining statistical stability. We are able to include a building FE for over half of all lease observations and either a building or a block FE for three quarters of lease observations.
Observable Characteristics
Variables controlled for include:
- Log lease size (square footage)
- Log building size (square footage)
- Log lease term (months)
- Log renovation-adjusted building age (years)
- Log average floor occupied
- Clear height (for industrial properties)
- Lease type (NNN, full service, modified gross, net, net of electric, gross)
- Transaction type (new lease, renewal, expansion, extension)
- Tenant industry
- Building class (A, B, C)
- Space subtype
- Central Business District (CBD) indicator
Index Normalization
To convert the time fixed effects into an index that tracks rent changes over time in nominal dollars per square foot, we compute:
where t₀ is the base period to which we normalize our index (2019Q4). We normalize the rent index in the base period equal to the average of the raw rent series for that space type in the base period. This ensures that the index reflects quality-adjusted rent movements relative to the average rent level in 2019Q4.
What "Quality-Adjusted" Means
"Quality-adjusted" means that our index controls for changes in the composition of properties being leased over time. Simple averages of rents can be misleading because they don't account for whether higher rents reflect genuine market movements or simply a shift toward leasing higher-quality properties. Our methodology holds quality constant by controlling for both observable characteristics (like building size, age, location) and unobservable characteristics (through hierarchical geographic fixed effects), allowing us to isolate true market-driven rent dynamics from compositional shifts.
For example, if office rents appear to increase because more leases are being signed in premium Class A buildings in prime locations, our quality-adjusted index will show a more modest increase (or even a decline) once we account for the shift toward higher-quality properties. This provides a more accurate measure of underlying market fundamentals.
Confidence Intervals
Each index value is accompanied by a 95% confidence interval, reflecting the statistical uncertainty in our estimates. Wider intervals typically indicate periods or markets with fewer transactions or greater variation in property characteristics.
Update Schedule and Data Revisions
National Indices: Updated monthly
Market-Level Indices (Cities and MSAs): Updated quarterly
Historical Revisions: Historical index values may be revised as new lease transactions are added to the dataset. CompStak's data collection is ongoing, and leases may be reported with a delay. As new data becomes available, we re-estimate the entire index series to incorporate these additional observations, which can result in revisions to previously published values. This ensures that our indices reflect the most complete and accurate data available.
Series End Dates: The index series ends at the most recent period for which we have sufficient data coverage. For national indices, this is typically the most recent month with adequate transaction volume. For market-level indices, this is typically the most recent quarter meeting our minimum data requirements (at least five observations per quarter). Series may end at different dates across markets depending on local data availability and transaction volume.
Research Team
Boaz Abramson
Assistant Professor of Business
Finance Division
Columbia Business School

Gaurav Choudhary
Staff Associate I in the Faculty of Business
Finance Division
Columbia Business School

Joel Kim
Staff Associate I in the Faculty of Business
Finance Division
Columbia Business School

Tomasz Piskorski
Edward S. Gordon Professor of Real Estate
Finance Division
Columbia Business School

Ziyi Qiu
Staff Associate I in the Faculty of Business
Finance Division
Columbia Business School

Stijn Van Nieuwerburgh
Earle W. Kazis and Benjamin Schore Professor of Real Estate
Finance Division
Earle W. Kazis and Benjamin Schore Professor of Real Estate
Paul Milstein Center for Real Estate
Co-Director
Paul Milstein Center for Real Estate
Columbia Business School

Wayne Yu
SVP, Data
CompStak

