My work focuses on developing new data infrastructure for empirical corporate governance, including large-scale datasets constructed from primary governance documents such as charters and bylaws. I integrate these empirical tools with doctrinal insight and computational methods to test theoretical claims about private ordering, regulatory spillovers, and the institutional design of law. In parallel, I develop frameworks for responsibly deploying natural language processing and large language models in legal research—clarifying where such tools can enrich legal inquiry and where their use requires careful governance and transparency.
​​​
My current projects include the following:
​“The Other Delaware Effect”
This paper examines the effects of Delaware’s 2015 ban on fee-shifting provisions in corporate charters and bylaws, a significant legislative intervention in corporate law aimed at curbing managerial powers. The Delaware Supreme Court had approved these provisions just one year earlier as part of a series of measures aimed at curbing shareholder litigation. Because of their perceived substantial potential to reduce wasteful litigation, the Delaware legislature’s ban led many to predict an exodus of corporations from Delaware and the continued spread of fee-shifting provisions in other states.
Contrary to these predictions, this study finds that the ban did not trigger a significant departure of corporations from Delaware. More importantly, it also documents a sudden decline in the adoption of fee-shifting provisions outside Delaware, where most states did not enact similar prohibitions. The paper argues that this decline is most likely due to a spillover effect: Adoptions declined not because Delaware’s ban on fee-shifting provisions coincided with a shift in market sentiment against these provisions, but because this ban stymied their adoption nationwide. The paper explores the mechanisms through which Delaware’s corporate law can influence governance practices elsewhere, including shareholder empowerment and law firms’ reluctance to recommend provisions that have been outlawed in Delaware. It concludes that Delaware’s legal leadership might be able to set informal norms that influence the behavior of important gatekeepers and other actors in the corporate governance ecosystem, effectively constraining the actions space of managers of corporations incorporated elsewhere.
By exploring these mechanisms, the paper contributes to a deeper understanding of the forces that shape corporate governance practices in the United States. It offers insights into the ways Delaware’s rules affect corporate governance practices that subsequently reverberate throughout the U.S. corporate landscape. Besides exploring important and previously overlooked aspects about Delaware’s role as a standard setter in today’s corporate law world, it also adds to our understanding of the diffusion of corporate governance innovations and regulatory competition. Additionally, the paper demonstrates the potential of artificial intelligence tools, such as large language models, to revolutionize empirical legal research by automating the extraction of legally significant provisions from corporate documents, allowing researchers to investigate previously underexplored questions at scale.
​
​“DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods”
Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from organizational documents. The benchmarks pair randomly sampled corporate charters and bylaws with high-quality human annotations covering a range of governance provisions commonly studied in empirical work.
Using these datasets, the paper evaluates several large-language-model extraction pipelines that vary in prompt design, task decomposition, and document handling. The underlying task consists of a set of document-level binary classification problems, one for each governance variable. The results show that automated extraction is feasible at a high level of accuracy for many provisions, with median performance near the upper bound across approaches. At the same time, performance varies systematically across variables, with a small number of provisions accounting for most of the remaining errors. More elaborate prompting strategies and cascading pipelines do not consistently improve performance for frontier models, but substantially narrow the gap between frontier and efficiency-oriented models in some settings, suggesting that pipeline design can partly substitute for model capability.
By providing a standardized benchmark and a systematic evaluation of extraction methods, the paper demonstrates that current frontier models can extract legally meaningful information from complex corporate documents with high accuracy and suggests an important future role for automated feature extraction in constructing corporate governance datasets.
​
​“Measuring Corporate Governance with Large Language Models”
Empirical corporate governance research depends on structured data extracted from legal text, a translation exercise that has traditionally required costly and difficult-to-scale human coding. Using the DECODEM benchmarks developed in companion work, this paper asks what role LLM-based extraction can play in governance data production: whether automated extraction routines perform on par with realistic human coding workflows and whether model-human disagreement can guide audits of human-coded data. The paper's main findings are based on a blind adjudication of observations on which human coding and the automated extraction routines disagree.
The analysis supports three main findings. First, for many governance variables, LLM-based extraction performs at levels at least comparable to feasible human coding workflows, at less than two percent of the marginal cost of human coding. Second, different extraction routines often make only partially overlapping errors, creating value from ensemble rules that aggregate multiple routines. In the setting studied here, labels generated through ensemble rules agree with the adjudicated labels substantially more often than initial human labels did. Third, substantial variation across variables remains, so LLM-generated governance variables should not be used for downstream research without validation against hand-collected data.
The results also carry an implication for settings where fully automated coding is not appropriate: because human and model errors also overlap only weakly, flagging observations on which any of six routines disagrees with the initial human label would surface more than ninety percent of observed human coding errors while requiring review of roughly seven percent of observations. Together, the results point towards a new model of data production in which human judgment defines and validates governance variables, while LLM-based extraction supplies scalable coding and targeted quality control. labels nor the model outputs.
​
