Reconciliation Testing: A Public-Records Methodology for Auditing Institutional Policy Coherence
Reconciliation Testing: A Public-Records Methodology for Auditing Institutional Policy Coherence
Joseph Phelan
Global Center for Advanced Studies, Dublin
F'nAround Media, Chicago, Illinois (fnaround.com)
Working paper draft, September 2026 prepared for comment, not for citation as published work
Developed in the public eye as a working research article at fnaround.com/academic
Abstract
Governments rarely speak through a single institutional voice. Legislatures enact statutory purposes; executives translate those purposes into political commitments; agencies interpret and enforce rules; attorneys defend state action in litigation; records officers decide what the documentary record will show. When those representations diverge, the researcher faces a prior question that the implementation literature does not answer: how to establish, as evidence rather than impression, that the divergence is real, how large it is, and whether any institutional mechanism reconciles it. This article introduces reconciliation testing, a qualitative public-records methodology for auditing the coherence of governmental meaning across institutional voices. The protocol proceeds in seven steps: fixing an authoritative commitment, mapping the voices authorized to interpret it, pre-registering observable implications, collecting representations through records requests and proceedings, comparing representations against the commitment, classifying gaps against a defined taxonomy, and attempting reconciliation before any divergence counts as a finding. Grounded in case-study logic, triangulation, and process tracing, the method treats persistent non-convergence as the result rather than the failure of triangulation, and builds the strongest rival explanations, lawful confidentiality, role-based narrowing, sequencing, resource constraint into the protocol itself. Two worked demonstrations from Illinois cannabis governance show the method in operation and show its restraint: gaps are reported exactly as strong as the surviving evidence, and no stronger.
Keywords
reconciliation testing; public records; freedom of information; institutional coherence; policy implementation; qualitative methods; triangulation; Illinois cannabis regulation
1. Introduction
Every researcher who has filed public-records requests against more than one office of the same government knows the experience this article formalizes. The statute says one thing. The governor's press release says something adjacent. The agency's enforcement file says something narrower. The litigating attorney says something narrower still. None of these actors is necessarily lying and none is necessarily wrong; each occupies a role that licenses its emphasis. The question is whether the roles add up to an account a citizen could follow and what it means, analytically, when they do not.
The implementation literature asks why policy changes as it travels through institutions. Reconciliation testing asks the prior question: whether the government's own representations of a single commitment can be laid side by side without contradiction, and what the researcher is entitled to conclude when they cannot. The distinction matters because most empirical methods in policy research sample one institutional voice and treat it as informative about the whole. A survey asks officials what the policy is; an interview asks implementers how it works; a statistical analysis treats agency outputs as realizations of legislative intent. Each of these approaches assumes, implicitly, a unitary object of study. When the phenomenon of interest is the relationship among voices, that assumption is not a simplification. It is the thing being tested.
The method proposed here uses the ordinary instruments of democratic transparency public-records requests, administrative complaints, pleadings, testimony, published agency materials as data-generating tools inside a structured protocol. It is qualitative but not impressionistic: every step produces a documented output, every finding must survive an explicit attempt at reconciliation, and every conclusion is stated in falsifiable form. Two demonstrations from Illinois cannabis governance, drawn from the author's primary-record research program (Phelan, 2026), show both what the method finds and what it refuses to claim.
2. The Methodological Problem
Policy coherence is usually treated as a property of instruments: do the tools fit together, do the agencies coordinate, does implementation reproduce intent. Those are important questions, but they skip the documentary one. Before asking whether a government's actions cohere, we can ask whether its statements do whether the legislature's purpose, the executive's framing, the agency's enforcement posture, and the litigator's characterization describe the same commitment in mutually compatible terms. If they do not, the incoherence is not a hypothesis about hidden behavior. It is visible in the record, provided someone assembles the record.
Public-records regimes make that assembly possible. Freedom-of-information laws compel governments to produce documents on request, generating an evidence base that is contemporaneous, attributable to specific institutional actors, and independent of researcher interpretation at the point of collection. Investigative journalists have worked this evidence for decades. What the scholarly literature lacks is the formalization: a protocol that converts records requests from an investigative tactic into a research design with stated steps, validity criteria, and outputs a third party could reproduce. Reconciliation testing is that formalization not a new way to file FOIA requests, but a way to make FOIA requests bear the weight of a method.
3. Theoretical Foundations
The method draws on three established traditions. What each supplies, and what each leaves for the present article to add, is stated explicitly.
Case-study logic supplies the architecture. Following Yin (2018), each policy episode is treated as a case and each institutional representation as an embedded unit of analysis; the chain of evidence questions to collection to findings, each link documented is the structural template. Eisenhardt (1989) licenses the move from a small number of deeply documented episodes to portable concepts, which is how the gap taxonomy in Section 5 earns its keep. Flyvbjerg (2006) supplies the defense of the revelatory case: a strategically chosen episode can falsify a general proposition, including the proposition that a government's stated commitments cohere. What this tradition does not supply is the diagnostic operation itself a repeatable procedure for laying representations side by side and ruling on their compatibility. That operation is the present contribution.
Triangulation supplies the evidentiary logic, inverted. Denzin (1978) distinguished data, investigator, theory, and methodological triangulation, all organized around the premise that convergence across sources increases confidence. Reconciliation testing is a data-triangulation design in which the multiple sources are not independent observations of one phenomenon but the phenomenon itself: the divergence among institutional voices is the finding. The inversion is deliberate. Where classical triangulation treats non-convergence as a problem to resolve, this method treats persistent non-convergence after genuine reconciliation attempts as the result.
Process tracing supplies the discipline of pre-registered expectations. Bennett and Checkel (2015) and George and Bennett (2005) formalized the practice of deriving observable implications from a claim and testing them against within-case evidence. Reconciliation testing borrows the logic directly: before any records are collected, the researcher states what each voice should produce if the commitment is coherent, then checks. The difference is that the claim under test is diagnostic rather than causal. The method asks whether representations align, not why they diverge. Causal explanation is a separate study, and the method is explicit about not doing its work.
The novel operation the one none of the three traditions provides is the reconciliation attempt. Before any divergence counts as a finding, the researcher must search for an institutional bridge: a legal distinction between roles, a sequencing explanation, a confidentiality constraint, a jurisdictional boundary, or any other mechanism that would render the representations compatible. Only divergence that survives that search is reported as a gap. This is what separates the method from adversarial journalism. The strongest rival explanation is not a footnote; it is Step 7.
4. The Protocol
Seven steps. Each produces a documented output, and the sequence is designed so that a third party with the same public-records regime could audit every one of them.
Step 1: Fix the commitment
Identify a policy commitment stated in an authoritative source statutory text, a signed executive order, an enacted regulation, an official gubernatorial statement of purpose. Record the exact language and its source. The commitment must be specific enough to generate observable implications. "Protect children" is not fixable; "no cannabis advertising within 1,000 feet of a school" is. Vague commitments produce vague findings, and the method refuses to launder the difference. Output: a commitment statement with citation.
Step 2: Map the voices
List the institutional actors authorized to interpret, implement, enforce, defend, or document the commitment: the legislature, the executive, the implementing agency, the enforcing agency, the government's litigating attorneys, oversight bodies, the public-records function itself. For each voice, note its legal role, because role differences are the most common legitimate source of divergence and the reconciliation attempt in Step 7 will need them. Output: a voice map with role descriptions.
Step 3: Generate observable implications
For each voice, state in advance what it should say or produce if the commitment is coherent. An agency charged with enforcing an advertising restriction should hold enforcement records, guidance documents, or complaint dispositions consistent with the restriction. A litigating attorney defending the statute's purpose should characterize that purpose in terms reconcilable with the statutory text. These implications are written before records are collected, which is what prevents the researcher from fitting expectations to whatever the agencies happened to produce. Output: a pre-registered implication table.
Step 4: Collect representations
Gather each voice's representations through public-records requests, administrative complaint filings and their dispositions, court pleadings, sworn testimony, published guidance, and official statements. Requests must be specific, dated, and reproducible: another researcher should be able to file the same request and receive the same response set. Log every request, response, denial, delay, and fee assessment, because the records process itself is data how an institution handles the asking is evidence about the institution. Where the research program maintains a public log of requests and filings, that log becomes part of the auditable record (F'nAround, 2025). Output: a dated collection log with all responsive materials.
Step 5: Compare
Place each voice's representations alongside the commitment and the pre-registered implications, and mark convergences and divergences with document-level citations. Maintain three evidentiary grades and do not confuse them: allegation (a claim made by one party), representation (a statement attributable to an institutional actor), and finding (an official determination). An allegation never becomes a finding by repetition. Output: a comparison matrix.
Step 6: Classify gaps
Assign each divergence a gap type from the taxonomy in Section 5. Classification is what keeps the method from collapsing every discrepancy into one undifferentiated charge of incoherence; it forces the researcher to say what kind of gap was found, which is what makes gap types comparable across cases. Output: a classified gap inventory.
Step 7: Attempt reconciliation
For each classified gap, search for the bridge. Read the statute for role-based distinctions. Check confidentiality rules, litigation sequencing, jurisdictional boundaries. Ask the agency for comment where feasible. Document the search whether or not it succeeds. Gaps that survive are reported as findings; gaps that reconcile are reported as reconciled, with the bridge described because a reconciled gap is also a result, and a method that only publishes divergence is selecting on its dependent variable. This distinction is the method's central validity claim. Output: a reconciliation log and final findings.
5. Gap Taxonomy
Five gap types, developed from the Illinois demonstrations below and proposed as portable. The taxonomy will need revision as the method travels; that revision is how it matures.
The Reconciliation Gap. Two or more institutional voices make materially divergent representations of the same commitment, and no observable mechanism explains how they fit together. The core finding type and the hardest to earn, because Step 7 must fail first.
The Enforcement Translation Gap. A commitment stated at the legislative or executive level has no observable enforcement pathway beneath it: no guidance, no complaint dispositions, no enforcement records. The policy exists in proclamation but not in practice. This is the gap between what the state says and what the state does, measured in documents rather than asserted in prose.
The Records Gap. A request that should on its face return responsive documents instead returns a nonexistence determination, a blanket exemption claim, or indefinite delay. The gap is between the documentary architecture the commitment implies a restriction that is enforced must generate enforcement records and the records the institution acknowledges holding.
The Temporal Gap. Representations that were coherent at one point diverge over time with no intervening authoritative change in the commitment. The analytical question is which moved: the policy or the telling.
The Role-Strain Gap. Divergence traceable to the conflicting demands of an actor's institutional roles a litigating attorney whose duty of zealous defense pulls against the statute's stated purpose, for example. Role-strain gaps are the most likely to reconcile, because the role conflict is itself the bridge. They are classified separately so that structural role explanations are not mistaken for incoherence, and so that genuine incoherence cannot hide behind role language.
6. Worked Demonstration I: Child Protection in Illinois Cannabis Governance
Illinois law restricts cannabis advertising likely to appeal to minors and bars cannabis advertising within 1,000 feet of schools, playgrounds, recreation facilities, child-care centers, public parks, public libraries, and certain arcades. In June 2026, the General Assembly sent and Governor J.B. Pritzker signed Senate Bill 3222, wide-ranging cannabis legislation whose centerpiece was the regulation of intoxicating hemp products delta-8 THC, THC-P, HHC previously sold outside the regulated market. The law immediately banned sales of intoxicating hemp to persons under 21, required child-resistant packaging, and prohibited packaging and advertising designed to appeal to children, including products packaged to resemble common snacks and candy (Gov. Pritzker Newsroom, 2026).
The Governor's framing fixed the commitment in unusually explicit terms, and it is worth quoting because the executive voice itself set the bar the method then tests. Celebrating the law on July 2, 2026, at Sway Dispensary in Chicago, Pritzker said: "While our regulated cannabis market has been operating under those strict standards, using a federal loophole, an entirely separate market emerged for intoxicating hemp products, creating real risks to the public, especially for our kids." He added: "Parents shouldn't have to worry that a product containing intoxicating levels of THC is deceptively packaged as a bag of candy in a convenience store" (Gov. Pritzker Newsroom, 2026; NPR Illinois, 2026). The commitment, fixed in Step 1, is therefore twofold: a statutory youth-protection rule for cannabis marketing, and an executive claim that the licensed cannabis market operates under strict child-protective standards strict enough to serve as the contrast case against the hemp market's failures.
The voice map (Step 2) included the legislature (the statutory restrictions and SB 3222), the executive (the Governor's framing), the Department of Agriculture and the Department of Financial and Professional Regulation (licensing and enforcement), and the public-records function. The pre-registered implications (Step 3) were straightforward: enforcement guidance interpreting the youth-appeal restriction, complaint intake and disposition records showing the restriction applied to licensed-cannabis marketing, and interagency communication about marketing enforcement.
In 2025, the author filed a public-records and complaint sequence concerning alleged licensed-cannabis marketing at an all-ages Chicago event the licensed market, held by the executive voice to strict standards, tested against a concrete allegation (Phelan, 2025, October–November; collection log in F'nAround, 2025). Collection (Step 4) produced a complaint record, agency correspondence, and records responses. Comparison (Step 5) showed the executive voice articulating the protective principle at full volume while the enforcement voice, confronted with a specific licensed-cannabis marketing allegation, produced investigation confidentiality, limited disposition information, and interpretive uncertainty about how the restriction applied.
Classification (Step 6) identified a candidate Reconciliation Gap between the executive's framing and the enforcement record, and a candidate Records Gap in the thinness of the enforcement file. The reconciliation attempt (Step 7) took the bridges seriously, because several are real. Investigation confidentiality is a lawful constraint, not an evasion. Hemp and licensed cannabis sit under different statutory provisions. A single complaint proves nothing systemic. And the Governor's "strict standards" remark described the regulatory framework, not any particular enforcement outcome a framework can be strict on paper while a given file goes quiet. The demonstration therefore reports the gaps as observed and unreconciled on the current record, not as proof of incoherence. The method's output here is deliberately unsatisfying: a documented divergence, a documented search for the bridge, and an honest statement that the bridge was not found in the materials at hand. That is what a diagnostic method owes its reader.
7. Worked Demonstration II: Social Equity Purpose and Litigation Posture
The Cannabis Regulation and Tax Act states social-equity purposes explicitly, including reducing barriers to ownership for communities disproportionately harmed by cannabis prohibition. The executive voice has repeatedly described equity as a design purpose of legalization not a side benefit, but part of what the statute is for. In litigation over the licensing process, however, the Illinois Department of Agriculture, represented by the Attorney General's office, characterized the advancement of Social Equity Applicants in terms narrower than the executive's design-purpose framing (Phelan, 2026).
The structure of the test is the same. The commitment (Step 1) is the statutory equity purpose. The voices (Step 2) are the legislature, the executive, the agency, and the litigating attorneys. The observable implication (Step 3): the State's litigating characterization of the statute's purpose should be reconcilable with its executive description of that same purpose not identical, reconcilable. Collection (Step 4) drew on pleading excerpts, the December 22, 2025 deposition testimony, and the October–November 2025 records sequence, held as primary records by the author and indexed in the public records log (F'nAround, 2025).
Comparison (Step 5) identified the divergence: purpose language at the executive level, narrower characterization at the litigation level. Classification (Step 6) produced a candidate Reconciliation Gap and a candidate Role-Strain Gap, and here the role-strain reading deserves its full weight. Litigating attorneys owe their client agency zealous defense; adversarial proceedings reward narrow readings of statutory purpose; a pleading is a weapon, not a policy speech. That is not a cynical observation it is how the role is designed, and the taxonomy classifies it separately for exactly this reason.
Whether the bridge carries the full span is the open question, and the method leaves it open on purpose. How narrow was the litigation characterization relative to the statutory text? The primary records are cited, the reasoning is shown, and the reader can disagree with the classification which is precisely what a falsifiable method permits. The output, again, is not an accusation. It is a disciplined claim about two sets of words describing one statute, with the strongest excuse for the difference built into the analysis before the conclusion was drawn.
8. Validity, Reliability, and Rival Explanations
Five challenges, each met by a design feature rather than a disclaimer.
Researcher bias. The researcher chooses the commitment, maps the voices, and classifies the gaps three discretionary acts before any evidence is weighed. The pre-registered implications of Step 3 constrain post hoc fitting, and the public collection log of Step 4 permits independent audit. Where resources allow, a second coder classifies gaps independently and disagreements are reported rather than resolved quietly.
Selection bias. Nothing in the protocol stops a researcher from choosing episodes already believed to show incoherence. The defense is comparative application: run the protocol on episodes where coherence is expected, and treat fully reconciled gaps as publishable findings. A method that only reports divergence is selecting on its dependent variable; the reconciliation log exists to prevent exactly that.
The confidentiality rival. Government silence is often lawful. Investigation confidentiality, attorney-client privilege, deliberative-process exemptions, litigation sequencing each can produce a records gap that reflects legal constraint rather than incoherence. Step 7 exists for this rival and must be worked in good faith, not performed as ritual. The demonstrations above show what good faith looks like: bridges considered, some found load-bearing, the remainder reported as unreconciled rather than unexplained.
The competence rival. Divergence may reflect error, bad records management, or understaffing rather than anything structural. The method does not adjudicate among these causes; it reports the divergence and leaves causal attribution to designs built for that purpose. This is a stated boundary of the method, not a defect in it.
Reactivity. Filing requests and complaints as a research instrument intervenes in the system under study, and agencies may handle a known researcher-litigant differently from an anonymous requester. The method cannot eliminate this; it can only disclose it, which Section 9 requires.
Reliability rests on documentation rather than on the researcher's judgment. Every step leaves a citable artifact the commitment statement, the voice map, the implication table, the collection log, the comparison matrix, the gap inventory, the reconciliation log. A third party with the same protocol and the same public-records regime should be able to reproduce the collection and argue with the classifications. This article was developed as a working paper in the public eye, with the research program's working articles maintained as a public repository open to critique, replication, and competing interpretations (F'nAround, n.d.). That is not a stylistic choice. It is the method's reliability claim, enacted.
9. Ethics and Positionality: The Researcher as Requester
The researcher in this method is not a passive observer but an active participant in the records system filing the requests, lodging the complaints, sometimes litigating the denials. Four principles govern that participation.
First, no deception. Requests must be genuine, complaints must allege genuinely held concerns, and the researcher must not misrepresent identity or purpose beyond what the law permits any requester. The method's findings depend on the records system functioning as it functions for everyone; gaming it corrupts the instrument.
Second, proportionality. Public-records systems are strained common resources, and requests should be narrowly drawn to the protocol's needs rather than weaponized as denial-of-service. The collection log lets the reader judge whether the burden was reasonable. The principle cuts both ways: where agencies answer lawful requests with improper delay or silence and the requester escalates through Public Access Counselor complaints or other lawful channels, that escalation history is part of the auditable record, not background grievance (F'nAround, 2025).
Third, positionality stated plainly. Where the researcher is also complainant, litigant, journalist, or market participant, that fact goes up front, not in a footnote. The demonstrations above involve the author's own complaints, litigation, and investigative reporting through F'nAround Media. The method treats that involvement as data about access rather than disqualification the researcher-as-requester sees the system from inside the request queue, which is a vantage, not just a bias but the reader is entitled to discount accordingly, and the disclosure must make that possible.
Fourth, care with persons. Gap findings concern institutional representations, not individual wrongdoing, and the method must not be used to imply corruption, bad faith, or illegality from divergence alone. Where individuals are named in underlying records, reporting follows the minimum necessary. The method audits voices, not people.
10. Limitations and Boundary Conditions
The method requires a functioning public-records regime. Where transparency law is unenforceable or nonexistent, Steps 4 and 7 cannot operate and the method does not apply. That excludes a large part of the world, and the exclusion should be stated rather than finessed.
The method is diagnostic, not explanatory. It establishes that representations diverge and that no bridge was found in the materials at hand. It does not establish why. Researchers who want causes need other designs.
The method is slow. A single episode can absorb dozens of requests, months of waiting, administrative appeals, and meticulous logging. The labor is the validity there is no shortcut version that keeps the guarantees but it bounds how many episodes one researcher can run.
The taxonomy is provisional. Five types from two demonstrations in one state is a beginning, not a system. Other domains and other countries will revise, merge, or extend it. The method expects this.
Finally, the method inherits its sources' limits. Records responses show what institutions disclose or are compelled to disclose. A coherent government with an obstructive records function would generate records gaps the method reports as gaps. Step 7 mitigates this, but the mitigation is interpretive, and interpretations can be wrong. The method's honesty about this is part of its claim to be taken seriously.
11. Conclusion
Reconciliation testing does not explain incoherence, reform institutions, or settle what policy should be. It does one thing: it converts the democratic intuition that a government's many voices should be mutually intelligible into a procedure for checking whether they are, with the check documented at every step and the benefit of every doubt given before doubt is reported as a finding. The seven steps make public-records work reproducible. The taxonomy gives divergence a vocabulary precise enough to argue with. The reconciliation requirement means no finding in this method was reached without first trying to defeat it.
Whether the method earns a place in the qualitative toolkit depends on whether other researchers can run it and whether their gap classifications survive scrutiny. It was built to be run, and built to be argued with. That is the whole of its ambition.
References
Bennett, A., & Checkel, J. T. (Eds.). (2015). Process tracing: From metaphor to analytic tool. Cambridge University Press.
Creswell, J. W., & Creswell, J. D. (2018). Research design: Qualitative, quantitative, and mixed methods approaches (5th ed.). SAGE.
Denzin, N. K. (1978). The research act: A theoretical introduction to sociological methods (2nd ed.). McGraw-Hill.
Eisenhardt, K. M. (1989). Building theories from case study research. Academy of Management Review, 14(4), 532–550.
Flyvbjerg, B. (2006). Five misunderstandings about case-study research. Qualitative Inquiry, 12(2), 219–245.
F'nAround. (n.d.). Academic research articles. Retrieved September 11, 2026, from https://www.fnaround.com/academic/
F'nAround. (2025). FOIA. Retrieved September 11, 2026, from https://www.fnaround.com/articles/foia
George, A. L., & Bennett, A. (2005). Case studies and theory development in the social sciences. MIT Press.
Gov. Pritzker Newsroom. (2026, July 2). Gov. Pritzker celebrates landmark legislation to keep the Illinois cannabis market safe and growing [Press release]. https://gov-pritzker-newsroom.prezly.com/gov-pritzker-celebrates-landmark-legislation-to-keep-the-illinois-cannabis-market-safe-and-growing
NPR Illinois. (2026, July 2). New regulations on intoxicating hemp are 'long overdue,' Pritzker says. https://www.nprillinois.org/health-harvest/2026-07-02/new-regulations-on-intoxicating-hemp-are-long-overdue-pritzker-says
Phelan, J. (2026). When the state speaks with multiple voices: Institutional policy incoherence in Illinois cannabis governance [Working article]. F'nAround. https://www.fnaround.com/academic/voices
Yin, R. K. (2018). Case study research and applications: Design and methods (6th ed.). SAGE.
────────────────────────────────────────
Author note. Joseph Phelan is a PhD researcher at the Global Center for Advanced Studies, Dublin, and the publisher of F'nAround Media (fnaround.com), a Chicago-centered investigative outlet. Suggested citation: Phelan, J. (2026). Reconciliation testing: A public-records methodology for auditing institutional policy coherence [Working paper]. F'nAround Media. Correspondence via fnaround.com. Primary records cited in the demonstrations are held by the author; a public index of records requests and filings is maintained at https://www.fnaround.com/articles/foia, and working versions of this article and related papers at https://www.fnaround.com/academic/. Web citations verified September 11, 2026.