What sampling methods do state sales tax auditors use?
At mid-market transaction volume, a full review is mechanically impossible. A $40M Shopify Plus brand processing 200,000 orders a year produces roughly 600,000 transactions over a three-year audit period. The auditor does not review them. They sample, compute an error rate, and project. Four sampling methods govern that work, and each one produces a different projection from the same underlying data.
The four methods, defined per the state procedure manuals auditors actually work from:
- Block sampling. The auditor selects a contiguous time period, typically one month or one quarter, reviews 100% of transactions in that block, computes an error rate, and extrapolates across the entire audit period.[1] Cheap to execute, statistically indefensible. A block over a promotional period or a since-fixed configuration error projects that anomaly across years.
- Statistical or random sampling. The auditor draws a random sample (often 250 to 400 transactions) from the entire audit-period population using a documented sampling plan, computes an error rate within a stated confidence interval (commonly 80% or 90%), and projects to the population using mathematical extrapolation. California CDTFA Publication 76, Texas Comptroller Audit Procedures Manual Chapter 8, and New York DTF Publication 130-D each codify procedures for statistical sampling.[1][2][3]
- Stratified sampling. A refinement of statistical sampling. The population is split into strata, typically by dollar amount (for example, $0 to $100, $101 to $1,000, $1,001 to $10,000, $10,001 and above), and each stratum is sampled independently. High-dollar strata may be sampled at 100% (a "detail" stratum) while smaller transactions are sampled at a lower rate. Stratification reduces the variance in the projection and is the methodology most state procedure manuals prefer for high-volume audits.[3]
- Judgmental sampling. The auditor uses professional judgment to select transactions deemed high-risk: large exempt sales, transactions with related parties, unusual customer types, manual overrides. Not random. Findings within a judgmental sample apply to the sampled transactions only and are not extrapolated to a population.[2]
The distinction that determines the assessment's size is not which method, but whether the method is extrapolated across the population. Block, statistical, and stratified samples produce a projected error rate that runs against the full audit-period revenue. Judgmental samples do not.
When can a taxpayer elect statistical over block sampling?
The single most consequential decision at the opening conference is the sampling-method election. The controllers we have watched succeed ask for statistical sampling before the auditor frames a block-test period, because a block test on a bad month gets extrapolated across the whole audit period and a statistical sample does not.
Most state audit procedure manuals grant the taxpayer the right to request statistical over block sampling, but the election typically must happen before fieldwork. California's CDTFA Publication 76 outlines statistical sampling procedures and indicates the taxpayer may request the method.[1] Texas's Audit Procedures Manual Chapter 8 establishes statistical sampling as the default for high-volume audits while allowing block sampling by agreement.[2] New York DTF Publication 130-D documents the agency's statistical sampling methodology in detail and requires a sampling plan agreement before sample selection.[3]
The trade-offs that matter at the opening conference:
| Dimension | Block sampling | Statistical or stratified sampling |
|---|---|---|
| Defensibility | Low. No confidence interval, no bounded projection. | High. Mathematically bounded, replicable by the taxpayer. |
| Volatility | High. One bad month projects across years. | Low. Stratification dampens single-period anomalies. |
| Audit cost (auditor time) | Lower. Single block, less sampling setup. | Higher. Sample design, stratification, documentation. |
| Election timing | Default in some states unless taxpayer requests statistical. | Must be requested and agreed before fieldwork in most states. |
| Challenge surface at protest | Period representativeness only. | Population definition, stratification design, projection math, sample size. |
A controller who arrives at the opening conference without a position on sampling is letting the auditor pick the method that is fastest to execute. That is typically block. The election is in writing, in the sampling plan agreement that becomes part of the audit file, and changing it later is materially harder than agreeing to it the first time.
The data point that makes the election concrete is the brand's own transaction-level record. TaxCloud's reporting API exposes transaction-level tax calculation logs across every channel (Shopify, Shopify Plus, BigCommerce, QuickBooks Online), which is the population definition you bring to the opening conference. Without that record, the auditor's population definition stands unchallenged.
How sample-period selection drives the assessment
Sampling is where a 2% error rate becomes a six-figure number. The auditor samples a few hundred transactions, computes an error rate, and projects it across hundreds of thousands of transactions in the population. The projection, not the sample, is the assessment.
The mechanics, in order:
- The auditor defines the population (for example, all Texas-delivery sales in the 36-month audit period).
- A sample is drawn (block, random, or stratified) per the agreed sampling plan.
- The auditor reviews each sampled transaction for taxability errors and computes an error rate as a percentage of taxable sales reviewed.
- The error rate is projected against the population revenue to produce a tax base subject to additional tax.
- State tax rate, penalties, and interest are applied to the projected base.
A worked example. A $40M Shopify Plus brand operating in Texas, three-year audit period covering 300,000 Texas transactions and $40M in cumulative Texas-delivered revenue. The auditor draws a 300-transaction sample, finds taxability errors on 8 transactions totaling $4,000 in undercollected tax on $200,000 in sampled taxable sales. The sampled error rate is 2.0%. Projected across $40M in population sales, that produces an additional tax base of $800,000. With a 10% penalty and roughly 10% accumulated interest, the assessment moves into the high six figures from an underlying sample error count of 8 transactions.
The math is brutal because of the leverage. A controller looking at "8 bad transactions" cannot see the assessment. The auditor's projection is the assessment.
What this means for sample-period selection. The brands that succeed at this insist on a representative period, not a favorable one. Asking the auditor to sample only a strong-controls month invites a counterargument that the period is unrepresentative. The better posture is to argue for stratified sampling across the full audit period, so anomalies in any single month wash out through the stratification.
Configuration-driven errors are the recurring source of asymmetric exposure. A misconfigured product-taxability rule in Shopify Tax, a calculation gap during a marketplace integration cutover, or a temporary mapping error during a BigCommerce-to-Shopify Plus migration can produce a high error rate concentrated in a specific period. If that period lands in the auditor's block, the projection runs that error across the entire audit. The same error caught by stratified sampling shows up in one stratum and is bounded by the stratum's share of population revenue.
How the auditor moves from sample to projected assessment
The projection math is what turns sampling into the assessment number a controller actually defends at protest. Three components drive it, and each is a separate challenge surface.
Component 1: Sample size and confidence interval
The auditor's sampling plan states the desired confidence level (commonly 80% or 90%) and acceptable precision (for example, plus or minus 10% of the projected error). Sample size is computed from those targets and the population variance. New York DTF Publication 130-D and the Texas Audit Procedures Manual Chapter 8 both publish sample-size tables and computation methods.[2][3] A 300-transaction sample at 80% confidence with plus-or-minus 10% precision is roughly the floor for a population of 200,000 transactions or more.
Component 2: The point estimate
The error rate observed in the sample (taxability-error dollars divided by taxable-sales dollars reviewed) becomes the projected error rate. Stratification refines this: each stratum's error rate is applied to that stratum's population revenue, and the components are summed. Stratified projections nearly always come in lower than unstratified projections on the same data, because high-dollar transactions, which auditors typically review at higher coverage, carry their own measured error rate rather than inheriting the small-transaction rate.
Component 3: Confidence-interval bounds
A statistically sampled projection has both a point estimate and a confidence interval. The point estimate is what the auditor proposes. The lower bound of the confidence interval is what a taxpayer can argue toward at protest if the precision is questionable. NY DTF's guidance acknowledges this: an assessment outside the precision range raises questions the auditor's sampling plan has to answer.[3]
The brands that defend these assessments well track all three components separately. The sample-size question is "did the auditor draw enough to support the precision they claim?" The point-estimate question is "is the error rate consistent with stratification?" The confidence-interval question is "what is the defensible lower bound, and how does it compare to the assessed amount?"
Each of these is a documented challenge surface at protest. Each requires the taxpayer's own population data to argue, not just the sampled transactions.
How to challenge a sampling result
The brands that challenge a bad sampling result successfully do it on population representativeness, not on individual transactions. If the sample period over-weights a promotional month or a since-fixed configuration error, the extrapolation is attackable. If the population definition includes transactions that should not be in scope, the projection runs against an inflated base.
Three challenge vectors carry weight at protest:
- Population definition. Did the auditor include transactions that do not belong in the population: marketplace-facilitated sales where the marketplace is the responsible collector, properly exempt transactions with valid certificates on file, sample-period transactions outside the actual audit window? A 5% error in the population definition runs straight through to a 5% error in the projection.
- Stratification design. Was the population stratified at all, and were the strata appropriately sized? An unstratified projection that applies a high-dollar-stratum error rate to small-dollar transactions (or the reverse) overstates exposure. NY DTF Publication 130-D specifies that the burden of demonstrating proper stratification rests with the auditor, and a sampling plan that does not address stratification is challengeable on its face.[3]
- Sample-period representativeness. Block sampling and short-window statistical samples are vulnerable to this challenge. A promotional period (Black Friday week, an unusual product launch), a known and remediated configuration error, or an integration cutover that produced a temporary calculation gap concentrates errors in one window. The auditor's projection assumes the sampled period's error rate is the population's error rate. Demonstrating non-representativeness shifts the projection.
The documentation that makes any of these challenges possible is the transaction-level record across the full audit period: every order, every calculated tax amount, every applied rate, every customer location, every exemption code. Without it, the taxpayer is arguing the audit on the auditor's data. TaxCloud calculates and logs every transaction across 13,000+ jurisdictions through one API, and exposes the resulting record through the reporting API. That record is the population the auditor's sample is supposed to represent, which is what makes representativeness, stratification, and population-definition arguments factual rather than rhetorical.
The protest workflow itself is a CPA and attorney exercise. The role of the data is not to win the protest on its own. The role is to give counsel a defensible factual base for the technical arguments, which are argued at the state administrative level long before any court is involved.
What this means at the opening conference for a $20M to $80M brand
A $20M to $80M Shopify or Shopify Plus brand sitting across the table from a state auditor is in a different position than a high-volume seller would be five years earlier. Transaction volume is too high for full review. Sampling is mechanical. The methodology decision at the opening conference, and the data the brand can produce to defend it, drives the eventual assessment more than any single transaction in the audit ever will.
Three operating-model consequences follow.
First, the sampling-method election is a leadership decision, not an auditor-led default. A controller or CFO who arrives without a position cedes that decision. The right position is statistical or stratified sampling across the full audit period, documented in the sampling plan agreement before fieldwork begins.
Second, the data the brand brings to the conference is the lever. Without a complete transaction-level record across the full audit window, the auditor's population definition stands. With it, every component of the projection (population definition, stratification, sample period, point estimate, confidence interval) is open to challenge at protest.
Third, sampling-defense competence is a CPA and attorney workflow. The data and tooling sit with finance and the tax provider. The protest itself is argued by counsel. The brand's job is to make the data ready before the audit begins, not after the assessment lands.
The reader here is past wondering whether sampling matters. The question is what the operating model looks like when an auditor proposes block sampling on a tough month at a brand running 200,000-plus annual transactions across multiple channels. TaxCloud is built for that: a complete transaction-level record across every channel, the reporting API counsel needs to argue any of the protest vectors above, and consolidated SST filing across the 23 full member states plus Tennessee as associate.