Every empirical paper says one thing caused another. Four things have to be true before you can believe it. This guide shows what they are, what breaks each one, and where to look in a paper to find the break.
Start here. The Libby boxes (Libby, Bloomfield and Nelson, 2002) lay out every empirical paper as two worlds at once. On top is the conceptual world, where your ideas and your theory sit. Underneath is the empirical world, where the data you can actually get sits. Your paper is only as good as the four arrows that connect them. Each arrow is a promise, and each one breaks in its own way. Click any arrow or box to see what it promises and what breaks it.
Click any arrow or box above. Showing: arrow 4
A paper usually defends one arrow carefully and leans on the others. Archival designs tend to concentrate on arrow 4; experiments tend to concentrate on arrows 2 and 3. When you write a referee report, say which arrow the paper defended and which it took for granted.
These are the names for the four ways a paper can break. The order depends on what you are doing. Building a paper, you work from the idea down: decide what you are measuring, then how to get a clean comparison, then how to estimate it, then who the answer applies to. Reading someone else, you work the other way, from the data outward. Switch below and watch the sequence flip.
Do the two things really move together? This depends on how you ran the test, how you computed the standard errors, and how many tests you tried before this one.
Failure looks likeA t statistic of 2.1, standard errors that were never clustered, three treated states, and the fourth measure you tried.
They move together, but is that because one caused the other? Or because firms chose their own treatment, something else happened at the same time, the trend was already running, or the outcome moved the cause?
Failure looks likeTreated firms were already diverging from controls two years before the mandate.
Your variables have to stand for the ideas in your title. You can have a clean causal effect and still be wrong about what caused what.
Failure looks likeThe bid ask spread went up, so the paper says information asymmetry went up. The spread moves for other reasons too.
Does the result hold for other firms, other sizes of the same change, other outcomes, other countries, other years? Designs that are clean on step 02 usually pay the price here.
Failure looks likeThe design only compares firms right at a cutoff, and the abstract talks about firms in general.
So which order is correct? Both, for different jobs, and this guide opens on the one you need first. When you build a paper, construct validity comes first. Decide what you are measuring and how, then find a clean comparison, then worry about estimation, then ask who the answer applies to. Start anywhere else and you get a clean coefficient for a question nobody asked. The textbook order, from Shadish, Cook and Campbell, is the reading order: is the pattern real, then did one cause the other, then are we measuring what we say, then who does it apply to. That is how you check a result someone hands you. Use it in referee reports. Use the building order for your own work.
Steps 02 and 04 fight each other. A clean design gives you a solid answer about a small group. A big sample gives you reach, but more things can confound it. Neither is better on its own. The question is whether the paper made the right trade for the question it asked.
Where time enters. Students ask whether temporal validity is a fifth type. It is not. Time enters at three points and the distinction matters. At step 01 it is the horizon: the window you chose decides the coefficient. At step 02 it is timing: anticipation, mis dated treatment, dynamic effects, and trends that were already running. At step 04 it is the regime: whether the relation itself has a half life. A paper can be flawless about the period it studies and still be describing a world that no longer exists.
What sits inside each of these. Construct validity contains the whole measurement family: content, face, convergent, discriminant, criterion and nomological validity, resting on reliability. That family has its own tab. External validity contains population validity, ecological validity, and the temporal question above. Statistical conclusion validity contains the identification assumptions you cannot test, including the exclusion restriction. Nothing here covers the qualitative tradition, where Maxwell splits validity into descriptive, interpretive, theoretical and evaluative. If a student is doing field work, that framework replaces this one.
Where the frameworks come from. The four way split and the named threats follow Campbell and Stanley, then Shadish, Cook and Campbell. The link diagram in the first tab follows Libby, Bloomfield and Nelson. The econometric vocabulary follows Angrist and Pischke. They describe the same problem in three dialects.
Arrows 2 and 3 are about construct validity — whether your variables measure the ideas you named. Showing that a variable measures what you say is not one test. It is six, and they only work if the measure is stable to begin with. This applies to an index you build by hand, a score from text, or anything a language model produces.
Reliability is a floor, not a check. A measure that gives a different answer on a rerun cannot be valid for anything. For hand collected data that means inter coder agreement. For a measure produced by a language model it means the same thing under a different seed, a different prompt wording, and a different model version. Report it before you report a correlation.
The check that is easy to overlook. Discriminant validity. Convergent evidence is easy to come by, because in accounting almost every firm level measure correlates with size, so a proxy that tracks the accepted proxy has not shown much. The harder question is what your measure should be unrelated to, and whether it is. Name that variable before you build the measure.
Forty seven things that can go wrong, each with an accounting example and what to do about it. Filter by type, or press Time to see the ten that are about timing. Time is not a fifth type. It shows up inside three of the four.
No design is best. Each one solves some problems and leaves others open. Click a row to see the weakness a referee will go to first.
| Design | History | Maturation | Selection | Regression to mean | Attrition | Omitted variable | Reverse causality | External reach |
|---|
Each row carries one weakness that referees go to first. Click a row to read it.
Work through these before you write your report. Order matters. If a paper fails the first block, it does not need your comments on the last one.
Writing the report. Lead with the link that carries the paper. State the threat, say why it is plausible in this setting, then name the test that would settle it. A threat you cannot turn into a test is a comment, not an objection.