This vignette covers high-dimensional inference in
fetwfe: what changes when the design has as many or more
columns than observations (\(p \geq
NT\)), and how debiasedATT() restores a uniformly
valid confidence interval for the overall ATT.
It is a companion to vignette("inference_vignette") (how
the package’s standard errors are computed in the low-dimensional
regime) and vignette("simultaneous_cis_vignette")
(family-wise simultaneous bands).
debiasedATT()The standard errors and simultaneous bands documented in the
companion vignettes assume a low-dimensional design
(\(p < NT\)), where the analytic
Gram inverse exists and the fit’s standard errors are the asymptotically
valid ones of vignette("inference_vignette"). When the
design is high-dimensional — the number of design
columns \(p\) (cohort and time fixed
effects, covariates, and all their treatment interactions) meets or
exceeds the number of observations \(NT\) — two things change:
The fused fit is a regularized (bridge) estimate, and
its pointwise standard errors (and the simultaneous bands of
vignette("simultaneous_cis_vignette")) are
post-selection: valid conditional on the selected
support, not uniformly valid. They can under-cover, because the
bridge’s regularization biases the point estimate (it shrinks and prunes
effects) and the post-selection standard error does not account for the
selection.
debiasedATT() returns a uniformly
valid confidence interval for the overall ATT by
desparsifying the fit (the nodewise relaxed-inverse
construction of the high-dimensional FETWFE theory in Faletto 2025),
which removes the regularization bias. This is “one estimator, two
regimes”: the call and the returned object are identical to the \(p < NT\) case; only the internal
nuisance changes.
The \(p \geq NT\) regime arises with
many covariates, short panels, or many cohorts.
simulateData() can generate such panels.
# d = 12 covariates, N = 40 units, T = 5 periods, G = 3 cohorts.
coefs_hd <- genCoefs(G = 3, T = 5, d = 12, density = 0.2, eff_size = 2, seed = 2)
sim_hd <- simulateData(
coefs_hd, N = 40, sig_eps_sq = 1, sig_eps_c_sq = 0.5, seed = 2
)
c(p = sim_hd$p, NT = sim_hd$N * sim_hd$T) # p >= NT: high-dimensional
#> p NT
#> 220 200fit_hd <- fetwfeWithSimulatedData(sim_hd, q = 0.5)
# The fused (post-selection, pointwise) overall ATT:
round(c(att = fit_hd$att_hat, se = fit_hd$att_se), 3)
#> att se
#> -2.416 0.186That fused estimate is post-selection. For uniformly valid inference,
pass the fit to debiasedATT(), which implements the
high-dimensional debiased construction of Theorem
debiased.highdim.thm (Faletto 2025). Its coverage
is robust across the penalty values studied — a cross-validated selector
below the theory scale and a fixed constant above it bracket the default
1.0 — so the default is fine; below we let
lambda_c = "cv" select the constant per fit as an
alternative.
db <- debiasedATT(fit_hd, lambda_c = "cv")
db
#> Debiased overall ATT (high-dimensional regime)
#> Estimate: -2.3684
#> Std. Error: 0.2507
#> 95% CI: [-2.8597, -1.8771]
#> High-dimensional (p >= NT) desparsified interval
#> [overall-ATT coverage validated near-nominally in simulation
#> (Theorem debiased.highdim.thm, Faletto 2025); diagnostics below]
#> nodewise penalty: lambda_c = 0.5 (CV-selected); lambda_node in [0.1836, 0.1836]
#> nodewise directions: 1/1 converged, 1/1 KKT-feasible (worst feasibility/lambda_node = 1)The print surfaces the high-dimensional diagnostics worth inspecting:
the resolved nodewise penalty (lambda_c /
lambda_node) and, per debiasing direction, whether it
converged and met its KKT feasibility
certificate. A direction that fails to converge or stay feasible flags a
potentially unreliable standard error — raise
riesz_max_iter and/or lambda_c and
re-inspect.
In any single draw the fused and debiased point estimates can be
close, so the value of debiasedATT() is not visible from
one fit — it is a repeated-sampling guarantee. A
coverage study at a \(p > NT\)
anchor (reported in Faletto 2025) found:
| Overall-ATT interval | Coverage (nominal 95%) |
|---|---|
Debiased (debiasedATT() with
lambda_c = "cv") |
≈ 94% |
| Fused plug-in (pointwise) | ≈ 77% |
The fused interval under-covers because of the post-selection
non-uniformity; the debiased interval is near-nominal — the direct
simulation confirmation of Theorem
debiased.highdim.thm, robust across the penalty
values studied (a CV selector below the theory scale and a fixed
constant above it, bracketing the default 1.0). The
per-direction converged / feasibility
diagnostics above remain worth a glance to confirm the nodewise
directions are well-behaved on your own fit.
method = "bootstrap" offers an alternative,
asymptotically-valid reference distribution for the overall-ATT
interval: a studentized score / influence-function wild cluster
bootstrap that re-signs the per-unit influence summands (no
refit per replicate), leaving the point estimate and the reported
se unchanged and changing only the critical value. It is
floored at the Gaussian quantile, so the interval is
never narrower than the analytic Wald interval.
It is not a few-clusters remedy. The analytic
sandwich SE is downward-biased with few clusters, but this bootstrap
does not fix that: as an unrestricted, no-refit score bootstrap
it does not reproduce the tail inflation the restricted
wild-cluster bootstrap-t uses for few-cluster coverage, and under
heterogeneous cluster influence its critical value (before the floor)
actually falls below the Gaussian — so the floor reduces it to
the analytic interval in exactly the small-N regime. Its
only effect is to widen the interval when cluster influence is
near-homogeneous. For genuine few-clusters inference a restricted
bootstrap-t or a CR2-type analytic correction is required (future
work).
# Works on both single-sample and two-sample (`indep_counts`) fits -- e.g. the
# fetwfeWithSimulatedData() fit from the worked example above (a two-sample fit).
debiasedATT(fit, method = "bootstrap", multiplier = "webb", seed = 1)The multiplier ("webb", the default;
"rademacher"; or "mammen") only affects the
near-homogeneous case where the bootstrap widens above the floor — in
the small-N / heterogeneous regime the interval is the
analytic one regardless of the weight. The bootstrap supports both
single-sample fits (real observational panels are single-sample) and
indep_counts two-sample fits (the cohort-weight channel is
perturbed on its own independent multiplier stream over the
reconstructed count-sample units).
simultaneousCIs() also switches to the desparsified
construction in the high-dimensional regime (via
method = "bootstrap"; the analytic Gram-inverse band does
not exist when \(p \geq NT\)).
Pass method = "bootstrap" for the valid
band: the analytic default — and
eventStudy(ci_type = "simultaneous") — instead return the
post-selection band on the selected support, which
under-covers. On a gls = FALSE fit only the bootstrap band
is available.
sc_hd <- simultaneousCIs(
fit_hd,
family = "cohort",
method = "bootstrap",
B = 500,
seed = 1
)
sc_hd$regime
#> [1] "high-dimensional"Both regimes are now simulation-validated: Theorem
debiased.highdim.thm covers the scalar
overall-ATT (debiasedATT(), ≈0.94), and
Theorem debiased.highdim.joint.thm covers
the family-wise band
(simultaneousCIs(method = "bootstrap")) — family-wise
coverage ≈0.92 over the event-study family (\(K = 7\)), tighter than Bonferroni — both at
the \(p > NT\) anchor of Faletto
(2025). Like the scalar, the band is validated at a single
high-dimensional anchor and rests on the sparsity /
restricted-eigenvalue assumptions (assumed, not proved); the cohort and
all-post-treatment families are covered by the same theorem but were not
separately simulated.
The default density placement spreads signal across all
coordinates, which when \(p\) greatly
exceeds the number of treatment-effect parameters rarely lands signal on
the treatment-effect block. For a controlled, sparse, non-degenerate
high-dimensional data-generating process, genCoefs() offers
a targeted-sparsity mode (n_signal_cohorts
or treat_base_levels) that places a heterogeneous
treatment-effect signal directly on the per-cohort base levels — the DGP
the #88 coverage study uses. See ?genCoefs.
Faletto, G. (2025). “Fused Extended Two-Way Fixed Effects for Difference-in-Differences with Staggered Adoptions.” arXiv preprint arXiv:2312.05985.