flexsynth generates new records intended to reproduce selected statistical
structure from a real dataset. The default Track A engine is a
utility-oriented sequential synthesiser. It may reproduce source values or even
records, carries no formal privacy guarantee, and must not be treated as
anonymisation. The package’s empirical utility and disclosure-risk diagnostics
inform a release review; they do not prove that a release is safe. The opt-in
differentially private Track B and the conditions required for its formal
guarantee are described in vignette("differential-privacy").
We start with a plain single-table dataset — one row per patient.
n <- 300
real <- data.frame(
id = seq_len(n),
age = round(rnorm(n, 62, 11)),
sex = sample(c("F", "M"), n, replace = TRUE, prob = c(0.45, 0.55)),
smoker = sample(c(FALSE, TRUE), n, replace = TRUE, prob = c(0.7, 0.3))
)
real$sbp <- round(0.6 * real$age + ifelse(real$smoker, 8, 0) + rnorm(n, 90, 10))
head(real)
| id | age | sex | smoker | sbp |
|---|---|---|---|---|
| 1 | 55 | F | TRUE | 128 |
| 2 | 64 | F | TRUE | 151 |
| 3 | 53 | M | FALSE | 127 |
| 4 | 80 | F | FALSE | 143 |
| 5 | 66 | F | FALSE | 128 |
| 6 | 53 | F | FALSE | 110 |
synth() takes the data and a structure formula naming the unit identifier.
Here every row is its own unit, so the structure is just ~ id.
res <- synth(real, structure = ~ id, seed = 1)
res
#> <synth_result>
#> track : A (high-utility; NOT differentially private)
#> datasets (m) : 1
#> rows each : 300 (input: 300 )
#> synthesised : age, sex, smoker, sbp
#> unit-level : age, sex, smoker, sbp (once per unit)
#> method : cart
#>
#> Get the data with as.data.frame(x)
syn <- as.data.frame(res)
head(syn)
| id | age | sex | smoker | sbp |
|---|---|---|---|---|
| 1 | 65.578 | M | FALSE | 143.902 |
| 2 | 59.392 | M | FALSE | 130.690 |
| 3 | 53.896 | F | FALSE | 113.032 |
| 4 | 53.052 | F | FALSE | 114.245 |
| 5 | 50.355 | F | FALSE | 116.001 |
| 6 | 72.592 | M | FALSE | 139.909 |
The unit identifier is regenerated, and every other column is synthesised in
turn from a model fitted on the real data (the default method = "cart" draws
from the matching leaf of a regression / classification tree).
diagnose()diagnose() compares real and synthetic data on four descriptive views:
per-variable marginals (a Kolmogorov-Smirnov statistic for numeric variables and
total-variation distance for categorical ones), numeric correlations,
categorical associations (Cramer’s V), and a propensity (pMSE) score. For the
default main-effects logistic model, a ratio near 1 is its null benchmark;
larger values indicate greater row-level distinguishability. The score is fitted
and evaluated in-sample, and repeated rows in longitudinal data are not
independent evidence, so do not interpret the ratio as a hypothesis test.
analysis_vars <- c("age", "sex", "smoker", "sbp")
d <- diagnose(real, res, vars = analysis_vars)
d
#> <flexsynth_diagnostics>
#> rows : real 300 synthetic 300
#> variables : 4
#>
#> Univariate fit (smaller = closer):
#> variable type metric distance
#> age numeric ks 0.0600
#> sex categorical tvd 0.0500
#> smoker logical tvd 0.0433
#> sbp numeric ks 0.0433
#> mean distance: 0.0492 worst: age (0.0600)
#>
#> Correlation structure (2 numeric vars):
#> Frobenius diff: 0.0618 mean |diff|: 0.0437 max |diff|: 0.0437
#>
#> Categorical association (Cramer's V, 2 vars):
#> mean |diff|: 0.0000 max |diff|: 0.0000
#>
#> Propensity utility (pMSE, logistic; descriptive, in-sample):
#> pMSE: 0.00123 expected: 0.00083 ratio: 1.48 (1 = indistinguishable)
The plot() method overlays each marginal.
plot(d)
disclosure_risk()Synthetic data is not anonymisation. disclosure_risk() provides four
empirical checks: replicated uniques; distance to the closest real record (DCR);
membership inference when a genuine, synthesis-excluded holdout is supplied;
and Target Correct Attribution Probability (TCAP) when a categorical sensitive
target is named. Pass genuinely identifying columns as quasi, excluding
surrogate keys such as the regenerated id. These diagnostics flag potential
identity, membership, and attribute-disclosure problems; no threshold certifies
a release as safe.
disclosure_risk(real, res, quasi = c("age", "sex", "smoker", "sbp"), seed = 1)
#> <flexsynth_disclosure>
#> rows : real 300 synthetic 300
#> quasi-identifiers: age, sex, smoker, sbp
#>
#> Replicated uniques (identity risk):
#> real sample-uniques : 282
#> reproduced in syn : 0 (0.00% of uniques, 0.00% of real rows)
#> syn rows copying a real row: 0 (0.00% of syn)
#>
#> Distance to closest record (Gower, 0 = exact copy):
#> syn->real : median 0.0075 5th pct 0.0018 exact copies 0.00%
#> real->real : median 0.0088 (baseline)
#> median syn distance is smaller than the real-neighbour baseline
#> descriptive only: inspect lower-tail distances and exact copies; this is not a safety guarantee
#>
#> Membership inference: not run (supply `holdout` of non-training records).
#>
#> Attribute disclosure: not run (supply `target` = a sensitive column).
Methods live in an extensible registry. Built-ins include sample, cart,
forest (a bagged CART ensemble), ctree (via partykit), and the parametric
numeric methods norm and normrank.
list_methods()
#> [1] "cart" "ctree" "forest" "norm" "normrank" "sample"
Set one method for all variables, or per variable via synth_control():
ctrl <- synth_control(method = c(sbp = "norm", age = "cart"))
res2 <- synth(real, ~ id, tuning = ctrl, seed = 1)
You can register your own with register_method() — supply a fit() and a
draw().
rule() declares a logical constraint the synthetic data must satisfy;
synth() enforces it by rejection sampling at the unit grain.
res3 <- synth(real, ~ id,
constraints = rule(sbp >= 80 & sbp <= 220),
seed = 1)
range(as.data.frame(res3)$sbp)
#> [1] 92 171
synth_control() also exposes smoothing (kernel-smooth numeric draws) and
predictor_matrix (restrict which variables may predict each target). See
?synth_control.
Track A output must never be described as differentially private. Use the diagnostics above to decide whether the utility / risk balance is acceptable for your release.