R package for clinical, epidemiological, and public-health analysis
gtstats
Understand your data, build a publication-ready Table 1, and answer focused statistical questions without hand-formatting every output.
Publication-ready descriptive and inferential statistics
gtstats helps students, researchers, and public-health analysts move from a new dataset to clear descriptive results, focused comparisons, and a publication-ready Table 1. It keeps the R syntax approachable while retaining the underlying results, denominators, assumptions, and automatic decisions for review.
| Stage | Start with | What you get |
|---|---|---|
| Understand | describe_data() |
Variable type, completeness, levels/range, and concise data overview |
| Assess | assess_distribution() |
Distribution diagnostics and descriptive guidance for continuous variables |
| Assess | assess_variance() |
Group SDs, variances, spread ratios, and median-centred Levene support |
| Describe | summary_table() |
Publication-ready participant-characteristics table |
| Present calculated results | as_stats_table() |
Value-preserving publication table for an already summarised data frame |
| Compare | compare_groups() |
One focused group comparison with a clearly identified test and automatic-selection rule |
| Audit |
assumptions_stats(), diagnostics_stats(), denominators_stats()
|
Transparent decisions and analysis population |
| Export |
customise_table(), to_gt(), save_output()
|
A modifiable, report-ready table |
| Guided | gtstats_app() |
A point-and-click companion that generates the matching R code |
Understand
describe_data()
assess_distribution()
Describe
summary_table()
add_*()
Compare
compare_groups()
correlation()
Report
customise_table()
save_output()
Why it exists
Statistics packages often make a user choose between an oversimplified default and a highly technical workflow. gtstats is designed to bridge that gap:
- publication-ready tables are the default, not an extra formatting task;
- automatic test selection is visible and can always be overridden;
- Welch methods are the conservative default for suitable independent continuous comparisons;
var_equal = TRUEis available when equal variances are prespecified and justified, not inferred from a variance test; - distribution, variance, assumptions, and denominators are inspectable;
- the table builder offers simple defaults with optional layers of control.
The automatic algorithm is documented in full: marked skewness—not a Shapiro-Wilk p-value alone—triggers rank-based continuous tests; sparse categorical tables use Fisher’s exact test when an expected count is below 1 or more than 20% are below 5; and paired/repeated designs use their dedicated methods. See the inferential tests and assumptions guide.
gtstats uses established R statistical engines such as stats, dplyr, ggplot2, gt, and flextable; it provides a coherent teaching and reporting workflow around them.
Built-in teaching data
gtstats includes five labelled datasets, so examples work without installing other packages:
data(
"birthwt", "trial_data", "paired_data", "outbreak_data",
"surveillance_data", package = "gtstats"
)-
birthwt: real low-birth-weight clinical data for descriptive tables, independent comparisons, and epidemiology. -
trial_data: a three-arm synthetic trial for Welch ANOVA, Kruskal-Wallis, chi-square, Fisher exact, correlations, and rates. -
paired_data: long-format paired follow-up data for paired t-tests, Wilcoxon signed-rank, and McNemar tests. -
outbreak_data: the real CDC Oswego foodborne-outbreak line list for attack rates and 2x2 epidemiology. -
surveillance_data: an archived CDC weekly hospital-admission extract for aggregate counts, population denominators, and rates per 100,000.
Prefer a guided interface?
The package remains code-first, but a Shiny companion is available when you prefer to learn by clicking through the workflow. It includes all five teaching datasets, data from the current R environment, CSV/Excel upload, a working data dictionary, distribution and variance checks, a layer-by-layer summary table, outbreak and surveillance tables, focused group comparisons, correlations and crosstabs. Tables can be downloaded as Word, HTML, PDF, or RTF, and every analysis has copyable and downloadable R code. Excel import uses the optional rio package.
install.packages("shiny") # once, if needed
install.packages("rio") # once, only for Excel files
gtstats_app() # opens in the RStudio Viewer when availableThe interface is intentionally a guide rather than a black box: use the code shown beside each result in an R script or Quarto document to keep the final analysis reproducible.
The app keeps three table routes separate. Choose Summary table for raw participant-level data, Epi table when events and denominators still need calculation, and Current data are final calculated results in Customise table only when every supplied row and value should be preserved exactly through as_stats_table().
Five-minute workflow
data("birthwt", package = "gtstats")
overview <- describe_data(birthwt)
distribution <- assess_distribution(birthwt, vars = c(age, lwt), by = low)
variance <- assess_variance(birthwt, vars = c(age, lwt), by = low)
# Use `test = "none"` for descriptive spreads only, or `test = "bartlett"`
# when the stronger normal-distribution assumption is justified.
table_one <- summary_table(
birthwt,
by = low,
include = c(age, lwt, race, smoke),
overall = "last",
layout = "separate"
) |>
add_ci(vars = c(age, race)) |>
add_p()
# Optional categorical-only presentation with distinct n and % columns
categorical_table <- summary_table(
birthwt,
by = low,
include = c(race, smoke),
categorical_layout = "separate"
)
# Optional compact binary presentation: one clinically relevant event per row
compact_binary_table <- summary_table(
birthwt,
by = low,
include = c(smoke, ht, race),
show_dichotomous = "single_row",
value = c(smoke = "Yes", ht = "Yes")
)
# Missing values are excluded from categorical percentage denominators by
# default. Use this explicit option only when Missing is itself a reported
# category and should contribute to 100%.
missing_as_category <- summary_table(
birthwt,
include = c(race, smoke),
missing = "as_category"
)
comparison <- compare_groups(birthwt, variable = lwt, group = low)
comparison$method$selection_rule
diagnostics_stats(comparison)Each object prints as a publication-ready flextable, ready for Word and PowerPoint. Use customise_table() to finish its appearance. Use to_gt() only when an HTML-focused gt table is specifically required.
table_one |>
customise_table(
theme = "journal",
title = "Table 1. Maternal characteristics",
spanning_header = list(
"Birth-weight groups" = c(
"low = Normal birth weight",
"low = Low birth weight"
)
),
borders = "horizontal",
pvalue_style = "threshold"
) |>
save_output("table-1.docx")
# Explicit HTML/gt route
to_gt(table_one, title = "Table 1. Maternal characteristics")Browse by task
| Task | Start here |
|---|---|
| First workflow | Getting started |
| Complete worked example | Birth-weight case study |
| Build Table 1 | Descriptive tables |
| Select and interpret comparisons | Inferential tests and assumptions |
| Missingness and denominators | Missing data and denominators |
| Confidence intervals and 2x2 tables | Epidemiology helpers |
| Every argument and default | Function options |
| All functions | Reference index |
Publication tables and console output
Core statistical functions produce a publication-ready table by default. In a terminal, automated test, or plain-text workflow, request a real tibble:
compare_groups(birthwt_data, variable = age, group = low)
compare_groups(
birthwt_data,
variable = age,
group = low,
format = "tibble"
)The same format = "table" or format = "tibble" choice is available for data description, distribution and variance assessment, effect sizes, correlations, proportions, rates, and cross-tabulations. A summary-table builder stays composable in either format, so add_p() and other layers still work before the completed tibble is printed.
Function map
| Workflow | Functions |
|---|---|
| Understand |
describe_data(), assess_distribution(), assess_variance()
|
| Build Table 1 |
summary_table(), add_ci(), add_p(), add_proportion(), add_rate(), add_total(), add_row()
|
| Compare |
compare_groups(), effect_size(), correlation()
|
| Outbreak and surveillance tables | epi_table() |
| Format an already calculated data frame | as_stats_table() |
| Focused epidemiology calculations |
proportion_stats(), rate_stats(), crosstabs()
|
| Inspect decisions |
assumptions_stats(), diagnostics_stats(), denominators_stats()
|
| Visualise |
plot_compare(), plot_correlation()
|
| Guided interface | gtstats_app() |
| Polish and export |
customise_table(), to_flextable(), to_gt(), save_output()
|
correlation() handles both a prespecified pair and an exploratory matrix. The matrix retains pair-specific denominators and p-values in $summary, while plot_correlation() provides a labelled coefficient heatmap.
correlation(mtcars, vars = c(mpg, disp, hp, wt)) |>
plot_correlation()Citation
If gtstats contributes to your analysis, tables, teaching, or publication, please cite:
Polani R, Kaviprawin M, Sakthivel M, Eliyas SK, Krishnamoorthy Y (2026). gtstats: Beginner-Friendly Statistics and Publication-Ready Tables. R package version 1.0.0. https://gtstats.thinkdenominator.com/.
R returns the authoritative citation and BibTeX entry installed with the package:
The pkgdown Citing gtstats article provides a copy-ready BibTeX record and reproducibility guidance. The reserved Zenodo DOI will be added here after its record is published and the DOI resolves publicly.