For everyone working with real-world data — pharma, biotech, MedTech, CROs, academia, and the hospitals — community, multi-specialty and research — that hold the data. From fit-for-purpose study design through causal theory to hands-on R labs on propensity scores and synthetic data.
Strategy morning · Paired hands-on R afternoonReal-world evidence now sits at the centre of regulatory, clinical and commercial decisions — and the hospitals holding EHR/EMR data — community, multi-specialty and research alike — are increasingly central to it. This workshop takes participants end-to-end: what RWE is and where the data comes from, how to judge whether a source is fit for a given question, the causal theory needed to analyse observational data credibly, and two hands-on R labs putting that theory to work.
Please bring a laptop with R installed — that is the minimum requirement. We will email full setup instructions (packages, datasets and scripts) beforehand, and confirm your setup in the R Environment Setup session. No coding needed in the morning; in the afternoon, non-technical attendees are paired with a practitioner and follow the analysis live.
RWD vs RWE and how they differ from RCT evidence; where the data comes from — hospital EHR/EMR systems, claims, registries and wearables — and why data owners are becoming central to the evidence chain; the regulatory landscape (FDA, EMA, CDSCO); what makes real-world data credible.
Is the data fit for purpose? The Structured Process to Identify Fit-For-Purpose Data: defining and ranking data-quality criteria, assessing a candidate source — whether an EHR/EMR, a claims database or a registry — against a research question, and running a feasibility assessment before committing to a study.
Confounding by indication; estimands; directed acyclic graphs (DAGs); the logic of propensity scores, IPTW, positivity & balance — why observational data needs more than a simple comparison, and the theory underpinning the afternoon labs.
Sharing data without sharing patients: what synthetic data is (fully vs partially synthetic); sequential synthesis with CART; the utility–privacy trade-off; the four-level validation framework (fidelity, structure, predictive, causal); ethics & governance.
We confirm everyone is analysis-ready. You will have had setup instructions by email; here we check R and RStudio, install the required packages (synthpop, MatchIt, survey), load the workshop dataset and scripts, and run a test. Non-coders pair up with a practitioner here.
Hands-on in R: estimate propensity scores, apply inverse probability weighting (IPTW), assess covariate balance (SMD) and fit weighted outcome models on a real EHR dataset. Paired seating — non-coders follow the analysis live. Led by Ashwini Mathur.
Hands-on in R: generate synthetic data with synthpop (sequential CART); validate utility across fidelity, structure, predictive and causal levels; privacy, ethics and responsible use. Led by Ashwini Mathur.
Onesto Consulting delivers practitioner-focused training in clinical data science and statistical methods. This workshop blends concise theory with live, code-along R sessions on a real EHR dataset, so participants leave able to apply causal methods and synthetic-data techniques directly to their own real-world data.
Places are limited to keep the lab sessions interactive. Email any of us to save your seat, request an invoice, or ask about paired places for non-coders.
Online payment gateway coming soon