Learn to describe a real biomedical dataset using two open-source tools: R for a code-based workflow and jamovi for a graphical workflow. You will inspect the variables, summarize the outcome, compare groups descriptively, and write a short results paragraph.
Level: Beginner. Data: 60 observations. Software: R, optionally with RStudio, or jamovi. No paid statistics package is needed for this exercise.
1. Get the free dataset
We use ToothGrowth, a dataset distributed with R. The official documentation describes measurements of odontoblast length, a tooth-growth response, in 60 guinea pigs receiving vitamin C through orange juice (OJ) or ascorbic acid (VC), at 0.5, 1, or 2 mg/day. There are six supplement-by-dose groups.
Download the ToothGrowth CSV from Rdatasets. If the link displays text in your browser, save it as ToothGrowth.csv. The public CSV mirror includes an extra row identifier; the R code below removes it and the jamovi steps ignore it.
Read the official R dataset documentation for the study description and original references, including Crampton’s 1947 paper and Bliss’s 1952 source. Dataset and documentation checked: 10 October 2026.
Context: This is an animal experiment used here to teach analysis. It does not establish a recommendation about vitamin C intake or treatment in humans. The documentation does not specify a measurement unit for len, so we report the recorded scale without inventing one.
2. Understand the variables
| Variable | Meaning | How to use it |
|---|---|---|
len |
Recorded tooth-growth response length | Continuous outcome |
supp |
OJ: orange juice; VC: ascorbic acid | Nominal grouping variable |
dose |
Vitamin C dose in mg/day: 0.5, 1, or 2 | Use as a grouping factor for this exercise |
rownames |
Row identifier added by the CSV mirror | Exclude from analysis |
Our descriptive question is: How does the recorded response vary across supplement and dose groups? This tutorial does not conduct a hypothesis test.
3. Follow the R workflow
Open R or RStudio, create a new script, and run the code below in order. It uses base R functions and needs no additional packages. If you prefer a local file, replace read.csv(url) with read.csv(file.choose()) and select the downloaded CSV.
# Import the same CSV used in the jamovi walkthrough
url <- "https://vincentarelbundock.github.io/Rdatasets/csv/datasets/ToothGrowth.csv"
dat <- read.csv(url)
dat$rownames <- NULL # Remove the CSV row identifier
# Define the grouping variables
dat$supp <- factor(dat$supp, levels = c("OJ", "VC"))
dat$dose_group <- factor(dat$dose, levels = c(0.5, 1, 2))
# Check the imported data
str(dat)
colSums(is.na(dat))
table(dat$supp, dat$dose_group)
# A reusable descriptive summary
summarize <- function(x) {
c(N = sum(!is.na(x)),
Missing = sum(is.na(x)),
Mean = mean(x, na.rm = TRUE),
Median = median(x, na.rm = TRUE),
SD = sd(x, na.rm = TRUE),
Minimum = min(x, na.rm = TRUE),
Maximum = max(x, na.rm = TRUE))
}
# Overall results
round(summarize(dat$len), 2)
# Results for each supplement-by-dose group
groups <- split(dat$len,
interaction(dat$supp, dat$dose_group,
sep = " | ", drop = TRUE))
round(do.call(rbind, lapply(groups, summarize)), 2)
# Explore the overall distribution
hist(dat$len, main = "Distribution of the recorded response",
xlab = "Response length (reported scale)", col = "lightblue")
# Compare the six groups visually
boxplot(len ~ interaction(supp, dose_group, sep = " / "),
data = dat, las = 2,
xlab = "Supplement / dose (mg per day)",
ylab = "Response length (reported scale)",
main = "ToothGrowth: descriptive comparison", col = "lightblue")
# Record the software environment for reproducibility
sessionInfo()
Import checks: Expect 60 rows, no missing values in the three study variables, and 10 observations in every supplement-by-dose cell. R’s sd() calculates the sample standard deviation.
Offline alternative: The data are also available as datasets::ToothGrowth. To use that version, replace the import lines with dat <- datasets::ToothGrowth, then continue with the grouping and summary steps.
4. Follow the jamovi workflow
- Open the CSV. Launch jamovi and open the downloaded file from its file menu.
- Check variable types. Open each column’s variable editor. Set
lento Continuous andsuppto Nominal. For this grouped exercise, setdoseto Nominal, keeping the levels 0.5, 1, and 2. Leaverownamesout of all analyses. - Start the summary. Choose Analyses → Exploration → Descriptives. Move
leninto Variables. - Select statistics. Request N, Missing, Mean, Median, Standard deviation, Minimum, and Maximum. Compare the overall output with the check table below.
- Inspect a plot. Under Plots, select a histogram or box plot to inspect the response distribution.
- Compare groups. Move both
suppanddoseinto Split by. This gives six groups, each with N = 10. Request box plots for a visual comparison. - Save your work. Save a jamovi
.omvproject. Add notes about your question, variable settings, and interpretation. Record your jamovi version.
Menu wording can vary by version. Consult jamovi’s getting-started guide and official Descriptives documentation for the available options.
5. Check your results
The following reference values were calculated from the downloaded CSV. Values are rounded to two decimal places; SD is the sample standard deviation.
| Overall statistic | Expected value |
|---|---|
| N | 60 |
| Missing response values | 0 |
| Mean | 18.81 |
| Median | 19.25 |
| SD | 7.65 |
| Minimum | 4.20 |
| Maximum | 33.90 |
| Supplement | Dose (mg/day) | N | Mean | SD |
|---|---|---|---|---|
| OJ | 0.5 | 10 | 13.23 | 4.46 |
| OJ | 1 | 10 | 22.70 | 3.91 |
| OJ | 2 | 10 | 26.06 | 2.66 |
| VC | 0.5 | 10 | 7.98 | 2.75 |
| VC | 1 | 10 | 16.77 | 2.52 |
| VC | 2 | 10 | 26.14 | 4.80 |
6. Write a careful interpretation
The dataset contained 60 guinea pigs, with 10 in each supplement-by-dose group. The overall recorded response had a mean of 18.81 and a standard deviation of 7.65. Group means increased across the three dose levels within both supplement groups. At 0.5 and 1 mg/day, the OJ group means were higher than the corresponding VC means; at 2 mg/day, the means were similar. These summaries describe the observed data and do not quantify uncertainty in the group differences or establish conclusions for humans.
The overall mean mixes all six groups. It is a useful data check, but the grouped summaries better address our question. A difference between means alone is not a significance test. SD describes the spread of individual observations; it is not the standard error of a mean.
7. Avoid these beginner mistakes
- Analyzing the row identifier. It is not a biological measurement.
- Treating dose as the outcome. In this exercise,
lenis the outcome and dose identifies groups. - Ignoring one grouping variable. Splitting only by supplement combines different dose levels.
- Deleting observations because a plot flags them. First investigate the value and document any justified decision.
- Overstating the findings. Keep the animal-study context and the descriptive scope visible.
Continue learning
For an introduction to the software, visit RStudio Education’s beginner resources. Continue with our Biostatistics guide, Academic Writing guide, or Health Research & Graduate Funding Hub.