library(tidyverse)2 Functional Testing
Functional testing (also called performance-based testing) is an objective measure of a participants physical functionality and complements self-reported pain and disability. Before surgery, A2CPS collects a short battery of standardized functional tasks, it records the pain a participant reports right before and right after each task. The change in pain across a task gives us a measure of movement-evoked pain (MEP), i.e., pain that is provoked by movement rather than pain at rest. MEP is increasingly recognized as its own meaningful construct.
The battery is tailored to each surgical cohort:
- Knee arthroplasty (TKA) cohort: lower-limb function is assessed with two tasks.
- The 10-meter walk test (10MWT), a standard measure of gait speed. Participants rate their pain before and after walking, and both the walk time and the pain ratings are recorded.
- The 5-times sit-to-stand test (5TSTS), which measures lower-limb strength and functional mobility by timing five repetitions of rising from a chair. Pain is again rated before and after the task.
- Thoracic (thoracic surgery) cohort: since the surgical site is the chest wall, function here is assessed with a deep breathing and coughing task, again with pain rated before and after the maneuver.
The pain ratings are double-entered (each rating is recorded twice) to guard against errors. The reformatted data also comprises derived MEP scores.
2.1 Starting Project
2.1.1 Locate Data
Where are the relevant files?
$ /corral-secure/projects/A2CPS/products/consortium-data/pre-surgery-release-2-1-0/functional-testing/reformattedThe reformatted data are split by cohort: reformatted_tka_func.csv (knee, with the walk and sit-to-stand measures) and reformatted_thor_func.csv (thoracic, with the deep breathing and coughing measures). Each one comes with an updated data dictionary, updated_func_dict.csv, which lists the original REDCap field names alongside the fields added during cleaning.
2.1.2 Extract Data
Here we use the tidyverse, and we focus on the TKA cohort, which has the two lower-limb tasks.
func <- read_csv("data/pre-surgery/functional-testing/reformatted/reformatted_tka_func.csv")The two derived movement-evoked pain scores are mep_walk (10-meter walk) and mep_5tsts (5-times sit-to-stand). Each one is the final pain rating minus the initial pain rating for that task, so higher values mean more pain was evoked by the movement.
func |>
select(record_id, walk10initialpainscl, walk10finalpainscl, mep_walk,
tstsprepainscl, tstspostpainscl, mep_5tsts) |>
head()| record_id | walk10initialpainscl | walk10finalpainscl | mep_walk | tstsprepainscl | tstspostpainscl | mep_5tsts |
|---|---|---|---|---|---|---|
| 10001 | 4 | 4 | 0 | 5.0 | 6 | 1.0 |
| 10003 | 0 | 2 | 2 | 0.0 | 3 | 3.0 |
| 10004 | 1 | 1 | 0 | 2.0 | 1 | -1.0 |
| 10005 | 5 | 5 | 0 | 5.0 | 6 | 1.0 |
| 10006 | 1 | 1 | 0 | 1.5 | 2 | 0.5 |
| 10007 | NA | NA | NA | NA | NA | NA |
Next, check the data types and make sure the numeric variables are stored as numbers and the categorical variables are not accidentally stored as numeric. glimpse() gives us a quick column-by-column view.
glimpse(func)Rows: 1,042
Columns: 40
$ record_id <dbl> 10001, 10003, 10004, 10005, 10006, 10007, …
$ guid <chr> "90F8FC45-5D53-0DE4-6853-284607A8C4E6", "3…
$ redcap_data_access_group <chr> "rush_university_me", "rush_university_me"…
$ redcap_event_name <chr> "baseline_visit_arm_1", "baseline_visit_ar…
$ redcap_repeat_instrument <chr> "functional_testing", "functional_testing"…
$ redcap_repeat_instance <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ walk10initialpainscl <dbl> 4.0, 0.0, 1.0, 5.0, 1.0, NA, 2.0, 1.5, 1.0…
$ walk10initialpainscl1 <dbl> 4.0, 0.0, 1.0, 5.0, 1.0, NA, 2.0, 1.5, 1.0…
$ walk10finalpainscl <dbl> 4.0, 2.0, 1.0, 5.0, 1.0, NA, 4.0, 3.0, 1.0…
$ walk10finalpainscl1 <dbl> 4.0, 2.0, 1.0, 5.0, 1.0, NA, 4.0, 3.0, 1.0…
$ walk10time <dbl> 8.1, 8.3, 5.1, 4.6, 7.6, NA, 7.4, 5.7, 5.1…
$ walk10time1 <dbl> 8.1, 8.3, 5.1, 4.6, 7.6, NA, 7.4, 5.7, 5.1…
$ walk10completeyn <dbl> 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, …
$ walk10incompletereason <dbl> NA, NA, NA, NA, NA, 1, NA, NA, NA, NA, NA,…
$ walk10assistyn <dbl> 0, 0, 0, 0, 1, NA, 0, 0, 0, 0, 0, 0, 0, 0,…
$ walk10assist_cane___1 <dbl> 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ walk10assist_crutch___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ walk10assist_walkder___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ walk10assist_perssuppt___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ walk10assist_other___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ walk10assist_othertxt <chr> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…
$ walk10comments <chr> NA, NA, NA, NA, NA, "re-consented to study…
$ tstsbpscreen <dbl> 1, 1, 1, 1, 1, NA, 1, 1, 1, 1, 1, 1, 1, 1,…
$ tstsprepainscl <dbl> 5.0, 0.0, 2.0, 5.0, 1.5, NA, 0.0, 0.5, 0.0…
$ tstsprepainscl1 <dbl> 5.0, 0.0, 2.0, 5.0, 1.5, NA, 0.0, 0.5, 0.0…
$ tstspostpainscl <dbl> 6.0, 3.0, 1.0, 6.0, 2.0, NA, 3.5, 3.5, 2.0…
$ tstspostpainscl1 <dbl> 6.0, 3.0, 1.0, 6.0, 2.0, NA, 3.5, 3.5, 2.0…
$ tststime <dbl> 22.6, 15.6, 10.0, 15.7, 25.5, NA, 12.3, 12…
$ tststime1 <dbl> 22.6, 15.6, 10.0, 15.7, 25.5, NA, 12.3, 12…
$ tstscompleteyn <dbl> 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 0, …
$ tstsnonreasonyn <dbl> NA, NA, NA, NA, NA, 2, NA, NA, NA, NA, NA,…
$ tstsnumbrepsyn <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…
$ tstsassistyn <dbl> 0, 0, 0, 0, 0, NA, 0, 0, 0, 0, 0, 0, 0, 1,…
$ tstsassist_1___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, …
$ tstsassist_2___1 <dbl> 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, …
$ tstsaddnotes <chr> NA, NA, "She said that the cartilage in he…
$ functional_testing_complete <dbl> 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, …
$ cohort <chr> "TKA", "TKA", "TKA", "TKA", "TKA", "TKA", …
$ mep_walk <dbl> 0.0, 2.0, 0.0, 0.0, 0.0, NA, 2.0, 1.5, 0.0…
$ mep_5tsts <dbl> 1.0, 3.0, -1.0, 1.0, 0.5, NA, 3.5, 3.0, 2.…
2.1.3 Data Quality
The reformatted files have already been curated, following are some of the important steps we used:
- Test records removed. Placeholder records used during site setup (e.g.,
record_idvalues like10000or15000) are dropped. - Completed tasks only, latest attempt. Only rows marked complete (
functional_testing_complete == 2) are kept, and where a task was repeated, we kept the most recent repeat instance, so we end up with one row per participant and visit. - Movement-evoked pain needs both endpoints.
mep_walkandmep_5tstsare only defined when both the initial and the final pain rating are present, and otherwise they are missing. So expect a modest number of missing MEP values even among completed tasks. - Error report. The cleaning workflow produces an error report that flags discrepancies between the double-entered pain ratings and other internal inconsistencies (for example, a task marked complete but missing a required pain rating). ### Cross-Modality Links
Every file includes record_id and guid, which are unique participant identifiers used to link their records across other A2CPS modalities (imaging, QST, psychosocial, biospecimen, etc.). To combine functional data with another modality, we join on the shared identifiers:
inner_join(func, other_modality, by = intersect(names(func), names(other_modality)))2.2 Exploratory data analysis
Both tasks evoke pain through movement. If walking and standing provoke pain through a common mechanism, then a participant who is sensitive to one should also be sensitive to the other, and the two MEP scores should track together.
We first compare the two tasks. On average, sit-to-stand evokes noticeably more pain than the 10-meter walk (mean mep_5tsts ≈ 1.85 vs. mep_walk ≈ 0.67), which makes sense since rising from a chair loads the knee more than walking on level ground does.
func |>
summarise(
across(
c(mep_walk, mep_5tsts),
list(n = ~sum(!is.na(.x)), mean = ~mean(.x, na.rm = TRUE), sd = ~sd(.x, na.rm = TRUE))
)
)| mep_walk_n | mep_walk_mean | mep_walk_sd | mep_5tsts_n | mep_5tsts_mean | mep_5tsts_sd |
|---|---|---|---|---|---|
| 1015 | 0.6660099 | 1.282478 | 991 | 1.852674 | 1.97982 |
Since both scores are bounded, discrete difference scores rather than smooth continuous measures, a spearman (rank-based) correlation is a better choice here than pearson. We add use = "complete.obs" to ignore rows that are missing either score.
cor(func$mep_walk, func$mep_5tsts, method = "spearman", use = "complete.obs")[1] 0.2527416
The correlation is positive but modest (≈ 0.25). So the two tasks are related, but not highly correlated i.e., each one captures a somewhat different aspect of movement-evoked pain. For MEP, it is worth looking at both rather than treating one as a surrogate for the other. Visualize the relationship with a boxplot of sit-to-stand MEP grouped by walk MEP.
func |>
filter(!is.na(mep_walk), !is.na(mep_5tsts)) |>
ggplot(aes(x = factor(mep_walk), y = mep_5tsts, fill = factor(mep_walk))) +
geom_boxplot(show.legend = FALSE) +
labs(
title = "Sit-to-stand movement-evoked pain by 10-meter walk movement-evoked pain",
x = "Walk MEP (final minus initial pain)",
y = "5TSTS MEP (post minus pre pain)"
) +
theme_minimal()
2.3 Considerations While Working on the Project
2.3.1 Data Generation
Functional tasks are administered in person by trained study staff, following the A2CPS Manual of Procedures, which specifies the walk course, the sit-to-stand protocol, the deep breathing and coughing maneuver, and the timing of the pre- and post-task pain ratings. Raw responses are captured in REDCap, and the pain ratings are double-entered to reduce error. The steps that turn the REDCap export into the reformatted files (filtering, deduplication, error checking, and computing the MEP derivatives) are documented in the cleaning workflow that comes with the release (FuntionalCRF_data_quality_checks_and_reformat.html).
2.3.2 Other
- Movement-evoked pain, not resting pain. The MEP derivatives are change scores tied to a specific task, so they behave differently from resting or clinical pain measures.
- Cohort-specific batteries. The TKA and thoracic cohorts do not share tasks, so
mep_walkandmep_5tstsshow up only for the TKA cohort, while the thoracic file carries the deep breathing and coughing measures instead. - Sample size after merging. Joining with other modalities can shrink the effective sample size, so the counts are worth a glance after each merge.
- Quality control. The cleaning and derivative code has been reviewed, and a derivative can be reproduced from the raw pain ratings using the accompanying workflow.
2.3.3 Citations
In publications or presentations including data from A2CPS, please include the following statement as attribution:
Data were provided (in part) by the A2CPS Consortium funded by the National Institutes of Health (NIH) Common Fund, which is managed by the Office of the Director (OD)/Office of Strategic Coordination (OSC). Consortium components and their associated funding sources include Clinical Coordinating Center (U24NS112873), Data Integration and Resource Center (U54DA049110), Omics Data Generation Centers (U54DA049116, U54DA049115, U54DA049113), Multi-site Clinical Center 1 (MCC1) (UM1NS112874), and Multi-site Clinical Center 2 (MCC2) (UM1NS118922).