2025 Puget Sound Regional Council Household Travel Study Data User’s Guide

The 2025 Puget Sound Regional Council Household Travel Study is a comprehensive study of household demographics, daily travel activities, and typical transportation patterns throughout the central Puget Sound region. Data were collected from March 04, 2025 to June 08, 2025. The full set of questions asked can be viewed in the 2025 Puget Sound Travel Study Questionnaire. The survey dataset and more information can be downloaded on the PSRC Household Travel Survey Program website.

This guide provides an overview of the Study and practical guidance for understanding, analyzing, and interpreting the survey data for a variety of research and planning purposes. It is intended to serve as a comprehensive reference for anyone working with the dataset, from first-time users to experienced analysts.

NoteWhat’s in this Guide?

The guide is organized as follows:

  • Executive Summary – Highlights on sample outcomes, participant demographics, trip rates, mode share, trip purpose, work arrangements, telework, and deliveries
  • Study Overview – Study objectives, study area, data collection approach, inclusive engagement, and weighting context
  • Sample Design – Sampling goals, surveyable population, sample geographies and strata, recruitment procedures, incentives, and field monitoring
  • Survey Instrument – Recruit survey, travel diary, daily surveys, topic areas, travel date assignment, survey modes, and questionnaire updates
  • Data Processing and Structure – Completion flags and trip-table processing, including speed checks, distance measures, mode, departure time, and purpose
  • Weighting – Weighting goals, targets and inputs, multi-stage weighting steps, weighted totals, design effects, and analyst guidance
  • Dataset Overview – Data hierarchy, core tables, person-trip structure, record counts, data types, outliers, and missing values
  • Codebook – Searchable metadata for variable definitions, table membership, value labels, and analysis-ready category ordering
  • Analyst Handbook – Software setup, data loading and joins, choosing analytic units, working with categorical and numeric variables, weights, survey inference, and example workflows

Executive Summary

This section presents a summary of key findings from the 2025 Puget Sound Regional Travel Study.

0.1 Sample Plan Evaluation

The Study achieved a total of 2,772 completed households, which is above the overall upper target of 2,670. Within Pierce County, the add-on geography contributed 1020 completed households, above its target range of 780 to 900. Table 1 summarizes achieved complete households, household travel days, and person-days by reported home county; detailed target information is already provided in Section 2.1 and Table 22.

Code
hh_sample_plan <- as.data.frame(hts$hh) %>%
  mutate(
    geography = case_when(
      home_county == "Pierce County" & pierce_uninc_bg == "Pierce County unincorporated" ~ "Pierce County Add-on*",
      home_county == "Pierce County" ~ "Other areas in Pierce County",
      TRUE ~ as.character(home_county)
    )
  )

day_sample_plan <- as.data.frame(hts$day) %>%
  inner_join(
    hh_sample_plan %>% select(household_id, geography),
    by = "household_id"
  ) %>%
  filter(summary_complete == "Yes", !is.na(geography))

sample_plan_summary <- hh_sample_plan %>%
  filter(!is.na(geography)) %>%
  group_by(geography) %>%
  summarise(
    complete_households = n(),
    .groups = "drop"
  ) %>%
  left_join(
    day_sample_plan %>%
      group_by(geography) %>%
      summarise(
        complete_household_travel_days = n_distinct(paste(household_id, travel_date)),
        complete_person_days = n(),
        .groups = "drop"
      ),
    by = "geography"
  ) %>%
  mutate(
    complete_household_travel_days = replace_na(complete_household_travel_days, 0L),
    complete_person_days = replace_na(complete_person_days, 0L)
  ) %>%
  mutate(
    geography = factor(
      geography,
      levels = c(
        "King County",
        "Kitsap County",
        "Pierce County Add-on*",
        "Other areas in Pierce County",
        "Snohomish County"
      )
    )
  ) %>%
  arrange(geography) %>%
  mutate(geography = as.character(geography))

sample_plan_summary <- bind_rows(
  sample_plan_summary,
  sample_plan_summary %>%
    summarise(
      geography = "TOTAL",
      complete_households = sum(complete_households, na.rm = TRUE),
      complete_household_travel_days = sum(complete_household_travel_days, na.rm = TRUE),
      complete_person_days = sum(complete_person_days, na.rm = TRUE)
    )
)

sample_plan_summary %>%
  gt() %>%
  cols_label(
    geography = "Geography",
    complete_households = "Complete Households",
    complete_household_travel_days = "Complete Household Travel Days",
    complete_person_days = "Complete Person Days"
  ) %>%
  fmt_number(
    columns = c(complete_households, complete_household_travel_days, complete_person_days),
    decimals = 0
  ) %>%
  cols_align(
    align = "left",
    columns = geography
  ) %>%
  cols_align(
    align = "center",
    columns = c(complete_households, complete_household_travel_days, complete_person_days)
  ) %>%
  tab_style(
    style = cell_text(weight = "bold"),
    locations = cells_body(rows = geography == "TOTAL")
  ) %>%
  cols_width(
    geography ~ px(250),
    complete_households ~ px(125),
    complete_household_travel_days ~ px(180),
    complete_person_days ~ px(160)
  ) %>%
  tab_options(
    table.width = pct(100),
    column_labels.font.weight = "bold",
    table_body.hlines.width = px(1)
  ) %>%
  tab_source_note(
    source_note = md("*Pierce County Add-on: block groups identified by Pierce County staff as representing unincorporated parts of the county.")
  ) %>%
  opt_row_striping()
Geography Complete Households Complete Household Travel Days Complete Person Days
King County 1,142 2,691 4,046
Kitsap County 134 238 428
Pierce County Add-on* 1,020 1,889 3,199
Other areas in Pierce County 162 291 486
Snohomish County 314 573 999
TOTAL 2,772 5,682 9,158
*Pierce County Add-on: block groups identified by Pierce County staff as representing unincorporated parts of the county.
Table 1: Achieved completes by reported home county

0.2 Participant Demographics

Sample Overview

The achieved sample spans travel from March 04, 2025 through June 08, 2025, giving the study a broad reporting window. In total, the dataset includes 2,772 households, 5,558 persons, 10,868 person-days, and 38,980 trips. Those counts show that the executive summary findings draw from a sizable base of both household- and trip-level records. Table 2 shows the key summary metrics for the achieved sample.

Code
sample_ov_pre <- list(
  hh = hts$hh,
  person = hts$person,
  day = hts$day,
  trip = hts$trip
)

sample_ov_sum <- data.frame(
  measure = c(
    "First travel date",
    "Last travel date",
    "Households (unweighted)",
    "Persons (unweighted)",
    "Person-days (unweighted)",
    "Trips (unweighted)"
  ),
  value = c(
    format(min(sample_ov_pre$hh$traveldate_start, na.rm = TRUE), "%B %d, %Y"),
    format(max(sample_ov_pre$hh$traveldate_end, na.rm = TRUE), "%B %d, %Y"),
    format(dplyr::n_distinct(sample_ov_pre$hh$household_id), big.mark = ","),
    format(dplyr::n_distinct(sample_ov_pre$person$person_id), big.mark = ","),
    format(dplyr::n_distinct(sample_ov_pre$day$day_id), big.mark = ","),
    format(dplyr::n_distinct(sample_ov_pre$trip$trip_id), big.mark = ",")
  ),
  stringsAsFactors = FALSE
)

sample_ov_fmt <- sample_ov_sum

sample_ov_fmt %>%
  gt() %>%
  tab_header(
    title = "2025 Sample Overview",
    subtitle = "Travel dates and unweighted record counts"
  ) %>%
  opt_row_striping()
2025 Sample Overview
Travel dates and unweighted record counts
measure value
First travel date March 04, 2025
Last travel date June 08, 2025
Households (unweighted) 2,772
Persons (unweighted) 5,558
Person-days (unweighted) 10,868
Trips (unweighted) 38,980
Table 2: 2025 sample overview, including travel dates and unweighted record counts

Weighted Sample Benchmarks

Table 3 compares the unweighted and weighted regional sample with benchmarks from the U.S. Census Bureau’s American Community Survey (ACS) for key demographic characteristics. Darker shading in the difference columns indicates larger departures from the ACS benchmark, with teal tones showing negative differences, purple tones showing positive differences, and near-zero differences fading to neutral. The weighted sample generally moves the survey closer to the ACS benchmarks for key characteristics, especially household size, vehicles, and age. Household income shows more notable deviations, particularly a slight overrepresentation of higher-income households and some middle-income categories. Overall, the weighting performs well.

Code
acs_benchmarks <- data.frame(
  section = c(
    rep("Household size", 3),
    rep("Household vehicles", 3),
    rep("Household income", 6),
    rep("Age", 4)
  ),
  category = c(
    "1-person", "2-person", "3+-person",
    "0 vehicles", "1 vehicle", "2+ vehicles",
    "Under $25,000", "$25,000-$49,999", "$50,000-$74,999",
    "$75,000-$99,999", "$100,000-$199,999", "$200,000 or more",
    "Under 18", "18-34", "35-64", "65+"
  ),
  acs_share = c(
    0.277, 0.343, 0.380,
    0.082, 0.320, 0.578,
    0.077, 0.091, 0.091, 0.107, 0.260, 0.187,
    0.202, 0.249, 0.395, 0.154
  ),
  stringsAsFactors = FALSE
)

benchmark_levels <- c(
  "1-person", "2-person", "3+-person",
  "0 vehicles", "1 vehicle", "2+ vehicles",
  "Under $25,000", "$25,000-$49,999", "$50,000-$74,999",
  "$75,000-$99,999", "$100,000-$199,999", "$200,000 or more",
  "Under 18", "18-34", "35-64", "65+"
)

benchmark_sections <- c("Age", "Household income", "Household size", "Household vehicles")

benchmark_inputs <- bind_rows(
  as.data.frame(hts$hh) %>%
    filter(!is.na(hhsize), !is.na(hh_weight), hh_weight > 0) %>%
    transmute(
      record_id = household_id,
      sample_segment,
      analysis_weight = hh_weight,
      section = "Household size",
      category = case_when(
        hhsize == "1 person" ~ "1-person",
        hhsize == "2 people" ~ "2-person",
        TRUE ~ "3+-person"
      )
    ),
  as.data.frame(hts$hh) %>%
    filter(!is.na(vehicle_count), !is.na(hh_weight), hh_weight > 0) %>%
    transmute(
      record_id = household_id,
      sample_segment,
      analysis_weight = hh_weight,
      section = "Household vehicles",
      category = case_when(
        vehicle_count == "0 (no vehicles)" ~ "0 vehicles",
        vehicle_count == "1 vehicle" ~ "1 vehicle",
        TRUE ~ "2+ vehicles"
      )
    ),
  as.data.frame(hts$hh) %>%
    filter(!is.na(hhincome_broad), hhincome_broad != "Prefer not to answer", !is.na(hh_weight), hh_weight > 0) %>%
    transmute(
      record_id = household_id,
      sample_segment,
      analysis_weight = hh_weight,
      section = "Household income",
      category = hhincome_broad
    ),
  as.data.frame(hts$person) %>%
    left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
    filter(!is.na(age), !is.na(person_weight), person_weight > 0) %>%
    transmute(
      record_id = person_id,
      sample_segment,
      analysis_weight = person_weight,
      section = "Age",
      category = case_when(
        age < 18 ~ "Under 18",
        age <= 34 ~ "18-34",
        age <= 64 ~ "35-64",
        TRUE ~ "65+"
      )
    )
) %>%
  mutate(
    section = factor(section, levels = benchmark_sections),
    category = factor(category, levels = benchmark_levels)
  )

benchmark_summary <- benchmark_inputs %>%
  as_survey_design(ids = record_id, weights = analysis_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(section, category) %>%
  summarise(
    estimate = survey_prop(vartype = "se", proportion = TRUE),
    unweighted_n = unweighted(n()),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  group_by(section) %>%
  mutate(
    unweighted_share = unweighted_n / sum(unweighted_n),
    se = estimate_se
  ) %>%
  ungroup() %>%
  select(section, category, unweighted_share, estimate, se) %>%
  mutate(
    section = factor(as.character(section), levels = benchmark_sections),
    category = factor(as.character(category), levels = benchmark_levels)
  ) %>%
  left_join(acs_benchmarks, by = c("section", "category")) %>%
  mutate(
    unweighted_gap = unweighted_share - acs_share,
    unweighted_gap_pp = unweighted_gap * 100,
    gap = estimate - acs_share,
    gap_abs = abs(gap),
    gap_pp = gap * 100,
    rse = se / estimate,
    rse_flag = case_when(
      rse > 0.5 ~ "**",
      rse > 0.3 ~ "*",
      TRUE ~ ""
    )
  )

benchmark_gap_limit <- max(
  abs(c(benchmark_summary$unweighted_gap_pp, benchmark_summary$gap_pp)),
  na.rm = TRUE
)

benchmark_gap_palette <- c(
  "#005753",
  "#5D8884",
  "#A6BBB9",
  "#F1F1F1",
  "#D4AFD0",
  "#B46FAF",
  "#92278F"
)

benchmark_summary %>%
  # arrange(section, category) %>%
  mutate(
    category = as.character(category),
    unweighted_display = percent(unweighted_share, accuracy = 0.1),
    estimate_display = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    acs_display = percent(acs_share, accuracy = 0.1)
  ) %>%
  select(
    section,
    category,
    unweighted_display,
    estimate_display,
    acs_display,
    unweighted_gap_pp,
    gap_pp
  ) %>%
  gt(groupname_col = "section") %>%
  cols_label(
    category = "Category",
    unweighted_display = "Unweighted sample",
    estimate_display = "Weighted sample",
    acs_display = "ACS benchmark",
    unweighted_gap_pp = "Difference (unweighted - ACS, percentage points)",
    gap_pp = "Difference (weighted - ACS, percentage points)"
  ) %>%
  fmt_number(columns = c(unweighted_gap_pp, gap_pp), decimals = 1, force_sign = TRUE) %>%
  cols_align(
    align = "left",
    columns = category
  ) %>%
  cols_align(
    align = "center",
    columns = c(unweighted_display, estimate_display, acs_display, unweighted_gap_pp, gap_pp)
  ) %>%
  data_color(
    columns = c(unweighted_gap_pp, gap_pp),
    fn = scales::col_numeric(
      palette = benchmark_gap_palette,
      domain = c(-benchmark_gap_limit, benchmark_gap_limit)
    )
  ) %>%
  tab_style(
    style = cell_text(weight = "bold"),
    locations = cells_body(columns = c(unweighted_gap_pp, gap_pp), rows = abs(unweighted_gap_pp) >= 5 | abs(gap_pp) >= 5)
  ) %>%
  opt_row_striping()
Category Unweighted sample Weighted sample ACS benchmark Difference (unweighted - ACS, percentage points) Difference (weighted - ACS, percentage points)
Age
Under 18 4.5% 7.1% 20.2% −15.7 −13.1
18-34 18.8% 22.4% 24.9% −6.1 −2.5
35-64 46.2% 50.2% 39.5% +6.7 +10.7
65+ 30.6% 20.4% 15.4% +15.2 +5.0
Household income
Under $25,000 7.9% 10.5% 7.7% +0.2 +2.8
$25,000-$49,999 13.3% 10.3% 9.1% +4.2 +1.2
$50,000-$74,999 14.0% 13.4% 9.1% +4.9 +4.3
$75,000-$99,999 14.9% 12.6% 10.7% +4.2 +1.9
$100,000-$199,999 33.3% 29.2% 26.0% +7.3 +3.2
$200,000 or more 16.6% 24.0% 18.7% −2.1 +5.3
Household size
1-person 35.7% 27.5% 27.7% +8.0 −0.2
2-person 42.5% 34.8% 34.3% +8.2 +0.5
3+-person 21.8% 37.8% 38.0% −16.2 −0.2
Household vehicles
0 vehicles 8.2% 8.7% 8.2% 0.0 +0.5
1 vehicle 38.1% 32.0% 32.0% +6.1 +0.0
2+ vehicles 53.8% 59.3% 57.8% −4.0 +1.5
Table 3: Regional sample compared with ACS benchmarks. Bold values indicate categories where the sample differs from the ACS benchmark by 5.0 or more percentage points in either difference column.

0.3 Trips Per Person-Day

Average Trips per Person-Day

Code
trip_dow_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, travel_dow, day_weight) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      group_by(day_id) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "day_id"
  ) %>%
  mutate(
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  ) %>%
  filter(
    travel_dow %in% c("Monday", "Tuesday", "Wednesday", "Thursday"),
    !is.na(day_weight),
    day_weight > 0
  )

trip_dow_sum <- trip_dow_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(travel_dow) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_dow_pre %>%
      group_by(travel_dow) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        unweighted_trip_rate = unweighted_trips / unweighted_days,
        .groups = "drop"
      ),
    by = "travel_dow"
  )

trip_dow_fmt <- trip_dow_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(
      is.na(rse) ~ "",
      rse > 0.5 ~ "**",
      rse > 0.3 ~ "*",
      TRUE ~ ""
    ),
    label = paste0(number(estimate, accuracy = 0.01), stability),
    travel_dow = factor(travel_dow, levels = c("Monday", "Tuesday", "Wednesday", "Thursday"))
  ) %>%
  arrange(travel_dow)

trip_dow_x_max <- max(trip_dow_fmt$ci_high, na.rm = TRUE)
trip_dow_label_buffer <- trip_dow_x_max * 0.08
trip_dow_fmt <- trip_dow_fmt %>%
  mutate(label_x = ci_high + trip_dow_label_buffer)

Weekday trip-making is notably consistent from Monday through Thursday. Thursday records the highest estimated rate at 4.15 trips per person-day, but midweek travel is stable without meaningful day-to-day swings.

Code
trip_dow_plot <- ggplot(
  trip_dow_fmt,
  aes(
    x = estimate,
    y = factor(travel_dow, levels = rev(levels(travel_dow))),
    text = paste0(
      "Day: ", travel_dow,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01),
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips))
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(limits = c(0, max(trip_dow_fmt$label_x, na.rm = TRUE) + trip_dow_label_buffer * 2)) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_dow_fig <- ggplotly(trip_dow_plot, tooltip = "text", height = 420)
trip_dow_fig$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
trip_dow_fig$x$data[[2]]$hoverinfo <- "skip"
trip_dow_fig$x$data[[3]]$hoverinfo <- "skip"

trip_dow_fig %>%
  layout(margin = list(l = 110, r = 50, t = 40, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 1: Average trips per person-day by day of week. Estimates use day_weight.
Code
trip_dow_tbl <- bind_rows(
  trip_dow_fmt %>%
    transmute(
      day_of_week = as.character(travel_dow),
      unweighted_days,
      unweighted_trips,
      weighted_days,
      weighted_trips,
      unweighted_trip_rate,
      weighted_trip_rate = estimate,
      ci_low,
      ci_high,
      rse,
      stability
    ),
  trip_dow_pre %>%
    summarise(
      day_of_week = "Total",
      unweighted_days = n(),
      unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
      weighted_days = sum(day_weight, na.rm = TRUE),
      weighted_trips = sum(weighted_trips, na.rm = TRUE),
      unweighted_trip_rate = unweighted_trips / unweighted_days,
      weighted_trip_rate = weighted_trips / weighted_days,
      ci_low = NA_real_,
      ci_high = NA_real_,
      rse = NA_real_,
      stability = ""
    )
)

trip_dow_tbl %>%
  transmute(
    `Day of week` = day_of_week,
    `Weighted trips per person-day` = paste0(number(weighted_trip_rate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = if_else(
      is.na(ci_low),
      "\u2014",
      paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01))
    ),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips),
    `Unweighted trips per person-day` = number(unweighted_trip_rate, accuracy = 0.01)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day
Day of week Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips Unweighted trips per person-day
Monday 3.80 3.51 to 4.08 843,223 6,627,953,871 1,820 5,977 3.28
Tuesday 4.05 3.62 to 4.47 1,264,477 16,123,686,845 1,977 6,690 3.38
Wednesday 4.13 3.73 to 4.52 1,034,938 12,059,524,901 1,883 6,588 3.50
Thursday 4.15 3.81 to 4.49 1,088,125 12,564,915,945 1,972 6,864 3.48
Total 4.04 — 4,230,763 17,100,397 7,652 26,119 3.41
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 4: Average trips per person-day by day of week
NoteWhat do confidence intervals and relative standard errors mean?

A 95% confidence interval (CI) is the range of values that likely includes the true population value. Wide confidence intervals indicate more uncertainty either because of small sample sizes or high variability in the data, while narrow confidence intervals indicate more precision.

A standard error (SE) shows how much the estimate could vary from sample to sample. A relative standard error (RSE) shows that uncertainty as a percentage of the estimate, so larger RSEs mean less stable estimates.

Additional Trips Per Person-Day

Trips Per Person-Day by Household Income

Person trip rates are broadly similar across household income groups, with estimates ranging from approximately 3.8 to 4.7 trips per person-day. While middle-income households show slightly higher point estimates, the confidence intervals overlap across all income categories, indicating no clear or statistically distinguishable pattern by income.

Code
trip_x_inc_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, person_id, day_weight) %>%
  left_join(
    as.data.frame(hts$person) %>% select(person_id, age, paid_work),
    by = "person_id"
  ) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment, hhincome_broad),
    by = "household_id"
  ) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      group_by(day_id) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "day_id"
  ) %>%
  mutate(
    grp = if_else(hhincome_broad == "Prefer not to answer", NA_character_, hhincome_broad),
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  ) %>%
  filter(!is.na(grp), !is.na(day_weight), day_weight > 0)

trip_x_inc_sum <- trip_x_inc_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(grp) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_x_inc_pre %>%
      group_by(grp) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "grp"
  )

trip_x_inc_fmt <- trip_x_inc_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(is.na(rse) ~ "", rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    label = paste0(number(estimate, accuracy = 0.01), stability)
  ) %>%
  mutate(grp = factor(grp, levels = income_group_levels[income_group_levels %in% grp])) %>%
  arrange(grp)

trip_x_inc_x_max <- max(trip_x_inc_fmt$ci_high, na.rm = TRUE)
trip_x_inc_label_buffer <- trip_x_inc_x_max * 0.08
trip_x_inc_fmt <- trip_x_inc_fmt %>%
  mutate(
    label_x = ci_high + trip_x_inc_label_buffer,
    grp_display = factor(
      dplyr::recode(as.character(grp), !!!income_group_plot_label_map),
      levels = unname(income_group_plot_label_map[income_group_levels[income_group_levels %in% grp]])
    ),
    hover_text = paste0(
      "Household income: ", grp_display,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01), stability,
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips)),
      "<br>Unweighted days: ", comma(unweighted_days),
      "<br>Unweighted trips: ", comma(unweighted_trips)
    )
  )

trip_x_inc_plot <- ggplot(
  trip_x_inc_fmt,
  aes(
    x = estimate,
    y = factor(grp_display, levels = rev(levels(grp_display))),
    text = hover_text
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(
    aes(x = label_x, label = label),
    hjust = 0,
    color = single_bar_color,
    size = 4.1,
    family = "Inter",
    fontface = "bold"
  ) +
  labs(x = NULL, y = NULL) +
  scale_x_continuous(limits = c(0, max(trip_x_inc_fmt$label_x, na.rm = TRUE) + trip_x_inc_label_buffer * 2)) +
  scale_y_discrete(drop = FALSE) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_x_inc_fig <- ggplotly(
  trip_x_inc_plot,
  tooltip = "text",
  height = 430,
  dynamicTicks = FALSE
)

for (i in seq_along(trip_x_inc_fig$x$data)) {
  if (identical(trip_x_inc_fig$x$data[[i]]$type, "bar")) {
    trip_x_inc_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  } else {
    trip_x_inc_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

trip_x_inc_fig %>%
  layout(margin = list(l = 175, r = 50, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 2: Trips per person-day by household income. Estimates use day_weight.
Code
trip_x_inc_fmt %>%
  transmute(
    Group = grp,
    `Weighted trips per person-day` = paste0(number(estimate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01)),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day by Household Income") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day by Household Income
Group Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips
Under $25,000 4.18 3.63 to 4.73 240,796 2,597,450,119 378 1,239
$25,000-$49,999 4.32 3.50 to 5.15 325,163 3,080,491,783 808 2,760
$50,000-$74,999 4.65 3.79 to 5.50 480,831 7,410,465,230 871 3,271
$75,000-$99,999 3.76 3.42 to 4.10 413,276 3,811,756,638 952 3,054
$100,000-$199,999 3.81 3.50 to 4.12 1,297,125 12,550,313,194 2,587 8,839
$200,000 or more 4.15 3.79 to 4.51 1,085,037 12,033,538,709 1,476 5,273
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 5: Trips per person-day by household income

Trips Per Person-Day by Age Group

Trip rates follow a clear life-cycle pattern. Working-age adults make the most trips on an average day, while children and older adults make fewer, reinforcing the central role of work and household-serving travel in shaping daily mobility.

Code
trip_x_age_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, person_id, day_weight) %>%
  left_join(
    as.data.frame(hts$person) %>% select(person_id, age),
    by = "person_id"
  ) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      group_by(day_id) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "day_id"
  ) %>%
  mutate(
    grp = case_when(
      is.na(age) ~ NA_character_,
      age < 18 ~ "Under 18",
      age <= 34 ~ "18-34",
      age <= 64 ~ "35-64",
      TRUE ~ "65+"
    ),
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  ) %>%
  filter(!is.na(grp), !is.na(day_weight), day_weight > 0)

trip_x_age_sum <- trip_x_age_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(grp) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_x_age_pre %>%
      group_by(grp) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "grp"
  )

trip_x_age_fmt <- trip_x_age_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(is.na(rse) ~ "", rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    label = paste0(number(estimate, accuracy = 0.01), stability)
  ) %>%
  mutate(grp = factor(grp, levels = c("Under 18", "18-34", "35-64", "65+"))) %>%
  arrange(grp)

trip_x_age_x_max <- max(trip_x_age_fmt$ci_high, na.rm = TRUE)
trip_x_age_label_buffer <- trip_x_age_x_max * 0.08
trip_x_age_fmt <- trip_x_age_fmt %>%
  mutate(label_x = ci_high + trip_x_age_label_buffer)

trip_x_age_plot <- ggplot(
  trip_x_age_fmt,
  aes(
    x = estimate,
    y = factor(grp, levels=rev(levels(grp))),
    text = paste0(
      "Age group: ", grp,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01), stability,
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips)),
      "<br>Unweighted days: ", comma(unweighted_days),
      "<br>Unweighted trips: ", comma(unweighted_trips)
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(limits = c(0, max(trip_x_age_fmt$label_x, na.rm = TRUE) + trip_x_age_label_buffer * 2)) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_x_age_fig <- ggplotly(trip_x_age_plot, tooltip = "text", height = 430)
trip_x_age_fig$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
trip_x_age_fig$x$data[[2]]$hoverinfo <- "skip"
trip_x_age_fig$x$data[[3]]$hoverinfo <- "skip"

trip_x_age_fig %>%
  layout(margin = list(l = 120, r = 50, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 3: Trips per person-day by age group. Estimates use day_weight.
Code
trip_x_age_fmt %>%
  transmute(
    Group = grp,
    `Weighted trips per person-day` = paste0(number(estimate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01)),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day by Age Group") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day by Age Group
Group Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips
Under 18 3.47 3.03 to 3.91 299,236 3,873,059,272 249 838
18-34 4.12 3.70 to 4.55 946,041 10,220,799,805 1,709 6,203
35-64 4.19 3.93 to 4.45 2,123,130 25,050,291,056 3,689 13,137
65+ 3.78 3.43 to 4.13 862,355 8,231,931,430 2,005 5,941
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 6: Trips per person-day by age group

Trips Per Person-Day by Employment Status

Employment status is generally related to daily trip-making. Adults who are working or otherwise attached to the labor force make more trips per day than adults who are not employed, but the difference is within the margins of error. Those employed and on leave seem to make even more trips, but this is a small sample with large margins of error.

Code
trip_x_emp_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, person_id, day_weight) %>%
  left_join(
    as.data.frame(hts$person) %>% select(person_id, age, paid_work),
    by = "person_id"
  ) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      group_by(day_id) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "day_id"
  ) %>%
  mutate(
    grp = case_when(
      paid_work == "Yes" ~ "Employed",
      paid_work == "On leave (e.g., medical, parental)" ~ "On leave",
      paid_work == "No" ~ "Not employed",
      TRUE ~ NA_character_
    ),
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  ) %>%
  filter(age >= 16, !is.na(grp), !is.na(day_weight), day_weight > 0)

trip_x_emp_sum <- trip_x_emp_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(grp) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_x_emp_pre %>%
      group_by(grp) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "grp"
  )

trip_x_emp_fmt <- trip_x_emp_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(is.na(rse) ~ "", rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    label = paste0(number(estimate, accuracy = 0.01), stability)
  ) %>%
  arrange(desc(weighted_days)) %>%
  mutate(grp = factor(grp, levels = grp))

trip_x_emp_x_max <- max(trip_x_emp_fmt$ci_high, na.rm = TRUE)
trip_x_emp_label_buffer <- trip_x_emp_x_max * 0.08
trip_x_emp_fmt <- trip_x_emp_fmt %>%
  mutate(label_x = ci_high + trip_x_emp_label_buffer)

trip_x_emp_plot <- ggplot(
  trip_x_emp_fmt,
  aes(
    x = estimate,
    y = grp,
    text = paste0(
      "Employment status: ", grp,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01), stability,
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips)),
      "<br>Unweighted days: ", comma(unweighted_days),
      "<br>Unweighted trips: ", comma(unweighted_trips)
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(limits = c(0, max(trip_x_emp_fmt$label_x, na.rm = TRUE) + trip_x_emp_label_buffer * 2)) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_x_emp_fig <- ggplotly(trip_x_emp_plot, tooltip = "text", height = 430)
trip_x_emp_fig$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
trip_x_emp_fig$x$data[[2]]$hoverinfo <- "skip"
trip_x_emp_fig$x$data[[3]]$hoverinfo <- "skip"

trip_x_emp_fig %>%
  layout(margin = list(l = 120, r = 50, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 4: Trips per person-day by employment status. Estimates use day_weight.
Code
trip_x_emp_fmt %>%
  transmute(
    Group = grp,
    `Weighted trips per person-day` = paste0(number(estimate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01)),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day by Employment Status") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day by Employment Status
Group Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips
Employed 4.24 4.01 to 4.47 2,265,541 25,238,894,666 4,138 15,506
Not employed 3.95 3.61 to 4.28 1,155,639 11,770,408,017 2,814 8,216
On leave 5.31 3.99 to 6.62 47,030 484,171,117 69 272
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 7: Trips per person-day by employment status

Trips Per Person-Day by Travel Mode

Mode-specific trip rates show a pronounced hierarchy led by driving. Walking is the next most common mode on a per-person-day basis, while transit, biking or micromobility, and other modes account for much smaller amounts of daily travel.

Code
trip_x_mode_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, day_weight) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  filter(!is.na(day_weight), day_weight > 0) %>%
  crossing(category = c("Walk", "Bike / Micromobility", "Drive", "Transit", "Other")) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      transmute(
        day_id,
        category = case_when(
          mode_class_5 == "Walk" ~ "Walk",
          mode_class_5 == "Bike/Micromobility" ~ "Bike / Micromobility",
          mode_class_5 == "Drive" ~ "Drive",
          mode_class_5 == "Transit" ~ "Transit",
          mode_class_5 == "Other" ~ "Other",
          TRUE ~ NA_character_
        ),
        trip_weight
      ) %>%
      filter(!is.na(category)) %>%
      group_by(day_id, category) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = c("day_id", "category")
  ) %>%
  mutate(
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  )

trip_x_mode_sum <- trip_x_mode_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_x_mode_pre %>%
      group_by(category) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "category"
  )

trip_x_mode_fmt <- trip_x_mode_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(is.na(rse) ~ "", rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    label = paste0(number(estimate, accuracy = 0.01), stability)
  ) %>%
  arrange(desc(weighted_trips)) %>%
  mutate(grp = factor(category, levels = rev(category)))

trip_x_mode_x_max <- max(trip_x_mode_fmt$ci_high, na.rm = TRUE)
trip_x_mode_label_buffer <- trip_x_mode_x_max * 0.08
trip_x_mode_fmt <- trip_x_mode_fmt %>%
  mutate(label_x = ci_high + trip_x_mode_label_buffer)

trip_x_mode_plot <- ggplot(
  trip_x_mode_fmt,
  aes(
    x = estimate,
    y = grp,
    text = paste0(
      "Travel mode: ", grp,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01), stability,
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips)),
      "<br>Unweighted days: ", comma(unweighted_days),
      "<br>Unweighted trips: ", comma(unweighted_trips)
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(limits = c(0, max(trip_x_mode_fmt$label_x, na.rm = TRUE) + trip_x_mode_label_buffer * 2)) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_x_mode_fig <- ggplotly(trip_x_mode_plot, tooltip = "text", height = 430)
trip_x_mode_fig$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
trip_x_mode_fig$x$data[[2]]$hoverinfo <- "skip"
trip_x_mode_fig$x$data[[3]]$hoverinfo <- "skip"

trip_x_mode_fig %>%
  layout(margin = list(l = 160, r = 50, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 5: Trips per person-day by travel mode. Estimates use day_weight and explicitly include person-days with zero trips by the given mode.
Code
trip_x_mode_fmt %>%
  transmute(
    Group = grp,
    `Weighted trips per person-day` = paste0(number(estimate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01)),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day by Travel Mode") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day by Travel Mode
Group Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips
Drive 3.29 3.08 to 3.49 4,230,763 40,262,156,423 7,652 20,085
Walk 0.39 0.33 to 0.45 4,230,763 3,191,287,632 7,652 3,520
Other 0.16 0.12 to 0.21 4,230,763 2,239,244,781 7,652 777
Transit 0.14 0.11 to 0.16 4,230,763 907,609,136 7,652 1,259
Bike / Micromobility 0.06 0.04 to 0.09 4,230,763 771,019,026 7,652 477
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 8: Trips per person-day by travel mode

Trips Per Person-Day by Trip Purpose

Daily travel is concentrated in a relatively small set of recurring purposes. Home-based movement, shopping and errands, and meals, social, or recreational travel account for much of everyday trip-making, while other purposes occur less frequently.

Code
trip_x_purp_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, day_weight) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  filter(!is.na(day_weight), day_weight > 0) %>%
  crossing(category = purpose_categories) %>%
  left_join(
    as.data.frame(hts$trip) %>%
      left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
      filter(!is.na(day_id), !is.na(trip_weight), trip_weight > 0) %>%
      transmute(
        day_id,
        category = case_when(
          dest_purpose == "Went home" ~ "Home",
          dest_purpose %in% c(
            "Went to primary workplace",
            "Went to work-related activity (e.g., meeting, delivery, worksite)",
            "Went to other work-related activity"
          ) ~ "Work",
          dest_purpose %in% c(
            "Attend K-12 school",
            "Attend daycare or preschool",
            "Attend college/university",
            "Attend other education-related activity (e.g., field trip)",
            "Attend other type of class (e.g., cooking class)",
            "Attend vocational education class",
            "Attended school/class"
          ) ~ "School",
          dest_purpose %in% c(
            "Grocery shopping",
            "Other shopping (e.g., mall, pet store)",
            "Personal business (e.g., bank, post office)",
            "Medical appointment (e.g., doctor, dentist)",
            "Got gas",
            "Other appointment/errands",
            "Appointment, shopping, or errands (e.g., gas)"
          ) ~ "Shopping / errands",
          dest_purpose %in% c(
            "Went to restaurant to eat/get take-out",
            "Exercise or recreation (e.g., gym, jog, bike, walk dog)",
            "Recreational event (e.g., movies, sporting event)",
            "Social event (e.g., visit friends, family, co-workers)",
            "Other social/leisure",
            "Went to another residence (e.g., someone else's home, second home)",
            "Religious/civic/volunteer activity",
            "Volunteering",
            "Social, leisure, religious, entertainment activity"
          ) ~ "Meals / social / recreation",
          dest_purpose %in% c(
            "Pick someone up",
            "Drop someone off",
            "BOTH pick up AND drop off",
            "Accompany someone only (e.g., go along for the ride)",
            "Dropped off, picked up, or accompanied another person"
          ) ~ "Escort / care",
          dest_purpose %in% c(
            "Changed or transferred mode (e.g., change from ferry to bus)",
            "Went to temporary lodging (e.g., hotel, vacation rental)",
            "Other activity only (e.g., attend meeting, pick-up or drop-off item)",
            "Other reason",
            "Not imputable"
          ) ~ "Other / change mode / overnight",
          TRUE ~ NA_character_
        ),
        trip_weight
      ) %>%
      filter(!is.na(category)) %>%
      group_by(day_id, category) %>%
      summarise(
        unweighted_trips = n(),
        weighted_trips = sum(trip_weight, na.rm = TRUE),
        .groups = "drop"
      ),
    by = c("day_id", "category")
  ) %>%
  mutate(
    unweighted_trips = replace_na(unweighted_trips, 0),
    weighted_trips = replace_na(weighted_trips, 0),
    wtd_trips_on_day = if_else(day_weight > 0, weighted_trips / day_weight, 0)
  )

trip_x_purp_sum <- trip_x_purp_pre %>%
  as_survey_design(ids = household_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(
    estimate = survey_mean(wtd_trips_on_day, vartype = c("se", "ci"), na.rm = TRUE),
    weighted_days = survey_total(vartype = NULL),
    weighted_trips = survey_total(weighted_trips, vartype = NULL, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  as.data.frame() %>%
  left_join(
    trip_x_purp_pre %>%
      group_by(category) %>%
      summarise(
        unweighted_days = n(),
        unweighted_trips = sum(unweighted_trips, na.rm = TRUE),
        .groups = "drop"
      ),
    by = "category"
  )

trip_x_purp_fmt <- trip_x_purp_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    stability = case_when(is.na(rse) ~ "", rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    label = paste0(number(estimate, accuracy = 0.01), stability)
  ) %>%
  arrange(desc(weighted_trips)) %>%
  mutate(grp = factor(category, levels = rev(category)))

trip_x_purp_x_max <- max(trip_x_purp_fmt$ci_high, na.rm = TRUE)
trip_x_purp_label_buffer <- trip_x_purp_x_max * 0.08
trip_x_purp_fmt <- trip_x_purp_fmt %>%
  mutate(label_x = ci_high + trip_x_purp_label_buffer)

trip_x_purp_plot <- ggplot(
  trip_x_purp_fmt,
  aes(
    x = estimate,
    y = grp,
    text = paste0(
      "Trip purpose: ", grp,
      "<br>Weighted trips per person-day: ", number(estimate, accuracy = 0.01), stability,
      "<br>95% confidence interval (CI): ", number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01),
      "<br>Weighted days: ", comma(round(weighted_days)),
      "<br>Weighted trips: ", comma(round(weighted_trips)),
      "<br>Unweighted days: ", comma(unweighted_days),
      "<br>Unweighted trips: ", comma(unweighted_trips)
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(limits = c(0, max(trip_x_purp_fmt$label_x, na.rm = TRUE) + trip_x_purp_label_buffer * 2)) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

trip_x_purp_fig <- ggplotly(trip_x_purp_plot, tooltip = "text", height = 430)
trip_x_purp_fig$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
trip_x_purp_fig$x$data[[2]]$hoverinfo <- "skip"
trip_x_purp_fig$x$data[[3]]$hoverinfo <- "skip"

trip_x_purp_fig %>%
  layout(margin = list(l = 170, r = 50, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 6: Trips per person-day by trip purpose. Estimates use day_weight and explicitly include person-days with zero trips for the given purpose; * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Code
trip_x_purp_fmt %>%
  transmute(
    Group = grp,
    `Weighted trips per person-day` = paste0(number(estimate, accuracy = 0.01), stability),
    `95% confidence interval (CI)` = paste0(number(ci_low, accuracy = 0.01), " to ", number(ci_high, accuracy = 0.01)),
    `Weighted days` = comma(round(weighted_days)),
    `Weighted trips` = comma(round(weighted_trips)),
    `Unweighted days` = comma(unweighted_days),
    `Unweighted trips` = comma(unweighted_trips)
  ) %>%
  gt() %>%
  tab_header(title = "Trips Per Person-Day by Trip Purpose") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trips Per Person-Day by Trip Purpose
Group Weighted trips per person-day 95% confidence interval (CI) Weighted days Weighted trips Unweighted days Unweighted trips
Home 1.53 1.46 to 1.60 4,230,763 18,339,580,589 7,652 9,220
Meals / social / recreation 0.68 0.61 to 0.75 4,230,763 7,381,369,408 7,652 5,301
Shopping / errands 0.66 0.58 to 0.73 4,230,763 7,132,390,591 7,652 4,905
Work 0.55 0.50 to 0.61 4,230,763 5,979,825,816 7,652 3,933
Escort / care 0.35 0.29 to 0.42 4,230,763 4,664,086,225 7,652 1,609
School 0.22 0.19 to 0.25 4,230,763 3,266,805,069 7,652 787
Other / change mode / overnight 0.06 0.04 to 0.07 4,230,763 612,023,864 7,652 364
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 9: Trips per person-day by trip purpose

0.4 Mode Share

Regional Mode Share

Driving dominates the regional mode profile by a wide margin. Walking is the second-largest mode, transit accounts for a smaller share, and all remaining modes represent comparatively limited portions of total trips, highlighting the region’s continued dependence on auto travel.

Code
mode_shr_pre <- as.data.frame(hts$trip) %>%
  left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(!is.na(trip_weight), trip_weight > 0) %>%
  transmute(
    trip_id,
    sample_segment,
    trip_weight,
    category = case_when(
      mode_class_5 == "Walk" ~ "Walk",
      mode_class_5 == "Bike/Micromobility" ~ "Bike / Micromobility",
      mode_class_5 == "Drive" ~ "Drive",
      mode_class_5 == "Transit" ~ "Transit",
      mode_class_5 == "Other" ~ "Other",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(category))

mode_shr_sum <- mode_shr_pre %>%
  filter(!is.na(category)) %>%
  as_survey_design(ids = trip_id, weights = trip_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(estimate = survey_prop(vartype = c("se", "ci"), proportion = TRUE), .groups = "drop") %>%
  as.data.frame()

mode_shr_fmt <- mode_shr_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    short_label = case_when(
      category == "Walk" ~ "Walk",
      category == "Bike / Micromobility" ~ "Bike",
      category == "Drive" ~ "Drive",
      category == "Transit" ~ "Transit",
      category == "Other" ~ "Other",
      TRUE ~ category
    )
  ) %>%
  arrange(desc(estimate)) %>%
  mutate(category = factor(category, levels = category))

mode_shr_colors <- sequential_teal_palette(5, direction = "dark_to_light")
mode_share_fill_values <- setNames(
  rep(unname(stacked_palette), length.out = dplyr::n_distinct(mode_shr_fmt$category)),
  levels(mode_shr_fmt$category)
)
mode_share_text_values <- setNames(
  contrast_text_for_fill(unname(mode_share_fill_values)),
  names(mode_share_fill_values)
)

plot_ly(
  mode_shr_fmt,
  labels = ~category,
  values = ~estimate,
  type = "pie",
  hole = 0.62,
  sort = FALSE,
  text = ~ paste0(short_label, "<br>", percent(estimate, accuracy = 1), rse_flag),
  textinfo = "text",
  hovertext = ~ paste0(
    category,
    "<br>Share: ", percent(estimate, accuracy = 0.1),
    "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)
  ),
  hoverinfo = "text",
  direction = "clockwise",
  textposition = "inside",
  textfont = list(family = "Inter", size = 13, color = contrast_text_for_fill(mode_shr_colors)),
  marker = list(
    colors = mode_shr_colors,
    line = list(color = divider_color, width = 2)
  ),
  hovertemplate = "%{hovertext}<extra></extra>",
  showlegend = TRUE
) %>%
  layout(
    height = 360,
    margin = list(l = 10, r = 10, t = 10, b = 20),
    legend = list(orientation = "h", x = 0.1, y = -0.02, font = list(family = "Inter", size = 12)),
    paper_bgcolor = "rgba(0,0,0,0)",
    plot_bgcolor = "rgba(0,0,0,0)"
  ) %>%
  config(displayModeBar = FALSE)
Figure 7: Trip-level weighted mode share. Estimates use trip_weight.
Code
mode_shr_fmt %>%
  transmute(
    Mode = category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1)
  ) %>%
  gt() %>%
  tab_header(title = "Mode Share") %>%
  tab_source_note(md("Estimates use `trip_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Mode Share
Mode Share 95% confidence interval (CI) Relative standard error (RSE)
Drive 81.3% 80.3% to 82.3% 0.6%
Walk 9.7% 9.0% to 10.4% 3.6%
Other 4.1% 3.5% to 4.7% 7.2%
Transit 3.4% 3.1% to 3.8% 5.6%
Bike / Micromobility 1.5% 1.2% to 1.9% 11.3%
Estimates use trip_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 10: Mode share

Weighted Travel Mode

Figure 8, Figure 9, and Figure 10 show weighted travel mode share by household income, age group, and time of day, respectively. These charts use trip-level weights to calculate mode share within each category, so shares reflect the distribution of trips rather than person-days. Across all three dimensions, driving is the dominant mode across most categories. Notable patterns include higher walking and transit shares among lower-income households and younger age groups, and a higher drive share during weekday peak periods.

Travel Mode by Household Income

Mode share differs by household income, but every income group remains predominantly auto-oriented. Lower-income households devote relatively larger shares of trips to walking and transit, while higher-income households are more heavily concentrated in driving, suggesting that transportation options and constraints vary meaningfully across the income spectrum.

Code
mode_x_inc_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, hhincome_broad, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = if_else(hhincome_broad == "Prefer not to answer", NA_character_, hhincome_broad),
    category = case_when(
      mode_class_5 == "Walk" ~ "Walk",
      mode_class_5 == "Bike/Micromobility" ~ "Bike / Micromobility",
      mode_class_5 == "Drive" ~ "Drive",
      mode_class_5 == "Transit" ~ "Transit",
      mode_class_5 == "Other" ~ "Other",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

mode_x_inc_group_levels <- income_group_levels

mode_x_inc_totals <- mode_x_inc_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

mode_x_inc_estimates <- lapply(mode_x_inc_group_levels, function(group_value) {
  mode_x_inc_group <- mode_x_inc_pre %>%
    filter(group_label == group_value)

  if (nrow(mode_x_inc_group) == 0) {
    return(NULL)
  }

  lapply(levels(mode_shr_fmt$category), function(category_value) {
    mode_x_inc_analysis <- mode_x_inc_group %>%
      mutate(indicator = as.numeric(category == category_value))

    mode_x_inc_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = mode_x_inc_analysis,
      nest = TRUE
    )

    mode_x_inc_estimate <- suppressWarnings(survey::svymean(~indicator, mode_x_inc_design, na.rm = TRUE))
    mode_x_inc_ci <- suppressWarnings(stats::confint(mode_x_inc_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(mode_x_inc_estimate)[1]),
      se = unname(survey::SE(mode_x_inc_estimate)[1]),
      ci_low = pmax(0, unname(mode_x_inc_ci[1])),
      ci_high = pmin(1, unname(mode_x_inc_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

mode_x_inc_fmt <- mode_x_inc_estimates %>%
  left_join(mode_x_inc_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(
      group_label,
      levels = income_group_levels[income_group_levels %in% group_label]
    ),
    group_display = factor(
      dplyr::recode(as.character(group_label), !!!income_group_short_map),
      levels = unname(income_group_short_map[income_group_levels[income_group_levels %in% group_label]])
    ),
    group_display_full = dplyr::recode(as.character(group_label), !!!income_group_plot_label_map),
    category = factor(category, levels = levels(mode_shr_fmt$category))
  ) %>%
  arrange(group_label, category) %>%
  mutate(
    hover_text = paste0(
      "Household income: ", group_display_full,
      "<br>Mode: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )

mode_x_inc_plot <- ggplot(
  mode_x_inc_fmt,
  aes(
    x = group_display,
    y = estimate,
    fill = category,
    text = hover_text
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = mode_share_fill_values) +
  scale_color_manual(values = mode_share_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted mode share within each household income group") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

mode_x_inc_fig <- ggplotly(mode_x_inc_plot, tooltip = "text", height = 430, dynamicTicks = FALSE)

for (i in seq_along(mode_x_inc_fig$x$data)) {
  mode_x_inc_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  mode_x_inc_fig$x$data[[i]]$showlegend <- identical(mode_x_inc_fig$x$data[[i]]$type, "bar")
  if (!is.null(mode_x_inc_fig$x$data[[i]]$name)) {
    mode_x_inc_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", mode_x_inc_fig$x$data[[i]]$name)
    mode_x_inc_fig$x$data[[i]]$legendgroup <- mode_x_inc_fig$x$data[[i]]$name
  }
  if (!identical(mode_x_inc_fig$x$data[[i]]$type, "bar")) {
    mode_x_inc_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

mode_x_inc_fig %>%
  layout(
    margin = list(l = 40, r = 20, t = 40, b = 110)
  ) %>%
  config(displayModeBar = FALSE)
Figure 8: Weighted share of trips by travel mode across household income groups. Shares use trip_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
mode_x_inc_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Travel Mode by Household Income (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Travel Mode by Household Income (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
Under $25,000
Drive 61.7% 2.5% 56.9% to 66.5% 4.0% 621,178 786
Walk 12.8% 1.5% 9.8% to 15.7% 11.9% 128,474 204
Other 9.0% 1.6% 5.9% to 12.1% 17.4% 90,680 63
Transit 10.0% 1.3% 7.4% to 12.5% 13.0% 100,264 142
Bike / Micromobility 6.6% 1.3% 4.1% to 9.1% 19.4% 66,552 44
$25,000-$49,999
Drive 79.0% 1.7% 75.7% to 82.2% 2.1% 1,110,366 2,128
Walk 13.0% 1.5% 10.1% to 16.0% 11.6% 183,218 351
Other 2.6% 0.6% 1.4% to 3.8% 23.7% 36,934 90
Transit 3.8% 0.5% 2.9% to 4.7% 11.9% 54,028 140
Bike / Micromobility 1.5% 0.4% 0.7% to 2.3% 27.6% 21,412 51
$50,000-$74,999
Drive 84.8% 1.3% 82.3% to 87.2% 1.5% 1,893,703 2,637
Walk 6.7% 0.7% 5.3% to 8.2% 11.1% 150,079 368
Other 2.6% 0.5% 1.7% to 3.5% 17.9% 58,092 80
Transit 3.8% 0.6% 2.5% to 5.0% 17.1% 84,439 152
Bike / Micromobility 2.1%* 0.7% 0.7% to 3.5% 33.7% 47,709 34
$75,000-$99,999
Drive 86.0% 1.1% 83.8% to 88.1% 1.3% 1,335,849 2,369
Walk 7.5% 0.7% 6.2% to 8.9% 9.3% 116,960 409
Other 2.1% 0.6% 1.0% to 3.2% 26.3% 32,685 60
Transit 3.9% 0.6% 2.7% to 5.2% 16.5% 61,114 174
Bike / Micromobility 0.5% 0.1% 0.3% to 0.7% 23.1% 7,454 42
$100,000-$199,999
Drive 83.4% 0.8% 81.7% to 85.0% 1.0% 4,120,859 6,937
Walk 8.4% 0.5% 7.4% to 9.4% 5.9% 415,171 1,110
Other 4.6% 0.6% 3.4% to 5.7% 12.5% 225,690 267
Transit 2.9% 0.4% 2.2% to 3.6% 12.4% 144,690 364
Bike / Micromobility 0.8% 0.2% 0.4% to 1.1% 26.0% 37,587 161
$200,000 or more
Drive 78.8% 1.1% 76.7% to 80.9% 1.3% 3,548,673 3,791
Walk 12.9% 0.9% 11.2% to 14.6% 6.6% 582,291 947
Other 4.2% 0.6% 3.0% to 5.4% 14.4% 188,919 162
Transit 2.4% 0.3% 1.9% to 3.0% 11.8% 110,259 236
Bike / Micromobility 1.6% 0.3% 1.0% to 2.3% 19.3% 74,061 136
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 11: Travel mode by household income (weighted)

Travel Mode by Age Group

Younger travelers make relatively greater use of non-auto modes, while older adults are the most drive-oriented, pointing to important differences in mobility patterns across the life course.

Code
mode_x_age_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$person) %>% select(person_id, age),
    by = "person_id"
  ) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = case_when(
      is.na(age) ~ NA_character_,
      age < 18 ~ "Under 18",
      age <= 34 ~ "18-34",
      age <= 64 ~ "35-64",
      TRUE ~ "65+"
    ),
    category = case_when(
      mode_class_5 == "Walk" ~ "Walk",
      mode_class_5 == "Bike/Micromobility" ~ "Bike / Micromobility",
      mode_class_5 == "Drive" ~ "Drive",
      mode_class_5 == "Transit" ~ "Transit",
      mode_class_5 == "Other" ~ "Other",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

mode_x_age_group_levels <- c("Under 18", "18-34", "35-64", "65+")

mode_x_age_totals <- mode_x_age_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

mode_x_age_estimates <- lapply(mode_x_age_group_levels, function(group_value) {
  mode_x_age_group <- mode_x_age_pre %>%
    filter(group_label == group_value)

  if (nrow(mode_x_age_group) == 0) {
    return(NULL)
  }

  lapply(levels(mode_shr_fmt$category), function(category_value) {
    mode_x_age_analysis <- mode_x_age_group %>%
      mutate(indicator = as.numeric(category == category_value))

    mode_x_age_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = mode_x_age_analysis,
      nest = TRUE
    )

    mode_x_age_estimate <- suppressWarnings(survey::svymean(~indicator, mode_x_age_design, na.rm = TRUE))
    mode_x_age_ci <- suppressWarnings(stats::confint(mode_x_age_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(mode_x_age_estimate)[1]),
      se = unname(survey::SE(mode_x_age_estimate)[1]),
      ci_low = pmax(0, unname(mode_x_age_ci[1])),
      ci_high = pmin(1, unname(mode_x_age_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

mode_x_age_fmt <- mode_x_age_estimates %>%
  left_join(mode_x_age_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(group_label, levels = c("Under 18", "18-34", "35-64", "65+")),
    category = factor(category, levels = levels(mode_shr_fmt$category))
  ) %>%
  arrange(group_label, category)

mode_x_age_plot <- ggplot(
  mode_x_age_fmt,
  aes(
    x = group_label,
    y = estimate,
    fill = category,
    text = paste0(
      "Age group: ", group_label,
      "<br>Mode: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = mode_share_fill_values) +
  scale_color_manual(values = mode_share_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted mode share within each age group") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

mode_x_age_fig <- ggplotly(mode_x_age_plot, tooltip = "text", height = 430)
for (i in seq_along(mode_x_age_fig$x$data)) {
  mode_x_age_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  mode_x_age_fig$x$data[[i]]$showlegend <- identical(mode_x_age_fig$x$data[[i]]$type, "bar")
  if (!is.null(mode_x_age_fig$x$data[[i]]$name)) {
    mode_x_age_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", mode_x_age_fig$x$data[[i]]$name)
    mode_x_age_fig$x$data[[i]]$legendgroup <- mode_x_age_fig$x$data[[i]]$name
  }
  if (!identical(mode_x_age_fig$x$data[[i]]$type, "bar")) {
    mode_x_age_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

mode_x_age_fig %>%
  layout(
    margin = list(l = 40, r = 20, t = 40, b = 70)
  ) %>%
  config(displayModeBar = FALSE)
Figure 9: Weighted share of trips by travel mode across age groups. Shares use trip_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
mode_x_age_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Travel Mode by Age Group (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Travel Mode by Age Group (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
Under 18
Drive 64.7% 3.0% 58.8% to 70.7% 4.7% 672,358 565
Walk 9.5% 1.9% 5.7% to 13.3% 20.5% 98,238 66
Other 17.4% 2.5% 12.6% to 22.2% 14.2% 180,873 158
Transit 5.6% 1.2% 3.2% to 8.0% 21.9% 57,918 43
Bike / Micromobility 2.8%* 1.3% 0.2% to 5.4% 47.2% 29,174 6
18-34
Drive 77.2% 1.0% 75.3% to 79.2% 1.3% 3,014,325 4,098
Walk 12.4% 0.8% 10.9% to 13.9% 6.1% 483,438 1,222
Other 1.7% 0.4% 0.9% to 2.5% 23.6% 67,100 115
Transit 6.2% 0.5% 5.1% to 7.2% 8.5% 240,136 599
Bike / Micromobility 2.5% 0.4% 1.7% to 3.3% 15.7% 97,182 169
35-64
Drive 83.3% 0.7% 82.0% to 84.6% 0.8% 7,408,875 10,262
Walk 9.0% 0.5% 8.1% to 10.0% 5.3% 801,852 1,703
Other 3.9% 0.4% 3.2% to 4.7% 9.6% 349,692 402
Transit 2.4% 0.2% 2.0% to 2.7% 8.1% 210,003 507
Bike / Micromobility 1.4% 0.2% 1.0% to 1.9% 16.1% 126,307 262
65+
Drive 86.1% 0.9% 84.4% to 87.9% 1.0% 2,808,583 5,160
Walk 8.4% 0.7% 7.1% to 9.7% 8.0% 274,451 529
Other 2.9% 0.5% 1.9% to 4.0% 18.4% 94,831 102
Transit 2.4% 0.4% 1.6% to 3.1% 15.8% 76,688 110
Bike / Micromobility 0.2% 0.0% 0.1% to 0.3% 21.8% 6,188 40
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 12: Travel mode by age group (weighted)

Travel Mode by Time of Day

Mode use varies over the day, but driving remains the leading mode in every time period. Midday and evening trips show a somewhat more varied modal mix than the peak periods, indicating that non-auto travel is more visible outside the core commute windows even as auto travel continues to dominate overall.

Code
mode_x_tod_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = case_when(
      is.na(depart_time_hour) ~ NA_character_,
      depart_time_hour < 6 ~ "Early AM (12-5)",
      depart_time_hour < 10 ~ "AM peak (6-9)",
      depart_time_hour < 16 ~ "Midday (10-3)",
      depart_time_hour < 19 ~ "PM peak (4-6)",
      TRUE ~ "Evening (7-11)"
    ),
    category = case_when(
      mode_class_5 == "Walk" ~ "Walk",
      mode_class_5 == "Bike/Micromobility" ~ "Bike / Micromobility",
      mode_class_5 == "Drive" ~ "Drive",
      mode_class_5 == "Transit" ~ "Transit",
      mode_class_5 == "Other" ~ "Other",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

mode_x_tod_group_levels <- c(
  "Early AM (12-5)",
  "AM peak (6-9)",
  "Midday (10-3)",
  "PM peak (4-6)",
  "Evening (7-11)"
)

mode_x_tod_totals <- mode_x_tod_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

mode_x_tod_estimates <- lapply(mode_x_tod_group_levels, function(group_value) {
  mode_x_tod_group <- mode_x_tod_pre %>%
    filter(group_label == group_value)

  if (nrow(mode_x_tod_group) == 0) {
    return(NULL)
  }

  lapply(levels(mode_shr_fmt$category), function(category_value) {
    mode_x_tod_analysis <- mode_x_tod_group %>%
      mutate(indicator = as.numeric(category == category_value))

    mode_x_tod_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = mode_x_tod_analysis,
      nest = TRUE
    )

    mode_x_tod_estimate <- suppressWarnings(survey::svymean(~indicator, mode_x_tod_design, na.rm = TRUE))
    mode_x_tod_ci <- suppressWarnings(stats::confint(mode_x_tod_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(mode_x_tod_estimate)[1]),
      se = unname(survey::SE(mode_x_tod_estimate)[1]),
      ci_low = pmax(0, unname(mode_x_tod_ci[1])),
      ci_high = pmin(1, unname(mode_x_tod_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

mode_x_tod_fmt <- mode_x_tod_estimates %>%
  left_join(mode_x_tod_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(group_label, levels = c(
      "Early AM (12-5)",
      "AM peak (6-9)",
      "Midday (10-3)",
      "PM peak (4-6)",
      "Evening (7-11)"
    )),
    category = factor(category, levels = levels(mode_shr_fmt$category))
  ) %>%
  arrange(group_label, category)

mode_x_tod_plot <- ggplot(
  mode_x_tod_fmt,
  aes(
    x = group_label,
    y = estimate,
    fill = category,
    text = paste0(
      "Time of day: ", group_label,
      "<br>Mode: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = mode_share_fill_values) +
  scale_color_manual(values = mode_share_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted mode share within each time-of-day bin") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

mode_x_tod_fig <- ggplotly(mode_x_tod_plot, tooltip = "text", height = 430)
for (i in seq_along(mode_x_tod_fig$x$data)) {
  mode_x_tod_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  mode_x_tod_fig$x$data[[i]]$showlegend <- identical(mode_x_tod_fig$x$data[[i]]$type, "bar")
  if (!is.null(mode_x_tod_fig$x$data[[i]]$name)) {
    mode_x_tod_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", mode_x_tod_fig$x$data[[i]]$name)
    mode_x_tod_fig$x$data[[i]]$legendgroup <- mode_x_tod_fig$x$data[[i]]$name
  }
  if (!identical(mode_x_tod_fig$x$data[[i]]$type, "bar")) {
    mode_x_tod_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

mode_x_tod_fig %>%
  layout(margin = list(l = 40, r = 20, t = 40, b = 70)) %>%
  config(displayModeBar = FALSE)
Figure 10: Weighted share of trips by travel mode across time-of-day periods based on trip departure hour. Shares use trip_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
mode_x_tod_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Travel Mode by Time of Day (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Travel Mode by Time of Day (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
Early AM (12-5)
Drive 78.7% 3.8% 71.3% to 86.1% 4.8% 383,220 554
Walk 5.5% 1.3% 2.9% to 8.1% 24.1% 26,951 54
Other 8.5%* 3.3% 2.1% to 14.9% 38.4% 41,333 20
Transit 3.1% 0.7% 1.7% to 4.5% 23.5% 15,043 55
Bike / Micromobility 4.2%* 2.0% 0.2% to 8.2% 48.4% 20,472 13
AM peak (6-9)
Drive 79.8% 1.1% 77.6% to 82.0% 1.4% 2,958,135 4,100
Walk 9.0% 0.8% 7.4% to 10.6% 8.9% 333,392 581
Other 5.7% 0.7% 4.4% to 7.1% 12.3% 213,030 263
Transit 4.1% 0.5% 3.2% to 4.9% 11.2% 150,121 311
Bike / Micromobility 1.4% 0.3% 0.8% to 2.0% 22.6% 50,883 107
Midday (10-3)
Drive 81.4% 0.8% 79.9% to 83.0% 1.0% 5,547,524 8,735
Walk 9.8% 0.6% 8.7% to 10.9% 5.8% 665,979 1,514
Other 4.3% 0.5% 3.4% to 5.2% 10.6% 294,219 336
Transit 3.1% 0.3% 2.5% to 3.7% 10.2% 212,294 393
Bike / Micromobility 1.4% 0.3% 0.8% to 1.9% 19.9% 92,296 159
PM peak (4-6)
Drive 81.0% 1.0% 79.2% to 82.9% 1.2% 3,374,966 4,538
Walk 11.3% 0.7% 9.8% to 12.7% 6.4% 468,567 926
Other 2.2% 0.5% 1.3% to 3.1% 20.9% 92,969 107
Transit 3.7% 0.4% 3.0% to 4.4% 10.1% 153,962 366
Bike / Micromobility 1.8% 0.4% 1.0% to 2.5% 20.8% 73,754 139
Evening (7-11)
Drive 85.0% 1.1% 82.8% to 87.2% 1.3% 1,640,296 2,158
Walk 8.5% 0.8% 6.9% to 10.0% 9.6% 163,092 445
Other 2.6% 0.6% 1.4% to 3.9% 23.4% 50,944 51
Transit 2.8% 0.4% 1.9% to 3.6% 15.0% 53,325 134
Bike / Micromobility 1.1% 0.3% 0.5% to 1.7% 27.4% 21,447 59
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 13: Travel mode by time of day (weighted)

0.5 Trip Purpose

Regional Trip Purpose

Trip purposes are concentrated in a few major categories rather than spread evenly across activities. Home is the single largest destination purpose, and shopping and errands plus meals, social, and recreational travel also account for substantial shares, emphasizing the importance of non-work travel in the regional pattern.

Code
purp_shr_pre <- as.data.frame(hts$trip) %>%
  left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(!is.na(trip_weight), trip_weight > 0) %>%
  transmute(
    trip_id,
    sample_segment,
    trip_weight,
    category = case_when(
      dest_purpose == "Went home" ~ "Home",
      dest_purpose %in% c(
        "Went to primary workplace",
        "Went to work-related activity (e.g., meeting, delivery, worksite)",
        "Went to other work-related activity"
      ) ~ "Work",
      dest_purpose %in% c(
        "Attend K-12 school",
        "Attend daycare or preschool",
        "Attend college/university",
        "Attend other education-related activity (e.g., field trip)",
        "Attend other type of class (e.g., cooking class)",
        "Attend vocational education class",
        "Attended school/class"
      ) ~ "School",
      dest_purpose %in% c(
        "Grocery shopping",
        "Other shopping (e.g., mall, pet store)",
        "Personal business (e.g., bank, post office)",
        "Medical appointment (e.g., doctor, dentist)",
        "Got gas",
        "Other appointment/errands",
        "Appointment, shopping, or errands (e.g., gas)"
      ) ~ "Shopping / errands",
      dest_purpose %in% c(
        "Went to restaurant to eat/get take-out",
        "Exercise or recreation (e.g., gym, jog, bike, walk dog)",
        "Recreational event (e.g., movies, sporting event)",
        "Social event (e.g., visit friends, family, co-workers)",
        "Other social/leisure",
        "Went to another residence (e.g., someone else's home, second home)",
        "Religious/civic/volunteer activity",
        "Volunteering",
        "Social, leisure, religious, entertainment activity"
      ) ~ "Meals / social / recreation",
      dest_purpose %in% c(
        "Pick someone up",
        "Drop someone off",
        "BOTH pick up AND drop off",
        "Accompany someone only (e.g., go along for the ride)",
        "Dropped off, picked up, or accompanied another person"
      ) ~ "Escort / care",
      dest_purpose %in% c(
        "Changed or transferred mode (e.g., change from ferry to bus)",
        "Went to temporary lodging (e.g., hotel, vacation rental)",
        "Other activity only (e.g., attend meeting, pick-up or drop-off item)",
        "Other reason",
        "Not imputable"
      ) ~ "Other / change mode / overnight",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(category))

purp_shr_sum <- purp_shr_pre %>%
  filter(!is.na(category)) %>%
  as_survey_design(ids = trip_id, weights = trip_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(estimate = survey_prop(vartype = c("se", "ci"), proportion = TRUE), .groups = "drop") %>%
  as.data.frame()

purp_shr_fmt <- purp_shr_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    rse_flag = case_when(
      rse > 0.5 ~ "**",
      rse > 0.3 ~ "*",
      TRUE ~ ""
    ),
    short_label = case_when(
      category == "Home" ~ "Home",
      category == "Work" ~ "Work",
      category == "School" ~ "Schl",
      category == "Shopping / errands" ~ "Shop",
      category == "Meals / social / recreation" ~ "Meal",
      category == "Escort / care" ~ "Care",
      category == "Other / change mode / overnight" ~ "Other",
      TRUE ~ category
    )
  ) %>%
  arrange(desc(estimate)) %>%
  mutate(category = factor(category, levels = category))

plot_ly(
  purp_shr_fmt,
  labels = ~category,
  values = ~estimate,
  type = "pie",
  hole = 0.62,
  sort = FALSE,
  text = ~ paste0(short_label, "<br>", percent(estimate, accuracy = 1), rse_flag),
  textinfo = "text",
  hovertext = ~ paste0(
    category,
    "<br>Share: ", percent(estimate, accuracy = 0.1),
    "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)
  ),
  hoverinfo = "text",
  direction = "clockwise",
  textposition = "inside",
  textfont = list(family = "Inter", size = 13, color = text_on_dark),
  marker = list(
    colors = unname(stacked_palette),
    line = list(color = divider_color, width = 2)
  ),
  hovertemplate = "%{hovertext}<extra></extra>",
  showlegend = TRUE
) %>%
  layout(
    height = 360,
    margin = list(l = 10, r = 10, t = 10, b = 20),
    legend = list(orientation = "h", x = 0.05, y = -0.02, font = list(family = "Inter", size = 12)),
    paper_bgcolor = "rgba(0,0,0,0)",
    plot_bgcolor = "rgba(0,0,0,0)",
    font = list(family = "Inter", color = brand_colors[["ink"]])
  ) %>%
  config(displayModeBar = FALSE)
Figure 11: Trip-level weighted purpose share. Estimates use trip_weight.
Code
purp_shr_fmt %>%
  transmute(
    Purpose = category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1)
  ) %>%
  gt() %>%
  tab_header(title = "Purpose Share") %>%
  tab_source_note(md("Estimates use `trip_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Purpose Share
Purpose Share 95% confidence interval (CI) Relative standard error (RSE)
Home 37.8% 36.5% to 39.2% 1.8%
Meals / social / recreation 16.8% 15.8% to 17.8% 3.1%
Shopping / errands 16.3% 15.3% to 17.3% 3.2%
Work 13.6% 12.7% to 14.6% 3.5%
Escort / care 8.7% 7.9% to 9.6% 4.9%
School 5.4% 4.8% to 6.1% 6.1%
Other / change mode / overnight 1.4% 1.1% to 1.8% 11.7%
Estimates use trip_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 14: Trip purpose share

Weighted Trip Purpose

Trip Purpose by Household Income

The mix of trip purposes shifts somewhat by household income, although the broad structure is similar across groups. Higher-income households tend to devote relatively larger shares of travel to work, while lower-income households show comparatively larger shares for maintenance and discretionary trips.

Code
purp_x_inc_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, hhincome_broad, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = if_else(hhincome_broad == "Prefer not to answer", NA_character_, hhincome_broad),
    category = case_when(
      dest_purpose == "Went home" ~ "Home",
      dest_purpose %in% c(
        "Went to primary workplace",
        "Went to work-related activity (e.g., meeting, delivery, worksite)",
        "Went to other work-related activity"
      ) ~ "Work",
      dest_purpose %in% c(
        "Attend K-12 school",
        "Attend daycare or preschool",
        "Attend college/university",
        "Attend other education-related activity (e.g., field trip)",
        "Attend other type of class (e.g., cooking class)",
        "Attend vocational education class",
        "Attended school/class"
      ) ~ "School",
      dest_purpose %in% c(
        "Grocery shopping",
        "Other shopping (e.g., mall, pet store)",
        "Personal business (e.g., bank, post office)",
        "Medical appointment (e.g., doctor, dentist)",
        "Got gas",
        "Other appointment/errands",
        "Appointment, shopping, or errands (e.g., gas)"
      ) ~ "Shopping / errands",
      dest_purpose %in% c(
        "Went to restaurant to eat/get take-out",
        "Exercise or recreation (e.g., gym, jog, bike, walk dog)",
        "Recreational event (e.g., movies, sporting event)",
        "Social event (e.g., visit friends, family, co-workers)",
        "Other social/leisure",
        "Went to another residence (e.g., someone else's home, second home)",
        "Religious/civic/volunteer activity",
        "Volunteering",
        "Social, leisure, religious, entertainment activity"
      ) ~ "Meals / social / recreation",
      dest_purpose %in% c(
        "Pick someone up",
        "Drop someone off",
        "BOTH pick up AND drop off",
        "Accompany someone only (e.g., go along for the ride)",
        "Dropped off, picked up, or accompanied another person"
      ) ~ "Escort / care",
      dest_purpose %in% c(
        "Changed or transferred mode (e.g., change from ferry to bus)",
        "Went to temporary lodging (e.g., hotel, vacation rental)",
        "Other activity only (e.g., attend meeting, pick-up or drop-off item)",
        "Other reason",
        "Not imputable"
      ) ~ "Other / change mode / overnight",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

purp_x_inc_group_levels <- income_group_levels

purp_x_inc_totals <- purp_x_inc_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

purp_x_inc_estimates <- lapply(purp_x_inc_group_levels, function(group_value) {
  purp_x_inc_group <- purp_x_inc_pre %>%
    filter(group_label == group_value)

  if (nrow(purp_x_inc_group) == 0) {
    return(NULL)
  }

  lapply(levels(purp_shr_fmt$category), function(category_value) {
    purp_x_inc_analysis <- purp_x_inc_group %>%
      mutate(indicator = as.numeric(category == category_value))

    purp_x_inc_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = purp_x_inc_analysis,
      nest = TRUE
    )

    purp_x_inc_estimate <- suppressWarnings(survey::svymean(~indicator, purp_x_inc_design, na.rm = TRUE))
    purp_x_inc_ci <- suppressWarnings(stats::confint(purp_x_inc_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(purp_x_inc_estimate)[1]),
      se = unname(survey::SE(purp_x_inc_estimate)[1]),
      ci_low = pmax(0, unname(purp_x_inc_ci[1])),
      ci_high = pmin(1, unname(purp_x_inc_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

purp_x_inc_fmt <- purp_x_inc_estimates %>%
  left_join(purp_x_inc_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(
      group_label,
      levels = income_group_levels[income_group_levels %in% group_label]
    ),
    group_display = factor(
      dplyr::recode(as.character(group_label), !!!income_group_short_map),
      levels = unname(income_group_short_map[income_group_levels[income_group_levels %in% group_label]])
    ),
    group_display_full = dplyr::recode(as.character(group_label), !!!income_group_plot_label_map),
    category = factor(category, levels = levels(purp_shr_fmt$category))
  ) %>%
  arrange(group_label, category)

purp_x_inc_fill_values <- setNames(
  rep(unname(stacked_palette), length.out = dplyr::n_distinct(purp_x_inc_fmt$category)),
  levels(purp_x_inc_fmt$category)
)
purp_x_inc_text_values <- setNames(
  contrast_text_for_fill(unname(purp_x_inc_fill_values)),
  names(purp_x_inc_fill_values)
)

purp_x_inc_fmt <- purp_x_inc_fmt %>%
  mutate(
    hover_text = paste0(
      "Household income: ", group_display_full,
      "<br>Purpose: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )

purp_x_inc_plot <- ggplot(
  purp_x_inc_fmt,
  aes(
    x = group_display,
    y = estimate,
    fill = category,
    text = hover_text
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = purp_x_inc_fill_values) +
  scale_color_manual(values = purp_x_inc_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted purpose share within each household income group") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

purp_x_inc_fig <- ggplotly(
  purp_x_inc_plot,
  tooltip = "text",
  height = 430,
  dynamicTicks = FALSE
)

for (i in seq_along(purp_x_inc_fig$x$data)) {
  purp_x_inc_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  purp_x_inc_fig$x$data[[i]]$showlegend <- identical(purp_x_inc_fig$x$data[[i]]$type, "bar")
  if (!is.null(purp_x_inc_fig$x$data[[i]]$name)) {
    purp_x_inc_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", purp_x_inc_fig$x$data[[i]]$name)
    purp_x_inc_fig$x$data[[i]]$legendgroup <- purp_x_inc_fig$x$data[[i]]$name
  }
  if (!identical(purp_x_inc_fig$x$data[[i]]$type, "bar")) {
    purp_x_inc_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

purp_x_inc_fig %>%
  layout(
    margin = list(l = 40, r = 20, t = 40, b = 140),
    xaxis = list(
      automargin = TRUE,
      tickangle = -30,
      tickfont = list(size = 10)
    )
  ) %>%
  config(displayModeBar = FALSE)
Figure 12: Weighted share of trips by destination purpose across household income groups. Shares use trip_weight; * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Code
purp_x_inc_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Trip Purpose by Household Income (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trip Purpose by Household Income (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
Under $25,000
Home 38.7% 2.8% 33.2% to 44.3% 7.2% 390,241 438
Meals / social / recreation 16.8% 2.3% 12.3% to 21.4% 13.8% 169,518 246
Shopping / errands 27.3% 2.4% 22.5% to 32.0% 9.0% 274,515 317
Work 6.1% 1.1% 4.0% to 8.2% 17.4% 61,587 122
Escort / care 6.0% 1.4% 3.3% to 8.6% 22.7% 59,998 63
School 4.6% 1.2% 2.3% to 7.0% 25.3% 46,790 38
Other / change mode / overnight 0.4%* 0.2% 0.1% to 0.8% 41.7% 4,499 15
$25,000-$49,999
Home 34.5% 2.1% 30.4% to 38.6% 6.1% 485,207 947
Meals / social / recreation 19.2% 1.8% 15.7% to 22.6% 9.2% 269,611 571
Shopping / errands 17.1% 1.6% 13.9% to 20.2% 9.4% 240,074 596
Work 15.2% 1.6% 12.0% to 18.3% 10.6% 213,144 374
Escort / care 10.6% 1.4% 7.8% to 13.5% 13.6% 149,476 175
School 2.6% 0.5% 1.6% to 3.6% 19.9% 36,804 67
Other / change mode / overnight 0.8%* 0.3% 0.3% to 1.4% 32.9% 11,642 30
$50,000-$74,999
Home 35.4% 2.0% 31.4% to 39.5% 5.8% 791,711 1,092
Meals / social / recreation 20.0% 1.7% 16.7% to 23.3% 8.5% 446,435 685
Shopping / errands 23.0% 1.9% 19.3% to 26.7% 8.3% 513,603 660
Work 12.5% 1.2% 10.1% to 14.9% 9.7% 279,048 563
Escort / care 5.8% 1.0% 3.9% to 7.8% 17.2% 130,134 171
School 2.8% 0.6% 1.5% to 4.0% 22.6% 61,967 59
Other / change mode / overnight 0.5% 0.1% 0.2% to 0.7% 25.5% 11,124 41
$75,000-$99,999
Home 38.7% 2.2% 34.4% to 42.9% 5.6% 600,919 1,109
Meals / social / recreation 15.0% 1.4% 12.3% to 17.7% 9.1% 233,325 640
Shopping / errands 17.9% 1.6% 14.7% to 21.1% 9.1% 278,224 595
Work 16.1% 1.5% 13.1% to 19.1% 9.5% 249,687 500
Escort / care 5.7% 1.1% 3.6% to 7.8% 18.7% 88,864 115
School 5.4% 1.1% 3.2% to 7.5% 20.2% 83,200 58
Other / change mode / overnight 1.3%** 0.6% 0.0% to 2.5% 50.1% 19,842 37
$100,000-$199,999
Home 38.9% 1.2% 36.5% to 41.3% 3.2% 1,921,629 3,113
Meals / social / recreation 16.4% 0.9% 14.7% to 18.1% 5.4% 811,949 1,844
Shopping / errands 12.8% 0.8% 11.3% to 14.3% 6.0% 632,417 1,587
Work 15.4% 1.0% 13.5% to 17.2% 6.2% 759,236 1,312
Escort / care 9.9% 0.8% 8.4% to 11.4% 7.8% 490,073 576
School 5.3% 0.6% 4.1% to 6.4% 11.1% 260,555 265
Other / change mode / overnight 1.4% 0.2% 0.9% to 1.9% 17.8% 68,138 142
$200,000 or more
Home 38.2% 1.3% 35.6% to 40.8% 3.4% 1,722,943 1,923
Meals / social / recreation 16.1% 0.9% 14.3% to 18.0% 5.9% 726,387 990
Shopping / errands 12.9% 0.8% 11.3% to 14.6% 6.4% 583,574 780
Work 12.7% 0.9% 11.0% to 14.4% 6.7% 571,476 874
Escort / care 10.4% 0.9% 8.6% to 12.2% 8.8% 468,118 403
School 7.1% 0.7% 5.7% to 8.6% 10.2% 321,158 225
Other / change mode / overnight 2.5% 0.5% 1.6% to 3.4% 18.8% 112,729 78
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 15: Trip purpose by household income (weighted)

Trip Purpose by Time of Day

Time of day produces some of the clearest differences in trip purpose. Morning travel is more concentrated in work and school trips, while midday and evening feature larger shares of shopping, social or recreational travel, and return-home trips, reflecting the daily rhythm of regional activity patterns.

Code
purp_x_tod_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = case_when(
      is.na(depart_time_hour) ~ NA_character_,
      depart_time_hour < 6 ~ "Early AM (12-5)",
      depart_time_hour < 10 ~ "AM peak (6-9)",
      depart_time_hour < 16 ~ "Midday (10-3)",
      depart_time_hour < 19 ~ "PM peak (4-6)",
      TRUE ~ "Evening (7-11)"
    ),
    category = case_when(
      dest_purpose == "Went home" ~ "Home",
      dest_purpose %in% c(
        "Went to primary workplace",
        "Went to work-related activity (e.g., meeting, delivery, worksite)",
        "Went to other work-related activity"
      ) ~ "Work",
      dest_purpose %in% c(
        "Attend K-12 school",
        "Attend daycare or preschool",
        "Attend college/university",
        "Attend other education-related activity (e.g., field trip)",
        "Attend other type of class (e.g., cooking class)",
        "Attend vocational education class",
        "Attended school/class"
      ) ~ "School",
      dest_purpose %in% c(
        "Grocery shopping",
        "Other shopping (e.g., mall, pet store)",
        "Personal business (e.g., bank, post office)",
        "Medical appointment (e.g., doctor, dentist)",
        "Got gas",
        "Other appointment/errands",
        "Appointment, shopping, or errands (e.g., gas)"
      ) ~ "Shopping / errands",
      dest_purpose %in% c(
        "Went to restaurant to eat/get take-out",
        "Exercise or recreation (e.g., gym, jog, bike, walk dog)",
        "Recreational event (e.g., movies, sporting event)",
        "Social event (e.g., visit friends, family, co-workers)",
        "Other social/leisure",
        "Went to another residence (e.g., someone else's home, second home)",
        "Religious/civic/volunteer activity",
        "Volunteering",
        "Social, leisure, religious, entertainment activity"
      ) ~ "Meals / social / recreation",
      dest_purpose %in% c(
        "Pick someone up",
        "Drop someone off",
        "BOTH pick up AND drop off",
        "Accompany someone only (e.g., go along for the ride)",
        "Dropped off, picked up, or accompanied another person"
      ) ~ "Escort / care",
      dest_purpose %in% c(
        "Changed or transferred mode (e.g., change from ferry to bus)",
        "Went to temporary lodging (e.g., hotel, vacation rental)",
        "Other activity only (e.g., attend meeting, pick-up or drop-off item)",
        "Other reason",
        "Not imputable"
      ) ~ "Other / change mode / overnight",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

purp_x_tod_group_levels <- c(
  "Early AM (12-5)",
  "AM peak (6-9)",
  "Midday (10-3)",
  "PM peak (4-6)",
  "Evening (7-11)"
)

purp_x_tod_totals <- purp_x_tod_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

purp_x_tod_estimates <- lapply(purp_x_tod_group_levels, function(group_value) {
  purp_x_tod_group <- purp_x_tod_pre %>%
    filter(group_label == group_value)

  if (nrow(purp_x_tod_group) == 0) {
    return(NULL)
  }

  lapply(levels(purp_shr_fmt$category), function(category_value) {
    purp_x_tod_analysis <- purp_x_tod_group %>%
      mutate(indicator = as.numeric(category == category_value))

    purp_x_tod_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = purp_x_tod_analysis,
      nest = TRUE
    )

    purp_x_tod_estimate <- suppressWarnings(survey::svymean(~indicator, purp_x_tod_design, na.rm = TRUE))
    purp_x_tod_ci <- suppressWarnings(stats::confint(purp_x_tod_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(purp_x_tod_estimate)[1]),
      se = unname(survey::SE(purp_x_tod_estimate)[1]),
      ci_low = pmax(0, unname(purp_x_tod_ci[1])),
      ci_high = pmin(1, unname(purp_x_tod_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

purp_x_tod_fmt <- purp_x_tod_estimates %>%
  left_join(purp_x_tod_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(group_label, levels = c(
      "Early AM (12-5)",
      "AM peak (6-9)",
      "Midday (10-3)",
      "PM peak (4-6)",
      "Evening (7-11)"
    )),
    category = factor(category, levels = levels(purp_shr_fmt$category))
  ) %>%
  arrange(group_label, category)

purp_x_tod_fill_values <- setNames(
  rep(unname(stacked_palette), length.out = dplyr::n_distinct(purp_x_tod_fmt$category)),
  levels(purp_x_tod_fmt$category)
)
purp_x_tod_text_values <- setNames(
  contrast_text_for_fill(unname(purp_x_tod_fill_values)),
  names(purp_x_tod_fill_values)
)

purp_x_tod_plot <- ggplot(
  purp_x_tod_fmt,
  aes(
    x = group_label,
    y = estimate,
    fill = category,
    text = paste0(
      "Time of day: ", group_label,
      "<br>Purpose: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = purp_x_tod_fill_values) +
  scale_color_manual(values = purp_x_tod_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted purpose share within each time-of-day bin") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

purp_x_tod_fig <- ggplotly(purp_x_tod_plot, tooltip = "text", height = 430)
for (i in seq_along(purp_x_tod_fig$x$data)) {
  purp_x_tod_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  purp_x_tod_fig$x$data[[i]]$showlegend <- identical(purp_x_tod_fig$x$data[[i]]$type, "bar")
  if (!is.null(purp_x_tod_fig$x$data[[i]]$name)) {
    purp_x_tod_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", purp_x_tod_fig$x$data[[i]]$name)
    purp_x_tod_fig$x$data[[i]]$legendgroup <- purp_x_tod_fig$x$data[[i]]$name
  }
  if (!identical(purp_x_tod_fig$x$data[[i]]$type, "bar")) {
    purp_x_tod_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

purp_x_tod_fig %>%
  layout(margin = list(l = 40, r = 20, t = 40, b = 70)) %>%
  config(displayModeBar = FALSE)
Figure 13: Weighted share of trips by destination purpose across time-of-day periods based on trip departure hour. Shares use trip_weight; * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Code
purp_x_tod_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Trip Purpose by Time of Day (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trip Purpose by Time of Day (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
Early AM (12-5)
Home 15.8% 3.3% 9.3% to 22.4% 21.1% 77,175 98
Meals / social / recreation 14.6% 3.1% 8.6% to 20.6% 21.0% 71,060 142
Shopping / errands 4.4%* 1.8% 0.9% to 8.0% 41.1% 21,628 37
Work 58.0% 4.3% 49.5% to 66.5% 7.5% 282,645 382
Escort / care 1.9%** 1.0% 0.0% to 3.9% 51.0% 9,422 20
School 3.3%** 1.7% 0.0% to 6.5% 50.7% 15,997 9
Other / change mode / overnight 1.9%** 1.4% 0.0% to 4.6% 75.0% 9,092 8
AM peak (6-9)
Home 14.6% 1.1% 12.5% to 16.7% 7.2% 541,087 807
Meals / social / recreation 10.9% 0.9% 9.1% to 12.7% 8.4% 404,401 795
Shopping / errands 9.7% 0.8% 8.1% to 11.3% 8.4% 359,072 714
Work 33.5% 1.4% 30.8% to 36.3% 4.2% 1,242,471 1,923
Escort / care 12.6% 1.1% 10.5% to 14.6% 8.4% 465,461 512
School 18.0% 1.2% 15.7% to 20.3% 6.6% 667,356 567
Other / change mode / overnight 0.7% 0.2% 0.3% to 1.1% 28.1% 25,713 44
Midday (10-3)
Home 37.1% 1.1% 35.0% to 39.2% 2.9% 2,527,580 3,670
Meals / social / recreation 17.1% 0.8% 15.6% to 18.7% 4.7% 1,167,242 2,264
Shopping / errands 24.2% 0.9% 22.4% to 26.1% 3.9% 1,652,431 2,999
Work 8.7% 0.6% 7.6% to 9.9% 6.6% 595,740 1,264
Escort / care 9.0% 0.7% 7.7% to 10.4% 7.6% 616,473 634
School 2.4% 0.4% 1.7% to 3.1% 15.2% 161,957 151
Other / change mode / overnight 1.4% 0.3% 0.9% to 1.9% 18.6% 93,071 156
PM peak (4-6)
Home 48.7% 1.4% 45.9% to 51.5% 3.0% 2,028,227 2,891
Meals / social / recreation 22.8% 1.2% 20.4% to 25.2% 5.3% 948,970 1,565
Shopping / errands 13.6% 1.0% 11.7% to 15.5% 7.3% 566,520 852
Work 3.9% 0.6% 2.8% to 5.0% 14.6% 161,093 272
Escort / care 7.3% 0.8% 5.8% to 8.9% 10.9% 305,914 334
School 1.7% 0.4% 0.9% to 2.5% 24.2% 71,343 53
Other / change mode / overnight 2.0% 0.4% 1.1% to 2.8% 22.0% 82,151 109
Evening (7-11)
Home 67.0% 1.9% 63.3% to 70.7% 2.8% 1,291,846 1,754
Meals / social / recreation 14.4% 1.3% 11.9% to 17.0% 9.0% 278,223 535
Shopping / errands 9.6% 1.2% 7.3% to 11.8% 12.0% 184,276 303
Work 2.5% 0.6% 1.4% to 3.6% 23.1% 47,962 92
Escort / care 4.7% 1.0% 2.9% to 6.6% 20.1% 91,244 109
School 0.3%** 0.2% 0.0% to 0.7% 55.2% 6,293 7
Other / change mode / overnight 1.5% 0.4% 0.8% to 2.2% 24.5% 29,261 47
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 16: Trip purpose by time of day (weighted)

Trip Purpose by Distance

Trip purpose also varies with trip length. Shorter trips are more likely to involve errands, meals, social activities, and nearby home-based travel, while longer trips are more work-oriented, linking trip distance to distinct types of daily activity.

Code
purp_x_dist_pre <- as.data.frame(hts$trip) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  mutate(
    group_label = case_when(
      is.na(distance_miles) ~ NA_character_,
      distance_miles < 1 ~ "<1 mile",
      distance_miles < 3 ~ "1-3 miles",
      distance_miles < 5 ~ "3-5 miles",
      distance_miles < 10 ~ "5-10 miles",
      TRUE ~ "10+ miles"
    ),
    category = case_when(
      dest_purpose == "Went home" ~ "Home",
      dest_purpose %in% c(
        "Went to primary workplace",
        "Went to work-related activity (e.g., meeting, delivery, worksite)",
        "Went to other work-related activity"
      ) ~ "Work",
      dest_purpose %in% c(
        "Attend K-12 school",
        "Attend daycare or preschool",
        "Attend college/university",
        "Attend other education-related activity (e.g., field trip)",
        "Attend other type of class (e.g., cooking class)",
        "Attend vocational education class",
        "Attended school/class"
      ) ~ "School",
      dest_purpose %in% c(
        "Grocery shopping",
        "Other shopping (e.g., mall, pet store)",
        "Personal business (e.g., bank, post office)",
        "Medical appointment (e.g., doctor, dentist)",
        "Got gas",
        "Other appointment/errands",
        "Appointment, shopping, or errands (e.g., gas)"
      ) ~ "Shopping / errands",
      dest_purpose %in% c(
        "Went to restaurant to eat/get take-out",
        "Exercise or recreation (e.g., gym, jog, bike, walk dog)",
        "Recreational event (e.g., movies, sporting event)",
        "Social event (e.g., visit friends, family, co-workers)",
        "Other social/leisure",
        "Went to another residence (e.g., someone else's home, second home)",
        "Religious/civic/volunteer activity",
        "Volunteering",
        "Social, leisure, religious, entertainment activity"
      ) ~ "Meals / social / recreation",
      dest_purpose %in% c(
        "Pick someone up",
        "Drop someone off",
        "BOTH pick up AND drop off",
        "Accompany someone only (e.g., go along for the ride)",
        "Dropped off, picked up, or accompanied another person"
      ) ~ "Escort / care",
      dest_purpose %in% c(
        "Changed or transferred mode (e.g., change from ferry to bus)",
        "Went to temporary lodging (e.g., hotel, vacation rental)",
        "Other activity only (e.g., attend meeting, pick-up or drop-off item)",
        "Other reason",
        "Not imputable"
      ) ~ "Other / change mode / overnight",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(group_label), !is.na(category), !is.na(trip_weight), trip_weight > 0, !is.na(sample_segment))

purp_x_dist_group_levels <- c(
  "<1 mile",
  "1-3 miles",
  "3-5 miles",
  "5-10 miles",
  "10+ miles"
)

purp_x_dist_totals <- purp_x_dist_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(trip_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

purp_x_dist_estimates <- lapply(purp_x_dist_group_levels, function(group_value) {
  purp_x_dist_group <- purp_x_dist_pre %>%
    filter(group_label == group_value)

  if (nrow(purp_x_dist_group) == 0) {
    return(NULL)
  }

  lapply(levels(purp_shr_fmt$category), function(category_value) {
    purp_x_dist_analysis <- purp_x_dist_group %>%
      mutate(indicator = as.numeric(category == category_value))

    purp_x_dist_design <- survey::svydesign(
      ids = ~trip_id,
      strata = ~sample_segment,
      weights = ~trip_weight,
      data = purp_x_dist_analysis,
      nest = TRUE
    )

    purp_x_dist_estimate <- suppressWarnings(survey::svymean(~indicator, purp_x_dist_design, na.rm = TRUE))
    purp_x_dist_ci <- suppressWarnings(stats::confint(purp_x_dist_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(purp_x_dist_estimate)[1]),
      se = unname(survey::SE(purp_x_dist_estimate)[1]),
      ci_low = pmax(0, unname(purp_x_dist_ci[1])),
      ci_high = pmin(1, unname(purp_x_dist_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

purp_x_dist_fmt <- purp_x_dist_estimates %>%
  left_join(purp_x_dist_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(group_label, levels = c(
      "<1 mile",
      "1-3 miles",
      "3-5 miles",
      "5-10 miles",
      "10+ miles"
    )),
    category = factor(category, levels = levels(purp_shr_fmt$category))
  ) %>%
  arrange(group_label, category)

purp_x_dist_fill_values <- setNames(
  rep(unname(stacked_palette), length.out = dplyr::n_distinct(purp_x_dist_fmt$category)),
  levels(purp_x_dist_fmt$category)
)
purp_x_dist_text_values <- setNames(
  contrast_text_for_fill(unname(purp_x_dist_fill_values)),
  names(purp_x_dist_fill_values)
)

purp_x_dist_plot <- ggplot(
  purp_x_dist_fmt,
  aes(
    x = group_label,
    y = estimate,
    fill = category,
    text = paste0(
      "Trip distance: ", group_label,
      "<br>Purpose: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted trips: ", comma(round(weighted_total)),
      "<br>Unweighted trips: ", comma(unweighted_n)
    )
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = purp_x_dist_fill_values) +
  scale_color_manual(values = purp_x_dist_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted purpose share within each trip-distance bin") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

purp_x_dist_fig <- ggplotly(purp_x_dist_plot, tooltip = "text", height = 430)
for (i in seq_along(purp_x_dist_fig$x$data)) {
  purp_x_dist_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  purp_x_dist_fig$x$data[[i]]$showlegend <- identical(purp_x_dist_fig$x$data[[i]]$type, "bar")
  if (!is.null(purp_x_dist_fig$x$data[[i]]$name)) {
    purp_x_dist_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", purp_x_dist_fig$x$data[[i]]$name)
    purp_x_dist_fig$x$data[[i]]$legendgroup <- purp_x_dist_fig$x$data[[i]]$name
  }
  if (!identical(purp_x_dist_fig$x$data[[i]]$type, "bar")) {
    purp_x_dist_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

purp_x_dist_fig %>%
  layout(margin = list(l = 40, r = 20, t = 40, b = 70)) %>%
  config(displayModeBar = FALSE)
Figure 14: Weighted share of trips by destination purpose across trip-distance bands. Shares use trip_weight.
Code
purp_x_dist_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted trips` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Trip Purpose by Distance (Weighted)") %>%
  tab_source_note(md("Shares use `trip_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Trip Purpose by Distance (Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted trips
<1 mile
Home 36.5% 1.4% 33.7% to 39.2% 3.9% 1,258,714 2,114
Meals / social / recreation 21.6% 1.2% 19.3% to 23.9% 5.4% 744,571 1,552
Shopping / errands 17.3% 1.1% 15.2% to 19.4% 6.3% 596,920 1,125
Work 8.4% 0.7% 7.1% to 9.7% 7.9% 288,739 868
Escort / care 8.1% 0.9% 6.3% to 9.9% 11.3% 278,086 305
School 6.6% 0.8% 5.0% to 8.3% 12.7% 227,851 206
Other / change mode / overnight 1.6% 0.3% 1.0% to 2.2% 19.5% 54,737 77
1-3 miles
Home 40.1% 1.4% 37.5% to 42.8% 3.4% 1,834,380 2,388
Meals / social / recreation 16.6% 1.0% 14.6% to 18.5% 6.0% 758,130 1,328
Shopping / errands 16.3% 1.0% 14.4% to 18.2% 5.9% 744,581 1,321
Work 7.6% 0.6% 6.3% to 8.8% 8.6% 346,287 629
Escort / care 10.9% 0.9% 9.1% to 12.8% 8.7% 498,911 509
School 7.3% 0.7% 5.8% to 8.7% 10.2% 332,209 270
Other / change mode / overnight 1.3% 0.4% 0.5% to 2.0% 29.1% 57,330 86
3-5 miles
Home 38.5% 1.8% 35.1% to 42.0% 4.6% 1,036,452 1,408
Meals / social / recreation 15.8% 1.3% 13.3% to 18.4% 8.1% 425,695 756
Shopping / errands 17.2% 1.3% 14.6% to 19.8% 7.8% 461,840 841
Work 12.2% 1.2% 9.9% to 14.5% 9.7% 328,123 446
Escort / care 9.3% 1.1% 7.2% to 11.5% 11.9% 251,201 268
School 6.4% 0.9% 4.6% to 8.2% 14.4% 171,945 121
Other / change mode / overnight 0.5% 0.1% 0.3% to 0.7% 22.8% 13,654 47
5-10 miles
Home 34.7% 1.7% 31.5% to 38.0% 4.8% 1,085,008 1,596
Meals / social / recreation 15.4% 1.3% 12.9% to 17.8% 8.1% 479,830 799
Shopping / errands 17.3% 1.3% 14.7% to 19.9% 7.6% 541,238 853
Work 16.9% 1.3% 14.4% to 19.4% 7.6% 527,592 777
Escort / care 10.8% 1.0% 8.8% to 12.8% 9.6% 337,139 307
School 3.6% 0.6% 2.5% to 4.8% 15.7% 113,876 118
Other / change mode / overnight 1.3% 0.4% 0.6% to 2.1% 29.2% 40,815 61
10+ miles
Home 38.3% 1.6% 35.2% to 41.4% 4.1% 1,251,313 1,713
Meals / social / recreation 14.1% 1.1% 12.1% to 16.2% 7.5% 461,672 866
Shopping / errands 13.5% 1.1% 11.3% to 15.6% 8.2% 439,348 765
Work 25.7% 1.4% 22.9% to 28.5% 5.5% 839,171 1,213
Escort / care 3.8% 0.5% 2.7% to 4.8% 14.4% 123,177 220
School 2.4% 0.5% 1.4% to 3.3% 20.9% 77,065 72
Other / change mode / overnight 2.2% 0.5% 1.3% to 3.1% 20.9% 72,753 93
Shares use trip_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 17: Trip purpose by distance (weighted)

0.6 Trip Substitutions

Work Arrangement

Most workers are still primarily working outside the home. Hybrid and fully remote arrangements remain important parts of the region’s work pattern, but they supplement rather than replace the dominant in-person work model.

Code
wfh_loc_pre <- as.data.frame(hts$person) %>%
  left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(age >= 16, paid_work %in% c("Yes", "On leave (e.g., medical, parental)"), !is.na(person_weight), person_weight > 0) %>%
  transmute(
    person_id,
    sample_segment,
    person_weight,
    category = case_when(
      work_from_home == "No" ~ "Only outside the home",
      work_from_home == "Yes, some of the time (less than 50% of the time)" ~ "Hybrid, less than half remote",
      work_from_home == "Yes, most of the time (at least than 50% but less than 100% of the time)" ~ "Hybrid, more than half remote",
      work_from_home == "Yes, all of the time (100% of the time)" ~ "Only from home",
      TRUE ~ NA_character_
    )
  )
wfh_loc_sum <- wfh_loc_pre %>%
  filter(!is.na(category)) %>%
  as_survey_design(ids = person_id, weights = person_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(estimate = survey_prop(vartype = c("se", "ci"), proportion = TRUE), .groups = "drop") %>%
  as.data.frame()

wfh_loc_fmt <- wfh_loc_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    short_label = case_when(
      category == "Only outside the home" ~ "Out",
      category == "Hybrid, less than half remote" ~ "H<50",
      category == "Hybrid, more than half remote" ~ "H>50",
      category == "Only from home" ~ "Home",
      TRUE ~ category
    )
  ) %>%
  arrange(desc(estimate))

wfh_loc_colors <- unname(stacked_palette[seq_len(4)])

plot_ly(
  wfh_loc_fmt,
  labels = ~category,
  values = ~estimate,
  type = "pie",
  hole = 0.58,
  sort = FALSE,
  text = ~ paste0(short_label, "<br>", percent(estimate, accuracy = 1), rse_flag),
  textinfo = "text",
  hovertext = ~ paste0(
    category,
    "<br>Share: ", percent(estimate, accuracy = 0.1),
    "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)
  ),
  hoverinfo = "text",
  textposition = "inside",
  textfont = list(family = "Inter", size = 13, color = contrast_text_for_fill(wfh_loc_colors)),
  marker = list(
    colors = wfh_loc_colors,
    line = list(color = divider_color, width = 2)
  ),
  hovertemplate = "%{hovertext}<extra></extra>"
) %>%
  layout(
    height = 360,
    margin = list(l = 10, r = 10, t = 10, b = 28),
    legend = list(orientation = "h", x = 0.02, y = -0.04, font = list(family = "Inter", size = 12)),
    paper_bgcolor = "rgba(0,0,0,0)",
    plot_bgcolor = "rgba(0,0,0,0)"
  ) %>%
  config(displayModeBar = FALSE)
Figure 15: Work arrangement among workers age 16 or older. Estimates use person_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
wfh_loc_fmt %>%
  transmute(
    Arrangement = category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1)
  ) %>%
  gt() %>%
  tab_header(title = "Work Location Pattern") %>%
  tab_source_note(md("Estimates use `person_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Work Location Pattern
Arrangement Share 95% confidence interval (CI) Relative standard error (RSE)
Only outside the home 57.1% 53.6% to 60.4% 3.0%
Hybrid, less than half remote 22.7% 19.9% to 25.7% 6.6%
Hybrid, more than half remote 10.1% 8.5% to 12.1% 9.1%
Only from home 10.1% 8.3% to 12.3% 9.9%
Estimates use person_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 18: Work location pattern

Telework Time

Telework-time results reinforce that many worker-days still involve no telework at all. When telework does occur, it is more often recorded as a substantial block of time than as a brief period, suggesting that remote work tends to be organized as a meaningful share of the day rather than a marginal activity.

Code
tele_tm_pre <- as.data.frame(hts$day) %>%
  left_join(as.data.frame(hts$person) %>% select(person_id, age, paid_work), by = "person_id") %>%
  left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(age >= 16, paid_work %in% c("Yes", "On leave (e.g., medical, parental)"), !is.na(day_weight), day_weight > 0) %>%
  transmute(
    day_id,
    sample_segment,
    day_weight,
    category = case_when(
      telework_time == "0 minutes" ~ "Did not telework",
      telework_time %in% c("30 minutes", "1 hour") ~ "0-1 hours",
      telework_time %in% c("1 hour 30 minutes", "2 hours", "2 hours 30 minutes", "3 hours", "3 hours 30 minutes", "4 hours", "4 hours 30 minutes", "5 hours", "5 hours 30 minutes", "6 hours") ~ "1-6 hours",
      telework_time %in% c("6 hours 30 minutes", "7 hours", "7 hours 30 minutes", "8 hours", "8 hours 30 minutes", "9 hours", "9 hours 30 minutes", "10+ hours") ~ "6+ hours",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(category))

tele_tm_sum <- tele_tm_pre %>%
  as_survey_design(ids = day_id, weights = day_weight, strata = sample_segment, nest = TRUE) %>%
  group_by(category) %>%
  summarise(estimate = survey_prop(vartype = c("se", "ci"), proportion = TRUE), .groups = "drop") %>%
  as.data.frame()

tele_tm_fmt <- tele_tm_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    short_label = case_when(
      category == "Did not telework" ~ "None",
      category == "0-1 hours" ~ "0-1h",
      category == "1-6 hours" ~ "1-6h",
      category == "6+ hours" ~ "6+h",
      TRUE ~ category
    )
  ) %>%
  arrange(desc(estimate))

tele_tm_colors <- unname(stacked_palette[seq_len(4)])

plot_ly(
  tele_tm_fmt,
  labels = ~category,
  values = ~estimate,
  type = "pie",
  hole = 0.58,
  sort = FALSE,
  text = ~ paste0(short_label, "<br>", percent(estimate, accuracy = 1), rse_flag),
  textinfo = "text",
  hovertext = ~ paste0(
    category,
    "<br>Share: ", percent(estimate, accuracy = 0.1),
    "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)
  ),
  hoverinfo = "text",
  textposition = "inside",
  textfont = list(family = "Inter", size = 13, color = contrast_text_for_fill(tele_tm_colors)),
  marker = list(
    colors = tele_tm_colors,
    line = list(color = divider_color, width = 2)
  ),
  hovertemplate = "%{hovertext}<extra></extra>"
) %>%
  layout(
    height = 360,
    margin = list(l = 10, r = 10, t = 10, b = 28),
    legend = list(orientation = "h", x = 0.04, y = -0.04, font = list(family = "Inter", size = 12)),
    paper_bgcolor = "rgba(0,0,0,0)",
    plot_bgcolor = "rgba(0,0,0,0)"
  ) %>%
  config(displayModeBar = FALSE)
Figure 16: Telework time on worker-days. Estimates use day_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
tele_tm_fmt %>%
  transmute(
    `Telework time` = category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1)
  ) %>%
  gt() %>%
  tab_header(title = "Telework Time") %>%
  tab_source_note(md("Estimates use `day_weight`. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Telework Time
Telework time Share 95% confidence interval (CI) Relative standard error (RSE)
Did not telework 60.1% 56.9% to 63.2% 2.7%
6+ hours 25.1% 22.5% to 27.8% 5.4%
1-6 hours 10.2% 8.4% to 12.3% 9.7%
0-1 hours 4.6% 3.4% to 6.2% 15.2%
Estimates use day_weight. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 19: Telework time on worker-days

Delivery Rates

Delivery activity is present across the sample, but it is concentrated most strongly in package delivery. Other delivery types are much less common, indicating that delivery substitution affects daily travel primarily through packages rather than through groceries, meals, or services.

Code
deliv_pre <- as.data.frame(hts$day) %>%
  left_join(as.data.frame(hts$hh) %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(!is.na(day_weight), day_weight > 0) %>%
  transmute(
    day_id,
    sample_segment,
    day_weight,
    packages = deliver_package,
    services = deliver_work,
    office = deliver_office,
    groceries = deliver_grocery,
    food = deliver_food,
    elsewhere = deliver_elsewhere,
    other = deliver_other,
    none = deliver_none
  ) %>%
  pivot_longer(cols = c(packages, services, office, groceries, food, other, none), names_to = "category", values_to = "selected") %>%
  mutate(category = case_when(
    category == "packages" ~ "Package delivery",
    category == "services" ~ "Service delivery (e.g., laundry, cleaning)",
    category == "office" ~ "Office delivery",
    category == "groceries" ~ "Grocery delivery",
    category == "food" ~ "Food delivery",
    category == "elsewhere" ~ "Other delivery",
    category == "other" ~ "Other delivery",
    category == "none" ~ "No delivery"
  )) %>%
  mutate(selected_clean = case_when(
    selected == "Yes" | selected == "Selected" ~ "Yes",
    selected == "No" | selected == "Not selected" ~ "No",
    selected == "Missing Response" ~ "Missing",
    TRUE ~ NA_character_
  ))

deliv_sum <- deliv_pre %>%
  as_survey_design(ids = day_id, weights = day_weight, strata = sample_segment) %>%
  filter(selected_clean %in% c("Yes", "No")) %>%
  group_by(category, selected_clean) %>%
  summarise(estimate = survey_prop(vartype = c("se", "ci"), na.rm = TRUE), .groups = "drop") %>%
  as.data.frame() %>%
  filter(selected_clean == "Yes") %>%
  filter(category != "No delivery")

deliv_fmt <- deliv_sum %>%
  mutate(
    se = estimate_se,
    ci_low = estimate_low,
    ci_high = estimate_upp,
    rse = se / estimate,
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ "")
  ) %>%
  arrange(desc(estimate)) %>%
  mutate(
    label = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    category = factor(category, levels = rev(category))
  )

deliv_x_max <- max(deliv_fmt$ci_high, na.rm = TRUE)
deliv_label_buffer <- deliv_x_max * 0.08
deliv_fmt <- deliv_fmt %>%
  mutate(label_x = ci_high + deliv_label_buffer)

plot_delivery <- ggplot(
  deliv_fmt,
  aes(
    x = estimate,
    y = category,
    text = paste0(
      "Delivery type: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)
    )
  )
) +
  geom_col(fill = single_bar_color, width = 0.68) +
  geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.14, color = error_bar_color) +
  geom_text(aes(x = label_x, label = label), hjust = 0, color = single_bar_color, size = 4.1, family = "Inter", fontface = "bold") +
  scale_x_continuous(
    labels = label_percent(accuracy = 1),
    limits = c(0, max(deliv_fmt$label_x, na.rm = TRUE) + deliv_label_buffer * 2)
  ) +
  labs(x = NULL, y = NULL) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.y = element_blank(),
    panel.grid.minor = element_blank()
  )

delivery_plotly <- ggplotly(plot_delivery, tooltip = "text", height = 430)
delivery_plotly$x$data[[1]]$hovertemplate <- "%{text}<extra></extra>"
delivery_plotly$x$data[[2]]$hoverinfo <- "skip"
delivery_plotly$x$data[[3]]$hoverinfo <- "skip"

delivery_plotly %>%
  layout(margin = list(l = 120, r = 40, t = 30, b = 20)) %>%
  config(displayModeBar = FALSE)
Figure 17: Delivery rates on eligible diary-days by delivery type. Delivery is multiple response, so percentages can sum to more than 100%. Estimates use day_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
deliv_fmt %>%
  transmute(
    `Delivery type` = category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1)
  ) %>%
  gt() %>%
  tab_header(title = "Delivery Rates") %>%
  tab_source_note(md("Estimates use `day_weight` and questionnaire eligibility for the delivery questions. Delivery is multiple response, so percentages can sum to more than 100%. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Delivery Rates
Delivery type Share 95% confidence interval (CI) Relative standard error (RSE)
Package delivery 33.9% 30.9% to 37.0% 4.6%
Food delivery 3.3% 2.2% to 4.8% 19.3%
Service delivery (e.g., laundry, cleaning) 2.9% 2.2% to 3.8% 13.6%
Grocery delivery 2.8% 1.8% to 4.3% 22.3%
Other delivery 0.9%* 0.5% to 1.7% 33.4%
Office delivery 0.2%* 0.1% to 0.6% 43.3%
Estimates use day_weight and questionnaire eligibility for the delivery questions. Delivery is multiple response, so percentages can sum to more than 100%. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 20: Delivery rates

Telework Time by Day of Week

Telework patterns vary across the week, with weekdays showing a mix of no telework and partial- or full-day telework, while weekends are much more concentrated in no telework. The strongest contrast is therefore between the weekday work schedule and the more limited role telework plays on weekends.

Code
tele_x_dow_pre <- as.data.frame(hts$day) %>%
  select(day_id, household_id, person_id, travel_dow, day_weight, telework_time) %>%
  left_join(
    as.data.frame(hts$person) %>% select(person_id, age, paid_work),
    by = "person_id"
  ) %>%
  left_join(
    as.data.frame(hts$hh) %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  filter(
    age >= 16,
    paid_work %in% c("Yes", "On leave (e.g., medical, parental)"),
    !is.na(travel_dow),
    !is.na(day_weight),
    day_weight > 0,
    !is.na(sample_segment)
  ) %>%
  mutate(
    group_label = travel_dow,
    category = case_when(
      telework_time == "0 minutes" ~ "Did not telework",
      telework_time %in% c("30 minutes", "1 hour") ~ "0-1 hours",
      telework_time %in% c("1 hour 30 minutes", "2 hours", "2 hours 30 minutes", "3 hours", "3 hours 30 minutes", "4 hours", "4 hours 30 minutes", "5 hours", "5 hours 30 minutes", "6 hours") ~ "1-6 hours",
      telework_time %in% c("6 hours 30 minutes", "7 hours", "7 hours 30 minutes", "8 hours", "8 hours 30 minutes", "9 hours", "9 hours 30 minutes", "10+ hours") ~ "6+ hours",
      TRUE ~ NA_character_
    )
  ) %>%
  filter(!is.na(category))

tele_x_dow_group_levels <- c("Monday", "Tuesday", "Wednesday", "Thursday", "Friday", "Saturday", "Sunday")
tele_x_dow_category_levels <- c("Did not telework", "0-1 hours", "1-6 hours", "6+ hours")

tele_x_dow_totals <- tele_x_dow_pre %>%
  group_by(group_label, category) %>%
  summarise(weighted_total = sum(day_weight, na.rm = TRUE), unweighted_n = n(), .groups = "drop")

tele_x_dow_estimates <- lapply(tele_x_dow_group_levels, function(group_value) {
  tele_x_dow_group <- tele_x_dow_pre %>%
    filter(group_label == group_value)

  if (nrow(tele_x_dow_group) == 0) {
    return(NULL)
  }

  lapply(tele_x_dow_category_levels, function(category_value) {
    tele_x_dow_analysis <- tele_x_dow_group %>%
      mutate(indicator = as.numeric(category == category_value))

    tele_x_dow_design <- survey::svydesign(
      ids = ~day_id,
      strata = ~sample_segment,
      weights = ~day_weight,
      data = tele_x_dow_analysis,
      nest = TRUE
    )

    tele_x_dow_estimate <- suppressWarnings(survey::svymean(~indicator, tele_x_dow_design, na.rm = TRUE))
    tele_x_dow_ci <- suppressWarnings(stats::confint(tele_x_dow_estimate))

    tibble(
      group_label = group_value,
      category = category_value,
      estimate = unname(stats::coef(tele_x_dow_estimate)[1]),
      se = unname(survey::SE(tele_x_dow_estimate)[1]),
      ci_low = pmax(0, unname(tele_x_dow_ci[1])),
      ci_high = pmin(1, unname(tele_x_dow_ci[2]))
    )
  }) %>%
    bind_rows()
}) %>%
  bind_rows()

tele_x_dow_fmt <- tele_x_dow_estimates %>%
  left_join(tele_x_dow_totals, by = c("group_label", "category")) %>%
  mutate(
    weighted_total = coalesce(weighted_total, 0),
    unweighted_n = coalesce(unweighted_n, 0L),
    rse = if_else(estimate > 0, se / estimate, NA_real_),
    rse_flag = case_when(rse > 0.5 ~ "**", rse > 0.3 ~ "*", TRUE ~ ""),
    share_label = if_else(estimate >= 0.05, percent(estimate, accuracy = 1), "")
  ) %>%
  mutate(
    group_label = factor(group_label, levels = c("Monday", "Tuesday", "Wednesday", "Thursday", "Friday", "Saturday", "Sunday")),
    category = factor(category, levels = c("Did not telework", "0-1 hours", "1-6 hours", "6+ hours"))
  ) %>%
  arrange(group_label, category)

tele_x_dow_fill_values <- setNames(
  rep(unname(stacked_palette), length.out = dplyr::n_distinct(tele_x_dow_fmt$category)),
  levels(tele_x_dow_fmt$category)
)
tele_x_dow_text_values <- setNames(
  contrast_text_for_fill(unname(tele_x_dow_fill_values)),
  names(tele_x_dow_fill_values)
)

tele_x_dow_plot <- ggplot(
  tele_x_dow_fmt,
  aes(
    x = group_label,
    y = estimate,
    fill = category,
    text = paste0(
      "Day of week: ", group_label,
      "<br>Telework time: ", category,
      "<br>Share: ", percent(estimate, accuracy = 0.1), rse_flag,
      "<br>SE: ", percent(se, accuracy = 0.1),
      "<br>95% CI: ", percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1),
      "<br>RSE: ", percent(rse, accuracy = 0.1),
      "<br>Weighted days: ", comma(round(weighted_total)),
      "<br>Unweighted days: ", comma(unweighted_n)
    )
  )
) +
  geom_col(position = "fill", width = 0.72) +
  geom_text(
    aes(label = share_label, color = category),
    position = position_fill(vjust = 0.5),
    show.legend = FALSE,
    family = "Inter",
    fontface = "bold",
    size = 3.3
  ) +
  scale_y_continuous(labels = percent_format(accuracy = 1)) +
  scale_fill_manual(values = tele_x_dow_fill_values) +
  scale_color_manual(values = tele_x_dow_text_values) +
  labs(x = NULL, y = NULL, fill = NULL, subtitle = "Weighted telework-time distribution by day of week among employed adults") +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.major.x = element_blank(),
    panel.grid.minor = element_blank(),
    plot.subtitle = element_text(color = subtitle_color),
    legend.position = "bottom"
  )

tele_x_dow_fig <- ggplotly(tele_x_dow_plot, tooltip = "text", height = 430)
for (i in seq_along(tele_x_dow_fig$x$data)) {
  tele_x_dow_fig$x$data[[i]]$hovertemplate <- "%{text}<extra></extra>"
  tele_x_dow_fig$x$data[[i]]$showlegend <- identical(tele_x_dow_fig$x$data[[i]]$type, "bar")
  if (!is.null(tele_x_dow_fig$x$data[[i]]$name)) {
    tele_x_dow_fig$x$data[[i]]$name <- sub("^\\((.*),[^,]*\\)$", "\\1", tele_x_dow_fig$x$data[[i]]$name)
    tele_x_dow_fig$x$data[[i]]$legendgroup <- tele_x_dow_fig$x$data[[i]]$name
  }
  if (!identical(tele_x_dow_fig$x$data[[i]]$type, "bar")) {
    tele_x_dow_fig$x$data[[i]]$hoverinfo <- "skip"
  }
}

tele_x_dow_fig %>%
  layout(margin = list(l = 40, r = 20, t = 40, b = 70)) %>%
  config(displayModeBar = FALSE)
Figure 18: Weighted distribution of telework time categories by weekday for employed adults and workers on leave age 16 or older. Shares use day_weight; * marks estimates with 30% < RSE <= 50%, and ** marks estimates with RSE > 50%.
Code
tele_x_dow_fmt %>%
  transmute(
    group_label,
    category,
    Share = paste0(percent(estimate, accuracy = 0.1), rse_flag),
    `Standard error (SE)` = percent(se, accuracy = 0.1),
    `95% confidence interval (CI)` = paste0(percent(ci_low, accuracy = 0.1), " to ", percent(ci_high, accuracy = 0.1)),
    `Relative standard error (RSE)` = percent(rse, accuracy = 0.1),
    `Weighted total` = comma(round(weighted_total)),
    `Unweighted days` = comma(unweighted_n)
  ) %>%
  gt(groupname_col = "group_label") %>%
  cols_label(category = "Category") %>%
  tab_header(title = "Summary of Telework Time by Day of Week (Employed Adults, Weighted)") %>%
  tab_source_note(md("Shares use `day_weight`. Totals shown are weighted counts in the corresponding analytic file. `*` marks estimates with `30% < relative standard error (RSE) <= 50%`, and `**` marks estimates with `RSE > 50%`.")) %>%
  opt_row_striping()
Summary of Telework Time by Day of Week (Employed Adults, Weighted)
Category Share Standard error (SE) 95% confidence interval (CI) Relative standard error (RSE) Weighted total Unweighted days
Monday
Did not telework 57.2% 3.3% 50.9% to 63.6% 5.7% 283,882 518
0-1 hours 5.9%* 1.9% 2.2% to 9.6% 32.2% 29,234 34
1-6 hours 10.2% 2.0% 6.3% to 14.2% 19.7% 50,776 108
6+ hours 26.6% 2.7% 21.3% to 31.9% 10.2% 132,035 324
Tuesday
Did not telework 64.6% 3.0% 58.7% to 70.5% 4.7% 468,060 583
0-1 hours 4.1% 1.2% 1.8% to 6.4% 28.6% 29,860 39
1-6 hours 9.6% 1.8% 6.1% to 13.2% 18.8% 69,722 123
6+ hours 21.6% 2.4% 16.8% to 26.4% 11.3% 156,428 342
Wednesday
Did not telework 61.9% 3.2% 55.7% to 68.1% 5.1% 313,168 535
0-1 hours 4.7% 1.1% 2.5% to 6.9% 23.5% 23,786 44
1-6 hours 9.7% 1.9% 5.9% to 13.5% 20.0% 49,163 112
6+ hours 23.7% 2.6% 18.7% to 28.8% 10.8% 120,122 330
Thursday
Did not telework 55.1% 3.2% 48.7% to 61.4% 5.9% 302,217 562
0-1 hours 4.0%* 1.4% 1.2% to 6.8% 35.9% 21,954 39
1-6 hours 11.5% 2.1% 7.3% to 15.7% 18.6% 63,081 132
6+ hours 29.5% 3.0% 23.6% to 35.3% 10.1% 161,656 333
Shares use day_weight. Totals shown are weighted counts in the corresponding analytic file. * marks estimates with 30% < relative standard error (RSE) <= 50%, and ** marks estimates with RSE > 50%.
Table 21: Summary of telework time by day of week (employed adults, weighted)

1 Study Overview

The 2025 Puget Sound Regional Travel Study is the second data collection year in a household travel survey program cycle that is expected to include four waves, with surveys in 2023, 2025, 2027, and 2029. The survey instrument, overall methodology, and study goals follow the previous three wave data collection program that began in 2017. The 2025 study collected household- and person-level activity and travel pattern information from residents throughout the Puget Sound Regional Council (PSRC) four-county region and was sponsored by PSRC and Pierce County.

1.1 Study Objectives

The overarching goal of the multiyear program is to maintain an updated source of household travel behavior data that supports and allows for the following:

  • Transportation and land-use modeling and planning needs.
  • Trend analysis over time.
  • Regular study design updates to integrate evolving data collection methods and emerging travel behaviors and transportation issues.

1.2 Study Area

Consistent with previous PSRC studies, the 2025 study encompassed the entire four-county PSRC region, which includes King, Kitsap, Pierce, and Snohomish counties. The region includes 82 cities and towns with a total population of over four million people. The study area comprises approximately 1,750,000 households.

1.3 Data Collection Overview

Survey data were collected through a mixed-mode design that combined:

  • Smartphone-based travel diary (rMove) – Participants recorded travel in real time for up to seven consecutive days.
  • Web-based travel diary (rMove for Web) – Participants reported travel for one assigned weekday.
  • Call center interviews – Participants reported travel for one assigned weekday. This mode supported participants without digital access or preferring phone-based completion.

Each household first completed a recruit survey describing household composition, demographics, and vehicles, followed by a travel diary describing all trips made during the assigned day(s).

Further details on survey instruments and question content are provided in Section 3.

1.4 Inclusive Engagement and Recruitment

Recognizing differences in survey accessibility and response rates across the four-county region, the project team used a combination of recruitment materials, incentives, language support, and follow-up procedures to improve participation among historically underrepresented groups and hard-to-reach households. Specific details about sampling frame, stratification, and recruitment procedures are provided in Section 2.

1.5 Weighting and Representation

All completed households were expanded and calibrated to represent the full population of the PSRC region. The weighting process incorporated demographic and household controls derived from external benchmarks and was designed so analysts can produce regionwide and subgroup estimates that reflect the underlying population rather than the realized sample mix.

Guidance on applying weights in analysis and interpreting weighted results is provided in Section 5 and Section 8.

2 Sample Design

The 2025 Puget Sound Regional Council Household Travel Study used a probability-based, geographically stratified sample of households across the central Puget Sound region. The sampling approach applied a uniform framework across the study geographies, including an add-on sample in unincorporated Pierce County, while ensuring coverage across county geographies, land-use contexts, and population groups of interest.

2.1 Sampling Goals

The 2025 study aimed to sample a minimum target of 2,250 complete responses with an upper target of 2,670 complete responses, which equates to a 0.15% target sample rate for the region, and was smaller than the 2023 target sample rate of 0.25%. Typical sample rates for similar studies range from approximately 0.5–1%. Across the 2017, 2019, and 2021 studies, the combined sample rate was 0.6%. The 2014 PSRC study (the last study prior to the recurrent data collection design) also had a sample rate of approximately 0.6%. Table 22 shows the sample minimums and targets by sponsor.

Code
sample_targets <- data.frame(
  sponsor = c("PSRC (Region Total)", "Pierce County"),
  minimum_target = c(2250, 780),
  upper_target = c(2670, 900),
  stringsAsFactors = FALSE
)


sample_targets %>%
  gt::gt() %>%
  gt::cols_label(
    sponsor = "Sponsor",
    minimum_target = "Minimum target complete households",
    upper_target = "Upper target complete households"
  ) %>%
  gt::fmt_number(
    columns = c(minimum_target, upper_target),
    decimals = 0
  ) %>%
  gt::sub_missing(
    columns = c(minimum_target, upper_target),
    missing_text = "-"
  ) %>%
  gt::cols_align(
    align = "left",
    columns = sponsor
  ) %>%
  gt::cols_align(
    align = "center",
    columns = c(minimum_target, upper_target)
  ) %>%
  gt::cols_width(
    sponsor ~ gt::px(250),
    minimum_target ~ gt::px(120),
    upper_target ~ gt::px(120)
  ) %>%
  gt::tab_options(
    table.width = gt::pct(100),
    column_labels.font.weight = "bold",
    table_body.hlines.width = gt::px(1)
  ) %>%
  gt::opt_row_striping()
Minimum target complete households Upper target complete households
PSRC (Region Total) 2,250 2,670
Pierce County 780 900
Table 22: Minimum and upper sample targets for the 2025 PSRC Household Travel Study.

2.2 Sampling Framework

Surveyable Population

In previous waves of the PSRC household travel survey, not all household members were surveyable: only persons related to Person 1 (the primary respondent) were considered surveyable. Non-surveyable members (e.g., guests, visitors, or unrelated roommates) did not have trip or day completion requirements and were excluded from household-level completeness determinations.

In the 2025 wave, all household members were considered surveyable regardless of their relationship to the primary respondent. This change reflected a desire to align household membership definitions to those used by the Census, and to capture atypical travel behavior present in households with non-traditional structures.

Sampling Frame

The sampling frame for the study was the list of all households, as defined by the Census, in King, Kitsap, Pierce, and Snohomish counties, excluding people living in group quarters. The project team used address-based sampling (ABS) to select and invite households to participate in the study. Under this approach, a random sample of residential addresses was drawn from the sampling frame so that households within each defined area had an equal chance of selection. The project team purchased household mailing addresses from Marketing Systems Group, which maintains the Computer Delivery Sequence file from the U.S. Postal Service.

The project team geographically stratified the ABS using Census Block Group (BG) data from the 2018-2022 American Community Survey (ACS) 5-year estimates, the most recent available at the time of sample planning. BGs are the smallest geography for which most Census and ACS tables are publicly available. According to those ACS data, the region contains 2,916 BGs, 1,671,727 households, and a total population of 4,268,571 persons, including children. Group quarters were excluded from the sampling frame, as were BGs with no reported households.

Primary Sampling Unit

The primary sampling unit was the household, selected through random sampling from the ABS frame. All household members were in scope for person-level and trip-level analysis, but only one member (the “primary respondent”) was required to complete the recruit survey on behalf of the household. Post-processing produced analysis files at multiple levels–household, person, day, vehicle, and trip. In this guide, the delivered trip table represents person-trips where transit access, egress and transfers are consolidated into a single record.

Though the primary sampling unit was the household, the data collected also represented the behavior of individual persons. For participants who reported data using the smartphone, data were collected across multiple days, representing a multitude of travel and daily activity data. As shown in Figure 19, the schematic illustrates the relationship between a mailed invitation and a household’s travel patterns.

Figure 19: Hierarchical structure of the data, from sampling to observed travel patterns measured across multiple days.

Sample Geographies

To reflect project sponsorship by PSRC and Pierce County and to ensure sufficient sample throughout the region, the survey team identified five geographic groups to stratify the sample based on data and analysis needs in the region. These groups were: 1. King County 2. Kitsap County 3. Pierce County (Incorporated) 4. Pierce County (Unincorporated) 5. Snohomish County

Code
knitr::include_graphics("images/Survey_Region_Map.png")
Figure 20: Map of the survey region and sample geographies.

Sample Strata

Within each geography, the study used the following mutually exclusive and collectively exhaustive sample strata to promote representation of groups that are typically hard to reach and provide a larger sample of groups of particular interest. These definitions aligned with those used in the 2023 sample plan.

  1. General population: Comprised of block groups that do not qualify for oversampling strata below.

  2. Hard-to-survey Oversample: Comprised of block groups that meet one or more of the following criteria:

  • At least 35% of households earn less than 200% of the Federal poverty level.
  • At least 60% of persons identify as Hispanic or Latino.
  • At least 40% of census tract persons identify as:
    • non-White and non-Asian, or
    • Asian and have a household income less than $50,000 per year.
  • At least 20% of households have limited English capabilities.
  1. Walk/Bike/Transit Oversample: Comprised of block groups with at least 55% of persons reporting walk, bicycle, or public transportation as their means of transportation to work, or at least 15% of households owning zero vehicles. Note: This stratum was present only in King County. While there were a handful of block groups that met this condition in other counties, their relatively small number (fewer than 10) led the project team to fold them into the General strata outside King County.

  2. Rural: Comprised of block groups that align with PSRC’s “Rural Areas and Natural Resources Lands” regional geography. These areas describe different types of unincorporated areas outside the urban growth area and include very low-density housing, working landscapes, and open space.

If a block group qualified for both the Hard-to-Survey and Walk/Bike/Transit stratum, it was classified as Hard-to-Survey. If a block group qualified for both Hard-to-Survey and Rural stratum, it was classified as Rural.

Code
knitr::include_graphics("images/Survey_Region_Map_Strata.png")
Figure 21: Map of the survey region and sample strata.

2.3 Recruitment and Field Procedures

Recruitment Channels

Households were recruited primarily by mailed invitation letters containing a unique survey access code. Reminder postcards were issued to nonrespondents. Invitations were mailed across 10 groups over the data collection period, with the first several mail groups used to evaluate observed response and inform subsequent mailings.

The first invitation letter group was scheduled to mail on February 14, 2025, with a reminder postcard following approximately one week later. Subsequent mail groups were planned at roughly weekly intervals through May, with a short late-March pause available for sample adjustments if needed. The final travel day was scheduled for June 8, 2025.

Participation Modes

Eligible households could participate via:

  • rMove smartphone app (7-day diary),
  • rMove for Web (1-day diary), or
  • Call-center interview (1-day diary).

App users were asked to record travel in real time for up to seven consecutive days, while web and call-center participants reported travel for one assigned weekday.

Mode assignment depended on household technology access and preference. Smartphone participation was encouraged, but households without smartphone access or who preferred not to use the app could choose the web-based diary or call-center interview.

Incentives

Households could choose a Starbucks e-gift card, an Amazon e-gift card, a Visa gift card by mail, or no gift card.

The incentive amount depended on diary platform and whether the household qualified for a higher incentive. For households completing by browser or call center, incentives were paid per household and set at $10 for households that did not qualify for a higher incentive and $25 for those that did. For households completing by rMove, incentives were paid per participating adult and set at $40 for those that did not qualify for a higher incentive and $50 for those that did.

Higher incentive eligibility was tied to selected household characteristics captured during recruitment, including household income below $50,000, household size of five or more people, Hispanic, Latino, or Spanish origin responses, and race responses identifying Black, American Indian or Alaska Native, Native Hawaiian or other Pacific Islander, or other race.

Incentive structure

Platform Payment unit Qualified for higher incentive Incentive amount
browser/call center Per household no $10
browser/call center Per household yes $25
rMove Per participating adult no $40
rMove Per participating adult yes $50
Table 23: Programmed incentive amounts by diary platform and qualification status.

Monitoring and Response Tracking

During data collection, the project team evaluated response patterns and adjusted invitation volumes as needed to keep the study on track toward its sample targets. In parallel with the mailing schedule, the project team repeated survey write-out testing during data collection and reviewed response patterns by sample segment as part of in-field quality control.

Code
knitr::include_graphics("images/monitoring_dashboard.png")
Figure 22: Example of response monitoring by sample geography and mail group during data collection.

2.4 Pre-Collection Testing

Before live data collection, the project team conducted topline survey testing and survey write-out testing. Topline testing was used to confirm that questions were clearly worded and that survey elements such as response options, links, buttons, and maps rendered and functioned as expected across devices. Survey write-out testing was then used to verify that recorded survey data followed the programmed survey logic and that responses appeared when expected and remained blank when not applicable.

Together, these testing steps were intended to identify issues early, confirm that survey responses were being written correctly, and support a smooth transition into live fielding.

3 Survey Instrument

The 2025 Puget Sound Regional Council Household Travel Study collected information about households, people, vehicles, and daily travel through a coordinated survey instrument implemented across smartphone, web, and call-center modes.The instrument used a common structure across modes while allowing mode-specific prompts and workflows appropriate to each reporting platform. View the 2025 Puget Sound Travel Study Questionnaire for the full set of questions asked in the study.

3.1 Instrument Structure

Part 1: Recruit Survey

The recruit survey collected household roster and baseline information used to determine eligibility, assign diary participation, and tailor follow-up prompts. Core modules included:

  • Household composition and member roster
  • Vehicle ownership and fuel type
  • Demographic characteristics, including age, gender, race, ethnicity, and household income
  • Employment, work-from-home arrangements, student status, and related work or school location questions
  • Smartphone access, participation mode selection, and incentive selection

The recruit survey was completed by the primary respondent on behalf of the household. Depending on household composition and participation mode, the recruit survey also assigned proxy-reporting responsibilities for children (under 18 years of age) and collected contact information needed for reminders and follow-up.

Part 2: Travel Diary

The travel diary collected information about travel made on the assigned reporting day or days. In the smartphone app, participants reviewed passively collected travel and completed prompted trip surveys. In rMove for Web and the call center, respondents reported travel directly through prompted diary instruments.

Across modes, the travel diary collected:

  • Day-begin and day-end location confirmation
  • Whether no travel occurred on the assigned day
  • Trip destinations, purposes, modes, and timing
  • Access, transfer, and egress details for transit trips
  • Companion and escort activity where applicable
  • Follow-up questions to confirm travel to reported home, work, or school locations

Adults could report their own travel, while proxy reporting was used where needed for children (under 18 years of age) and other eligible household members.

TipAnalyst Tip: Multi-Day Travel Data

Only participants who used the rMove smartphone app recorded travel for multiple days, including weekends. Weekend days are not weighted in the final dataset; weights represent Monday-Thursday travel.

Part 3: Daily Surveys

The daily survey collected information about conditions and activities associated with each reporting day. Core daily questions asked where the day began and ended, whether any travel occurred, how much time was spent teleworking, and whether deliveries or household services occurred that day.

Additional daily modules collected information on work and school schedules, commute frequency, telework frequency, commute mode, school mode, employer commute benefits, deliveries, moving history, and transportation attitudes. Some of these modules were asked on the first reporting day, while others were reserved for later days in the seven-day smartphone protocol.

For rMove for Web and call-center respondents, daily questions were incorporated into the one-day reporting workflow. For smartphone participants, daily questions were distributed across the reporting period and paired with prompted trip review.

3.2 Topic Areas

The survey instrument covered standard household travel survey topics as well as several special-topic modules.

Category Description
Household Household composition, vehicles, income, home and second-home context
Person Demographics, employment, student status, smartphone access, and proxy assignment
Travel day Assigned travel date, begin/end-of-day location, no-travel confirmation, and daily context
Trip Destinations, purposes, modes, timing, transfers, and related travel details
School and work context Workplace and school locations, commute and telework frequency, school travel, and employer commute benefits
Special topics Deliveries, household services, moving history, and transportation attitudes
Table 24: Major survey instrument topic areas.

Question wording and skip logic were aligned across modes to preserve comparability in the delivered analysis variables.

3.3 Travel Date Assignment

Households were assigned a Monday, Tuesday, Wednesday, or Thursday travel date during the study period. Households participating through rMove then received a seven-day reporting period beginning with the assigned start date. Households completing by rMove for Web or by call center reported travel for one assigned day and completed the survey after that travel date.

3.4 Survey Modes and Language Support

The survey instrument was administered through three participation modes:

  • rMove smartphone app – a seven-day smartphone-based diary with passive trip collection and prompted trip review
  • rMove for Web – a one-day web diary completed after the assigned travel date
  • Call center – a one-day diary completed by telephone

Households eligible to use rMove were offered the option to participate through the smartphone app. If a household did not use the smartphone app, travel could be reported online through rMove for Web or through the call center.

The questionnaire included an activation-language step at the beginning of the survey. The survey also provided call-center completion for households reporting travel by telephone.

The survey instruments were available only in English. Households that spoke Spanish, Chinese, Russian, Korean, Tagalog, Vietnamese, or Somali had the option to call the toll-free line to complete the survey over the phone in their preferred language. The call center received ten calls in Spanish, three calls in Vietnamese, and one call in Russian.

3.5 Questionnaire Changes: 2023 to 2025

Relative to the 2023 instrument, the 2025 questionnaire placed greater emphasis on measuring work arrangements, telework, deliveries, and other behaviors that have become more important for regional travel analysis. At the same time, the instrument streamlined several lower-priority or redundant items, simplified portions of the trip diary, and updated household and person logic to better reflect the full household roster.

Overall, the 2025 instrument expanded detail where travel behavior and policy context have changed most, while reducing respondent burden in parts of the survey where additional detail provided less analytic value.

Key themes in the 2025 update included:

  • Expanded work and telework measurement, including leave status, work-from-home arrangements, workplace availability, and travel-for-work logic
  • Additional daily modules on deliveries, household services, transportation attitudes, and moving-related factors
  • Broader household coverage by removing earlier related-member restrictions and applying the instrument to all household members
  • Simplified or removed lower-value items such as toll transponder ownership, transit pass ownership, commute duration, broadband access, and number of jobs
  • More modular daily survey design and closer alignment across rMove and web-based reporting workflows
Table 25: Major questionnaire changes from 2023 to 2025.
Domain Question or module Status Change type Summary Detail
Household roster
Signup Household relationship filter
define_related
Deleted Logic / sample universe Removes related-member restriction for household participation. 2023 restricted participation to members defined as related. 2025 includes all household members. This modifies the sample universe and reporting eligibility logic.
Signup Age
(2023: age; 2025: age_detailed)
Modified Response options / measurement resolution Modifies age measurement to detailed values instead of categories. 2023 used grouped age categories. 2025 records detailed ages across a full numeric range. This increases measurement resolution without changing concept.
Signup Proxy assignment
proxy
Modified Logic / sample universe Modifies proxy assignment independent of related-member definition. 2023 limited proxy selection to related adults. 2025 allows any adult. This reflects removal of related-member constraints.
Employment
Signup Employment status
(2023: employment; 2025: paid_work + employment)
Modified Structural redesign Modifies employment measurement by separating work activity and classification. 2023 captured employment in a single question. 2025 introduces a paid_work screener followed by employment classification. This restructures logic and separates recent activity from status.
Signup Nonworker classification
employment_followup
Added Logic / sample universe Adds follow-up classification for nonworking individuals. 2025 introduces employment_followup for persons without paid work. It distinguishes unemployed, retired, student, and unpaid roles. This refines classification logic for nonworkers.
Signup Work location and telework
(2023: job_type; 2025: work_from_home + job_type)
Modified Structural redesign Modifies work location by separating telework from job type. 2023 embedded telework within job_type categories. 2025 separates work_from_home from job_type. This clarifies telework frequency and work location logic.
Signup Drive for work
drive_for_work
Added Logic / sample universe Adds follow-up identifying workers who travel to access vehicles. 2025 adds drive_for_work for mobile workers. It captures whether travel is required to pick up a vehicle. This expands logic for nontraditional work patterns.
Signup Workspace availability
office_available
Modified Logic / sample universe Modifies workspace concept to include any available work location. 2023 asked about dedicated private workspace. 2025 asks about any available work or business space. This broadens applicability across worker types.
Signup Return-to-work expectations
return_from_leave
Added Logic / sample universe Adds expectations for workplace return among workers on leave. 2025 adds return_from_leave for individuals on leave. It captures expected future commuting frequency. This supports modeling of post-leave behavior.
Signup Primary work location
work_loc
Modified Logic / sample universe Modifies work location to reflect broader work space definition. 2023 focused on primary workplace location. 2025 refers to work or business space away from home. This broadens applicability for hybrid workers.
Signup Available workspace location
workspace_loc
Added Logic / sample universe Adds location of available work space separate from usage. 2025 introduces workspace_loc for available work locations. It captures location even if not regularly used. This supports analysis of latent workplace access.
Daily core Telework duration
telecommute_time
Modified Response options / measurement resolution Modifies telework duration increments to coarser intervals. 2023 used 15-minute increments. 2025 uses 30-minute increments. This reduces burden while slightly lowering precision.
Day 1 Work frequency
work_freq
Added Logic / sample universe Adds typical number of days worked per week. 2025 introduces work_freq to capture weekly work patterns. It provides direct measurement of work frequency. This supports improved modeling inputs.
Day 1 Commute frequency
commute_freq
Modified Logic / sample universe Modifies commute frequency routing to include additional cases. 2023 excluded certain worker types. 2025 includes teleworkers and workers on leave under specific conditions. This expands applicability of commute frequency.
Day 1 Telework frequency
telework_freq
Modified Logic / sample universe Modifies telework frequency logic to align with new employment structure. 2023 linked telework to job type. 2025 uses paid_work and work_from_home variables. This aligns routing with revised employment design.
Day 1 Commute tenure
commute_dur
Deleted Burden reduction / simplification Removes commute tenure question. 2023 captured years commuting to a workplace. 2025 removes this question. This reduces burden and removes nonessential contextual detail.
Day 1 Commute benefits
commute_subsidy
Modified Question text / wording Modifies commute benefits wording to include employer or business. 2023 referenced employer-provided benefits. 2025 includes employer or business. This broadens applicability to self-employed individuals.
Day 1 Use of commute benefits
commute_subsidy_use
Modified Question text / wording Modifies benefit usage wording to align with broader employer definition. 2023 referenced employer benefits. 2025 includes employer or business benefits. This ensures consistency across employment types.
Day 1 Typical mode to work
work_mode
Modified Response options / measurement resolution Modifies answer options to distinguish between walking and driving onto a ferry. 2023 captured a single ferry mode. 2025 offers two separate options to walk or drive onto a ferry.
Education
Signup Adult student status
(2023: student; 2025: adult_student)
Modified Question text / wording Modifies student variable naming for clarity and consistency. 2023 used student for adults. 2025 renames to adult_student. This aligns naming with PSRC conventions without changing meaning.
Signup School type
school_type
Modified Response options / measurement resolution Modifies school categories by consolidating response options. 2023 used more granular school categories. 2025 combines several categories into broader groups. This simplifies responses while preserving key distinctions.
Day 1 School travel frequency
school_freq
Deleted Structural redesign Removes adult school travel frequency question. 2023 asked how often adults traveled to school. 2025 removes this item. This simplifies the adult student module.
Day 1 Typical mode to school
school_mode
Added Logic / sample universe Adds typical mode to school 2023 did not ask typical mode to school. 2025 added this question to support school trip impuation.
Day 1 Remote class frequency
remote_class_freq
Deleted Structural redesign Removes adult remote class frequency question. 2023 captured frequency of remote classes. 2025 removes this item. This reduces redundancy in student reporting.
Travel
Signup Toll pass ownership
toll_transponder
Deleted Burden reduction / simplification Removes toll transponder ownership question. 2023 asked whether vehicles had a toll pass. 2025 removes this question. This reduces respondent burden and removes noncore detail.
Signup Driving status
hh_licenses
Modified Logic / sample universe Modifies driving eligibility threshold and simplifies wording. 2023 applied to persons age 16+. 2025 includes persons age 15+. This expands eligibility and simplifies response options.
Daily core No travel indicator
no_travel
Modified Structural redesign Modifies no-travel question to binary indicator instead of reasons. 2023 captured reasons for no travel. 2025 uses a yes/no indicator. This simplifies the question and removes explanatory detail.
Day 2 Mode use (30-day)
share
Modified Question text / wording Modifies mode-use question to specify 30-day reference period. 2023 did not specify timeframe. 2025 explicitly references past 30 days. This improves consistency in reported behavior.
Day 2 Transit pass ownership
transit_pass
Deleted Burden reduction / simplification Removes transit pass ownership question. 2023 asked about transit pass ownership. 2025 removes this item. This reduces burden and removes fare-media detail.
Day 4 EV charging location
ev_typical_charge
Modified Logic / sample universe Modifies EV charging eligibility to include all EV households. 2023 applied to primary EV drivers. 2025 applies to households with any EV. This broadens the sample universe.
Trip Detailed mode follow-ups
mode_*
Deleted Structural redesign Removes detailed trip-mode subtypes and associated follow-up questions. 2023 captured detailed subtypes for vehicle, transit, and micromobility modes. 2025 removes these follow-ups and records mode at a higher level. This simplifies the instrument and reduces mode-specific detail.
Delivery
Daily core Delivery and services
delivery
Modified Logic / sample universe Modifies eligibility for deliveries to workplace In 2025, eligibility for certain responses—particularly workplace deliveries—is based on broader employment criteria rather than detailed job-type conditions, improving consistency and coverage across worker types.
Demographics
Day 3 Disability
disability
Modified Question text / wording Modifies disability question to function-based wording. 2023 referenced disability affecting travel. 2025 uses functional difficulty wording. This aligns with federal survey standards.
Day 4 Primary vehicle used
vehicle
Deleted Structural redesign Removes primary vehicle identification question. 2023 linked respondents to a primary vehicle. 2025 removes this question. This eliminates driver-vehicle linkage detail.
Housing
Day 4 Months at residence
res_months
Deleted Burden reduction / simplification Removes months-at-residence question. 2023 captured seasonal residence duration. 2025 removes this item. This reduces burden and noncore detail.
Day 4 Broadband access
broadband
Deleted Burden reduction / simplification Removes broadband access question. 2023 captured internet access availability. 2025 removes this item. This eliminates a contextual but nonessential variable.
Attitudes
Day 5 Attitudinal battery
transportation_statement_*
Added Structural redesign Adds attitudinal battery on travel preferences and behaviors. 2025 introduces multiple Likert-scale statements on travel preferences. Topics include car ownership, transit, telework, environment, and active travel. This adds a new behavioral and attitudinal dimension.

4 Data Processing and Structure

RSG first arranged raw survey responses into tabular format and then applied a series of processing steps to clean, validate, and enhance the data. This chapter describes key aspects of the data processing pipeline.

Data processing is primarily focused on the trip and day tables. The household and person tables are also processed but contain fewer derived variables. The processing steps described below apply to all records in the dataset, but only records that meet certain completeness criteria (see Section 4.1) are eligible for weighting and analysis.

4.1 Completion Flags

Several processing tables in the pipeline include harmonized completion flags that determine whether a household, person, or day is eligible for weighting and analysis. These indicators ensure that only fully usable records contribute to weighted estimates (see Some Weights are Zero).

The data contains records from complete households, but not all records are themselves complete. Incomplete records can be identified by zero weights (see Some Weights are Zero) or is_complete flags.

Completion works upwards from the trip level to the household level:

  1. In the trip table, trip_survey_complete = 1 if the respondent answered all questions about their trip in the travel diary.
  2. In the day table, is_complete = 1 if all trip surveys on that person-day are complete and the proxy/daily survey is complete; or if a respondent reports no trips and responds affirmatively that no trips occured on their travel day in the daily survey. When all surveyable (see Section 2.2.1) household members are complete on that day, hh_day_complete = 1. Households with children will only have up to one complete day.
  3. The dataset is filtered to households that have one complete concurrent day that falls on a weighted weekday (Monday-Thursday) (hh_day_complete on travel_dow 1-5).

In the study, households with children were only required to report travel for children on one day, even if the household participated in a multi-day rMove survey. Meanwhile, only days with concurrent complete travel diaries across all household members were complete-eligible. This means that rMove households with children may have up to six unweighted, incomplete days even if all adults in the household were complete on non-proxy days.

4.2 Trip Table Processing

RSG applied special treatments, which are described below, to trips.

Speed Plausibility Flag (speed_flag)

RSG created a simple plausibility check for trip speeds to help identify likely GPS or processing errors. First, RSG computed an average trip speed in miles per hour using reported distance and duration:

speed_mph = distance_miles / (duration_seconds / 3600)

Because realistic speeds depend heavily on travel mode, RSG then defined a mode-specific maximum plausible speed (max_speed) for each value of mode_type, using the ranges in Table 26:

Mode Description mode_type(s) Max Plausible Speed (mph)
Walk 1 13
Bike, Bike-Share 2, 3 40
Taxi, TNC, Other, Car, Car-Share, School bus 5, 6, 7, 8, 9, 10 100
Shuttle/Vanpool, Transit 11, 13 160
Long Distance Passenger 14 600
Table 26: Maximum Plausible Speeds by Mode Type. These thresholds were used to flag trips with implausible speeds.

For each trip, RSG compared the observed speed_mph to the corresponding max_speed:

  • speed_flag = 1 if speed_mph exceeded the mode-specific max_speed
  • speed_flag = 0 otherwise
  • Trips with modes that did not receive a max_speed (e.g., missing response) were assigned speed_flag = 0 and treated as not flagged by this check

These thresholds are intentionally generous and are meant to flag implausible speeds, not enforce strict behavioral limits. Analysts can use speed_flag to (a) exclude clearly invalid trips when analyzing distances and speeds, or (b) inspect flagged trips as part of broader data quality checks.

Trip Routing and Distance Measures

For this study, RSG delivered three distance fields for each trip: distance_beeline_meters, distance_meters and distance_miles. These represent measures of trip length, based on either straight-line geometry or a combination of raw GPS traces and network routing.

Straight-Line (“Beeline”) Distance

distance_beeline_meters provides the great-circle (Haversine) distance between the trip’s start and end coordinates. It does not use any of the intermediate GPS points and does not consider the street network. This metric is included so analysts have access to a simple, geometry-based reference distance that is comparable across all trips, regardless of whether trace data were available or valid.

Cleaned / Network-Routed Distance

distance_meters represents the primary trip distance used in analysis and reporting. Its definition depends on the trip type:

  • Trips with GPS trace data (rMove app trips): distance_meters reflects a cleaned trace distance, computed by summing the point-to-point distances after removing invalid points, smoothing artifacts, reversed traces, and other trace-quality issues. This produces a more accurate representation of how far the participant actually traveled based on the observed location data.

  • Trips without GPS trace data (Browser, Call Center and manually added rMove app trips): Because these trips do not include a usable trace, distance_meters is obtained using RSG’s Open Street Maps-based routing engine to compute a network-routed distance between the trip’s start and end points along the transportation network. This ensures these trips are comparable to traced trips and not underestimated by relying solely on straight-line distance.

Distance in Miles

distance_miles is a simple unit conversion of distance_meters into statute miles. No additional processing or logic is applied. This field is provided for convenience so downstream workflows, summaries, and model inputs can use miles directly without performing conversions.

Trip Mode Type

In the survey instrument, users can select multiple modes for a single trip. The variable mode_type synthesizes mode_1 to mode_4 down to a single, easier-to-use variable for analytical purposes using an established hierarchy. Higher values of mode_type are prioritized over lower values in derivation. For example, transit trips, with mode_type 13, are prioritized over walk trips, with mode_type 1.

When transit trips were unlinked using the Google API during cleaning, and thus did not have a reported mode_1, mode_2, or mode_3, the non-transit legs of the trip were recoded using Google’s suggested mode (most frequently walk or bike).

In PSRC post-processing, mode_type is replaced by mode_class that is also created by a hierarchy logic. See following steps for how mode_class is created:

  1. prepare mode group assignment: assign all trip modes in mode_1, mode_2, mode_3, mode_4 to mode groups

    Mode group assignment for all modes in 2025
    mode in mode_1~mode_4 mode group
    Household vehicle (or motorcycle) Drive
    Other vehicle (e.g., friend's car, rental, carshare, work car)
    Bus (public transit) Transit
    Ferry or water taxi
    Rail (e.g., train, subway)
    Walk (or jog/wheelchair) Walk
    Bicycle or e-bicycle Bike
    Scooter, moped, skateboard Micromobility
    Uber/Lyft, taxi, or car service Ride Hail
    School bus School Bus
    Airplane or helicopter Airplane or helicopter
    Other bus, shuttle, or vanpool (private service or shuttles for older adults and people with disabilities) Other
    Other
    Missing Response Missing Response
  2. assign mode_class with hierarchy logic: go through each level of the hierarchy from 1 to 10 and assign mode_class (once records are assigned, they are removed from the later assignment)

    `mode_class` hierarchy logic
    level mode group (if exist in any of mode_1~mode_4) mode_class assignment
    1 Airplane or helicopter Other
    2 School Bus School Bus
    3 Transit AND mode_acc is NOT "Drove onto ferry" (*) Transit
    4 Ride hail Ride hail
    5 Other Other
    6 Drive OR (* Transit AND mode_acc is "Drove onto ferry") Drive SOV/HOV 2/HOV 3+ (**)
    7 Bike Bike
    8 Micromobility Micromobility
    9 Walk Walk
    10 Missing Response Missing Response

(* In 2025, we have “Drove onto ferry” as a new transit access mode. Those trips that people brought their vehicle onto the ferry should be characterized as Drive)

(** Drive SOV/HOV 2/HOV 3+ is assigned by considering travelers_total. If travelers_total==1, assign Drive SOV. If travelers_total==2, assign Drive HOV 2. If travelers_total>=3, assign Drive HOV 3+. All Drive SOV made by children under 16 has been replaced with Drive HOV2)

  1. assign mode_class_5: create more aggregated trip mode for general use in data analysis projects

    `mode_class` hierarchy logic
    mode_class mode_class_5
    Drive SOV Drive
    Drive HOV2
    Drive HOV3+
    Transit Transit
    Walk Walk
    Bike Bike/Micromobility
    Micromobility
    Ride hail Other
    School Bus
    Other
    Missing Response Missing Response

Trip Departure Time

The rMove app occasionally detects the start of a trip a few minutes after it actually occurs. When this happens, the reported departure time can be too late, which in turn produces unrealistic trip durations or speeds. To correct for these late pickup cases, the fields depart_date, depart_hour, and depart_minute were adjusted using the following procedure.

1. Primary imputation using speed and distance

For most trips, departure time is imputed based on observed movement:

  • rMove estimates the median speed across all recorded trip locations (excluding the origin point).
  • The distance between the origin and the next recorded point is divided by this median speed to infer how long the device was likely in motion before the first ping.
  • This inferred time is subtracted from the first recorded timestamp to generate the imputed departure time.

This method works well when the trip has enough GPS points to characterize speed reliably.

2. Special rule for trips with very few GPS points

Some trips include fewer than three recorded locations. In those cases, speed cannot be estimated:

  • The departure time is instead set to three minutes earlier than the originally reported departure time, reflecting rMove’s typical 3–5 minute ping interval.
  • Exception: Split loop trips may also contain only a few points, but they reuse the imputed departure time generated before the loop was split. These are not subject to the three-minute rule.

3. Logical consistency checks

Two additional adjustments ensure that imputed times remain valid:

  • If the imputed departure time overlaps with the arrival time of the previous trip, the previous trip’s arrival time is used instead.
  • If the imputed departure time is later than the originally reported time, the system defaults back to the original timestamp; departure times should never move forward in time through imputation.

4. Trips that never receive imputed departure times

Two types of trips are always assigned the original reported departure time:

  • User-added trips, since they are not subject to rMove late-pickup behavior.
  • Long-distance passenger-mode trips (e.g., air travel), because rMove cannot collect accurate traces while the phone is in airplane mode. In these cases, GPS-derived speed information is inherently unreliable.

5. Downstream effects

Once a consistent departure time is established—whether original or imputed—all trip duration and speed calculations use the adjusted departure time.

Trip Purpose

Respondents report the purpose of the trip destination in each trip survey. The origin purpose is derived from the destination purpose of the previous trip, except for the first trip in the travel period or where an rMove trip occurs after a trip where the respondent did not provide a destination purpose. For the first trip in the travel period, the origin purpose can be inferred from begin_day in the day table.

When purpose was not asked because an analyst split a user-reported trip during data cleaning and created new destination along a trip, purpose values are derived where possible based on proximity (within 150 meters) to estimated home, work, or school locations. If the location is not proximate to home, work, or school locations, the purpose is set to other.

The purpose category variables (origin_purpose_cat, dest_purpose_cat) contain aggregated purpose values based on the type of purpose at the origin or destination of each trip. Dataset users are welcome to perform their own recoding of the purpose categories as well.

Trip purposes have also been imputed in cases where a purpose reported by the user is assumed to be inaccurate based on information about that person’s reported habitual locations and other trips (primarily to home, work, and school locations). The trip purpose imputation approach was applied to all rMove trips in person-days with at least 1 complete trip and no more than 10 incomplete trips. (Incomplete trips are trips for which the respondent did not answer the trip-specific survey questions about purpose, mode, etc. for the given trip.)

Various tests were applied in logical sequence to trips for which the stated purpose was not consistent with the location type based on the reported habitual locations. In general terms, the tests were designed to:

  • Check the respondent’s reported destination purpose when it conflicts with the destination location type. (The details of the tests depend on the trip purpose, with different criteria used for change-mode trips, escort trips, linked transit trips, trips with home destinations but other reported purposes, etc.)
  • Identify cases where respondents swapped the order of two or more trips when reporting their details.
  • Identify cases where respondents may have omitted a trip and shifted remaining reported trip details by one trip when reporting the rest of their trips.
  • Fill in missing data by sampling destination purposes from other trips made to the same locations, either by the same respondent or by other respondents.

No trips were removed from the dataset based on purpose imputation. Instead, the original reported purpose is retained in dest_purpose and origin_purpose, while the imputed purpose is stored in dest_purpose_cat and origin_purpose_cat. This allows analysts to choose whether to use the imputed purposes or the original detailed reported purposes in their analyses.

PSRC in-house data cleaning

PSRC has programmed procedures to revise missing or blatantly incorrect survey data, including:

  • trip insertion when the distance between a trip destination and the following trip origin are over 500m
  • purpose revision for trips with contradictory data where reporting matches narrow, observed error patterns
  • missing purpose imputation when Google Places API location type corresponds unambiguously to a trip purpose
  • missing mode imputation when trip duration matches Google Routes API results
  • travel time revision using Google Routes API when speed is unrealistic (travel duration and speed revised correspondingly)
  • trip linking when trip components were reported as if they were complete trips

Afterwards, in cases where trip attributes are still missing or are ostensibly contradictory, PSRC staff examine and, where warranted, edit the data to resolve potential reporting errors. The potential errors we identify include:

  • trip attributes that contradict personal attributes (e.g. child driving to work; adult attending elementary school)
  • trip purposes that contradict the location (e.g. “went home”, when the destination is not home)
  • trip purposes that contradict timing (e.g. grocery shopping in less than a minute; spending the night at the gas station)
  • impossible or improbable timing (travel that is much too fast or slow, or trips that overlap temporally)

5 Weighting

This section summarizes the weighting and expansion procedures used in the 2025 Puget Sound Regional Council Household Travel Study dataset. The goal of weighting is to expand the survey sample so that it represents the full resident population across the central Puget Sound region.

RSG’s weighting procedures follow a multi-stage approach that begins with design-based base weights and proceeds through several rounds of adjustment to correct for demographic nonresponse, diary-platform reporting bias, and trip-type under-reporting.

NoteWhat are survey weights?

To produce statistics that represent an entire population without surveying every household or individual, survey researchers assign weights to each completed observation. In household travel surveys, the survey weight indicates how many people, households, days, or trips in the population a given respondent or record is estimated to represent. By applying these weights, analysts can generate unbiased regional estimates even when the sample is only a small fraction of the full population.

5.1 Overview of Weighting Goals

The weighting process aligns weighted survey estimates with external population totals and distributions across key household, person, day, and trip characteristics. Weighting corrects for differential sampling, differences in survey completion across demographic groups, and systematic differences in trip reporting that arise from the method respondents used to report their travel (smartphone app, web diary, or call center). The final weights allow analysts to produce household, person, day, and trip estimates that reflect the true resident population.

5.2 What do these weights represent?

Across all steps, the weighting process produces four types of final weights:

  • Household weight: expands each surveyed household to represent the total number of households in its weighting zone group (Figure 23) and the full study region.
  • Person weight: expands each person to represent the population of persons.
  • Day weight: expands each weekday person-day to represent one average weekday of travel.
  • Trip weight: expands trips between an origin and destination regardless of transit transfers, access, and egress legs.
NoteSome Weights are Zero

The final dataset contains many weights equal to zero. When a weight is equal to zero, it means that the record was not eligible for weighting for the following reasons:

  • Partially complete records. For example, if a household participated for seven days, but only provided three days of complete diary data, the remaining days would be flagged as incomplete and not weighted.
  • For households with children, non-proxy days. Households with children were only required to report travel for one day, on which a proxy reporter could report on behalf of children. Additional days reported for children without a proxy reporter were not weighted.
  • Days outside of the “typical weekday” definition. Data were weighted to represent a “typical weekday.” Friday, Saturday and Sunday data are provided only in their raw form and were not weighted.

5.3 Inputs to Weighting

The 2025 Puget Sound Regional Council Household Travel Study used two primary datasets as inputs to the weighting process:

  • Survey data, consisting of cleaned and imputed household, person, day, and trip files that exclude incomplete households and include imputed demographic variables. Only complete diary days were included in the weighting process.
  • Target data, constructed using 2023 ACS 5-year estimates and ACS 1-year PUMS data. These data provide total household and population counts for each weighting zone group, and detailed demographic distributions used as control totals in weighting.

Survey sample households are assigned to weighting zone groups defined for this project. These were made to align with the sample segment geographies and were built from grouped 2022 PUMAs using the household field home_puma_2022. PUMS records are allocated to these same grouped-PUMA geographies to establish household and person control totals. The weighting zone groups used in this project were defined at the PUMA (Public Use Microdata Area) level, with groupings chosen to align the survey with ACS/PUMS control geography while still maintaining enough sample for stable weighting. The final weighting zone groups are shown in Figure 23.

Code
knitr::include_graphics("images/weighting_zone_groups.png")
Figure 23: Weighting Zone Groups Used in the 2025 Puget Sound Regional Council Household Travel Study. Gray outlines are Public Use Microdata Areas (PUMAs) used to construct the weighting zones.

Targets

Targets are the control totals used to calibrate the survey weights so that the weighted survey matches the known population across the study area. For the 2025 Puget Sound Regional Council Household Travel Study, these controls were developed from ACS 1-year target estimates and ACS 1-year PUMS microdata, then aggregated from 2022 PUMA geography to the project’s five weighting zone groups shown in Figure 23:

  • King County - Seattle
  • King County - Other
  • Kitsap County - Expanded
  • Pierce County
  • Snohomish County

The target preparation process established both overall household and person totals and a set of household-level and person-level marginal distributions for each weighting zone group. These grouped-PUMA targets were then used as the control tables in RSG’s weighting engine, PopulationSim.

At the highest level, the weighting process was constrained to match the total number of households and total number of persons in each weighting zone group. The target documentation reports the following regional totals across all weighting zone groups combined:

  • Total households: approximately 1,745,353
  • Total persons: approximately 4,243,682

Within those totals, PSRC’s weighting controls included the following target variables and categories.

Household-level targets

  • Household size: 1 person, 2 persons, 3 persons, 4 or more persons
  • Household income: less than $25,000; $25,000-$49,999; $50,000-$74,999; $75,000-$99,999; $100,000-$199,999; $200,000 or more
  • Household workers: 0 workers, 1 worker, 2 or more workers
  • Vehicle availability / sufficiency: 0 vehicles; fewer vehicles than workers; more vehicles than workers
  • Presence of children: no children, 1 or more children
  • Total households

These controls align the weighted survey with the household structure of each weighting zone group, not just the total number of households.

Person-level targets

  • Gender: male, female
  • Age: 0-4, 5-15, 16-17, 18-24, 25-44, 45-64, 65 or older
  • Employment status: non-worker, part-time, full-time
  • Commute mode: work from home, transit, walk, bike, other (including auto), none
  • University student status: yes, no
  • Educational attainment: no college, some college
  • Race: White; Black or African American; Asian or Pacific Islander; other
  • Ethnicity: Hispanic, not Hispanic
  • Total persons

These controls align the weighted person file with the known population profile of each weighting zone group across major demographic dimensions.

In practice, PSRC’s weighting controls were designed to do two things simultaneously:

  • match the total households and total persons in each weighting zone group; and
  • match the marginal distributions of these household and person characteristics within each weighting zone group.

Some target categories were simplified, combined, or selectively applied to maintain stable estimation in smaller geographies and to avoid over-constraining the weighting process. Analysts should therefore interpret these categories as the effective levels at which the survey was calibrated to known population totals.

TipAnalyst Tip: Interpreting Weighting Zone Geography

The practical implication is important for analysis. If an analyst summarizes the data to these same five weighting zone groups, the weighted household and person totals and the controlled marginal distributions will add up as intended. If an analyst instead summarizes to a geography that cuts across the grouped PUMAs, such as a city boundary or another custom area that does not nest inside the weighting zones, the weighted totals do not need to line up exactly with external benchmarks for that geography.

This does not limit the usefulness of the data for sub-regional analysis, but it does mean that analysts should be cautious when interpreting weighted totals for geographies – especially small geographies – that do not align with the weighting zones. In particular, analysts should check the unweighted sample size and the distribution of weights within those geographies to understand how representative the estimates are likely to be. Analysis at the County level is likely to be more stable than analysis at the city level, and analysis at the city level is likely to be more stable than analysis at the neighborhood level, but this will depend on the specific geography and the distribution of the sample and weights within it.

2025 Puget Sound Regional Council Household Travel Study Combined Weighting Targets

The categories listed above summarize the household- and person-level controls used in weighting. Analysts can use them as a quick reference for the dimensions and levels at which the weighted survey was calibrated to known population totals across the five weighting zone groups.

TipAnalyst Tip: What Weights Can and Cannot Correct

Weighting targets define the population dimensions used to calibrate the survey. These controls make the marginal distributions (e.g., age, gender, income groups) in the weighted data match known population totals. However, there are important limitations:

1. Weighting improves representativeness only within defined categories.
Estimates are most reliable at the level of the weighting targets. More detailed breakdowns (e.g., finer income bins) were not controlled and may still reflect sampling variability or bias. In practice, targets define the finest level of safe aggregation.

2. Joint distributions are not guaranteed to match the population.
Weights align individual targets, not combinations of them. For example, age and race may each match population totals, but age x race may still be misrepresented. Be cautious with highly disaggregated cross-tabulations.

3. Non-targeted variables and small cells may be unstable.
Variables not included in weighting controls are not explicitly bias-corrected. Small or sparse groups remain unstable after weighting, especially when weights are large or variable. We recommend checking cell sizes and relative standard errors (RSEs, see Section 8.8.5) before interpreting results, especially when sample sizes are small.

4. Weighting does not correct measurement error.
Targets adjust who is represented, not what was reported. Misreporting or limitations in survey design (e.g., coarse mode categories) are not fixed through weighting.

Other useful diagnostics include the effective sample size, which reflects the equivalent number of equally weighted observations, and the design effect, which captures how weighting inflates variance (see Section 5.6.3 below).

Bottom line:
Weighting improves representativeness along specific dimensions, but it does not guarantee reliable estimates for all subgroups. Use targets as a guide to where estimates are most trustworthy.

5.4 Weighting Process

Base Weights

Weighting begins with base weights, which reflect the probability that a household was included in the survey. For each sample segment, RSG calculated a base weight as the inverse of the probability of inclusion, which depends on both the probability of selection and the probability of response. For segment s with H total households and R responding households, the base weight is:

\[ w_{s} = \frac{H_s}{R_s} \]

Base weights provide the initial expansion from the sample to the population and serve as the seed weights for subsequent rounds of weighting adjustments (see below).

Round 1 Weighting: Adjusting for Demographic Bias

Round 1 and 2 of weighting use PopulationSim to adjust base weights so that weighted survey estimates match demographic control totals derived from PUMS. PopulationSim performs constrained entropy maximization, adjusting the household weights in the smallest way necessary to match a set of household- and person-level targets.

Reference: Paul et al. (2018), PopulationSim Technical Paper (PDF)

NoteWhat is Entropy Maximization?

Entropy maximization is a statistical method used to adjust survey weights so that the weighted survey data matches known population totals (such as the number of households, adults, workers, or children in a region). It is widely used in official statistics, including by the U.S. Census Bureau.

The key idea is simple: Change the initial (base) weights as little as possible while forcing the final weighted totals to match external control totals.

Think of it as a balancing procedure:

  • Base weights reflect who responded to the survey and how the sampling effort was stratified.
  • We also have population counts from sources like the ACS.
  • Entropy maximization finds the smallest set of adjustments to the base weights that makes the survey line up with the known population.

This approach helps support a weighting process where:

  • Groups that were underrepresented in the sample get slightly higher weights
  • Groups that were overrepresented get slightly lower weights
  • The final weights stay close to their starting values, avoiding extreme or unstable adjustments

The result is a set of survey weights that preserves the structure of the collected data while helping the survey reflect the true population.

During Round 1, target categories were refined to improve stability, particularly in smaller geographies. For example, household size was top-coded at four or more persons, and person age categories were consolidated for children. Vehicle sufficiency, worker counts, income, educational attainment, and race/ethnicity were also included as targets.

PopulationSim constraints on minimum and maximum expansion factors and absolute weight bounds were tuned iteratively. The final configuration used a maximum expansion factor of seven, a minimum expansion factor of 0.1, and a cap of 700. This balance provided a strong fit to targets while maintaining stable weight distributions.

The output of Round 1 consists of:

  • Round 1 household weights, which align with demographic targets.
  • Round 1 person weights, created by assigning household weights to related household members and redistributing weights away from unrelated members.
  • Round 1 day weights, created by dividing each person weight across the number of complete diary days.

Round 1 weights are used as inputs to the day-pattern model in Round 2.

Round 2 Weighting: Adjusting for Day-Pattern Bias

Survey trip rates differ across diary platforms, in part because smartphone app users report more complete travel. To correct for this, RSG estimated a multinomial logit model that predicts whether a person-day is a no-travel day, a mandatory-travel day, or a non-mandatory-travel day. The model was estimated using weighted Round 1 day-level data and demographic predictors.

The model was then used to predict, for each person-day, the probability of each day-pattern as if all respondents had used the smartphone app. These predicted totals were summed within each weighting zone group and supplied as additional control totals for Round 2 PopulationSim.

Round 2 applies PopulationSim again, using the same demographic targets as Round 1 but now also requiring the weighted data to match these day-pattern targets. This step produces:

  • Final household weights

Adjusting Person and Day Weights

Next, person and day weights are adjusted for non-surveyable household members and incomplete diary days. Person weights are adjusted by redistributing weights from non-surveyable members (e.g., roommates) to surveyable members within the same household. Day weights are adjusted by dividing each person weight across only the complete diary days for that person. This step produces:

  • Final person weights
  • Final day weights

The final day weights sum to the total population and reflect one average weekday per person.

Round 3 Weighting: Adjusting for Trip-Type Reporting Bias

The final weighting step corrects for under-reporting of specific trip types across diary platforms. Trip records in hts$trip were grouped into work, school, and other trip categories. For each trip type, RSG estimated a weighted Poisson regression model predicting the number of linked person-trips per person-day. The models controlled for diary platform and demographic variables.

Two sets of predicted trip rates were generated: one using the fitted diary-platform effects and another with diary-platform effects removed. The ratio of these predictions forms the trip-type adjustment factor. These factors were scaled so that the minimum factor equals one, reflecting an assumption that differences across platforms represent under-reporting rather than over-reporting.

The adjustment factor is applied to each record in the delivered trip table, using the Round 2 day weight as the starting point, to produce the final trip weights.

These trip weights expand the survey to represent weekday trip totals.

%%{init: {"theme":"default","flowchart":{"htmlLabels":true}}}%%

flowchart TD

  %% Nodes
  Survey[("Survey Data<br/>(households, persons, days, trips)")]
  Census[("Census Data<br/>(ACS 5-yr + PUMS)")]
  Targets[["Census Targets"]]
  Base{"Base Weight Estimation"}
  DP{"Day Pattern Modeling"}
  P1{"Round 1: Demographic Nonresponse Weighting"}
  P2{"Round 2: Day-Pattern Weighting"}
  TripAdj{"Round 3: Trip Weight Adjustment"}
  WeightedSeed[["Base Weights<br/>(Weighted Seed)"]]
  R1[/"Round 1 Weights<br/>(HH, Person, Day)"/]
  DayTargets[["Day-pattern Targets"]]
  R2[/"Round 2 Weights<br/>(Final Household, Person, Day Weights)"/]
  FinalWeights[/"Round 3 Weights<br/>(Final Trip Weights)"/]

  %% Edges
  Survey --> Base
  Survey --> DP
  Survey --> TripAdj
  Census --> Targets
  Census --> Base
  Targets --> P1
  Targets --> P2
  Base --> WeightedSeed
  WeightedSeed --> P1
  WeightedSeed --> P2
  P1 --> R1
  R1 --> DP
  DP --> DayTargets
  DayTargets --> P2
  P2 --> R2
  R2 --> TripAdj
  TripAdj --> FinalWeights

  %% Styling
  classDef data fill:#005753,stroke:#005753,color:#ffffff,font-family:Inter,stroke-width:2px
  classDef proc fill:#A3D063,stroke:#A3D063,color:#000,font-family:Inter,stroke-width:2px
  classDef weight fill:#55B5B0,stroke:#55B5B0,color:#fff,font-family:Inter,stroke-width:2px

  class Survey,Census,Targets,WeightedSeed,DayTargets data
  class Base,DP,TripAdj,P1,P2 proc
  class R1,R2,FinalWeights weight

5.5 Weighted Totals

The final weights expand the survey to the following population totals:

  • Households: 1,750,591 households across the central Puget Sound region.
  • Persons: 4,230,763 persons.
  • Person-days: 4,230,763 person-days.
  • Trips: 17,100,397 weekday linked person-trips from the delivered trip table.

These totals are constructed so that analysts can directly compute estimates using the corresponding weight type.

This chapter summarizes the weighting process used for this guide. Refer to the delivered project weighting documentation when additional methodological detail is needed.

5.6 Additional Guidance for Analysts

Choosing the Right Weight

Different analyses require different weight types. Analysts should select the weight that matches the level of measurement:

  • Household weights should be used when households are the unit of analysis or when studying household-level characteristics (e.g., vehicles, income, housing type).
  • Person weights should be used for demographic characteristics, person-level behaviors, and analyses where individuals, not days or trips, are the unit.
  • Day weights should be used when analyzing travel made on a given weekday, including day patterns, trip rates, and average daily travel.
  • Trip weights should be used when analyzing complete trips from origin to destination.

Using the wrong weight type can lead to biased estimates. For example, applying person weights to trip tables will underestimate total travel. For more information, see Section 8.3 in the Analyst Handbook.

What the Weights Can and Cannot Correct

The weighting process corrects for several forms of bias:

  • Differences in sampling likelihood across geographies
  • Differential response rates across demographic groups
  • Reporting differences across diary platforms (e.g., smartphone vs. web)
  • Under-reporting of specific trip types

However, weighting cannot correct for:

  • Misreported or miscoded trip purposes
  • Missing data not captured through imputation
  • Recall errors unrelated to diary platform
  • GPS or map-matching errors
  • Weekends or seasons not included in the survey period

Analysts should interpret highly granular or rare-behavior results with caution.

Design Effects and Effective Sample Size

Unequal weights reduce the statistical precision of estimates compared with a simple random sample of the same size. This reduction is summarized by the design effect (DEFF) and the effective sample size (ESS). DEFF reflects how much weight variability inflates variance; ESS reflects the size of an unweighted sample that would yield equivalent precision. As a rule of thumb, when DEFF exceeds 2.0, analysts should expect a noticeable loss of precision—particularly when estimates are based on small subgroups, where limited sample size and weight variability compound.

The 2025 Puget Sound Regional Council Household Travel Study dataset shows moderate design effects regionwide and larger effects in small counties. A condensed summary is shown below in Table 27. While DEFF values exceed 2.0 across all geographies, this does not mean that all estimates are unreliable. For the full region and larger counties, the effective sample sizes remain sufficiently large to support stable estimates for many common analyses. However, in smaller geographies—or when results are further segmented (e.g., by mode, income, or demographic group)—the combination of higher DEFF and smaller sample sizes can lead to increased uncertainty. Analysts should use caution when interpreting highly disaggregated results and consider reporting standard errors or relative standard errors where possible.

Unweighted N Effective Sample Size Design Effect
King County 1,142 466 2.45
Kitsap County 134 29 4.56
Pierce County 1,182 206 5.74
Snohomish County 314 68 4.59
Total 2,772 673 4.12
Table 27: Design Effect and Effective Sample Size Summary

Distribution of Weights

Weight variability differs by dataset level. In general:

  • Household and person weights range from small values (approximately 5-10) to capped values in the hundreds.
  • Day weights have more variability because weights are divided across multiple diary days.
  • Trip weights inherit day-weight variability and incorporate trip-type adjustments, producing the widest distributions.
Figure 24: Weight Distributions by Dataset Level

Analysts should be cautious when conducting analyses in which a small number of high-weight observations dominate the estimates.

Geographic Considerations and Small-Area Estimates

Because weighting was performed to the five grouped-PUMA weighting zone groups, those are the geographies at which the weighted data is internally consistent. This means an analyst can:

  • Analyze any weighting zone group independently, with weighted household and population totals expected to align with that zone group’s control totals.
  • Report weighting-zone-group statistics directly, since the weights were calibrated at that geography.

However, analysts should note the following:

  • Cities, towns, neighborhoods, counties, and other analyst-defined geographies are not individually weighted unless they align with the weighting zone groups. Weighted totals will match the weighting controls only where those geographies nest within the grouped-PUMA zones.
  • Any analysis for sub geographies must account for the full survey design. This includes understanding sampling variability, weight variability, and the fact that the weighting was not optimized for those small domains.
  • For fine-scale estimates, uncertainty will generally be larger, and in some cases pooling multiple areas or using model-based methods may be necessary.
TipAnalyst Tip: Can I Analyze My City?

Because the survey was weighted to grouped-PUMA weighting zones, those five weighting zone groups can be treated as the primary standalone geographies for weighted household- and person-level summaries. Within those zones, weighted totals for households and persons are expected to align with the corresponding weighting controls.

However, cities, towns, neighborhoods, counties, and other custom areas were not individually weighted unless they align with the grouped-PUMA zones. Their weighted totals will not necessarily match true local population counts, and small-area estimates may have substantial sampling and weighting variability. Analysts working at fine geographic scales should interpret results carefully and consider pooling areas or using model-based approaches when possible.

Small Population Groups

Similarly, rare population groups—such as university students, young children, zero-vehicle households, and active transportation commuters—may have limited representation in the weighted data and may not be stable in their raw form.

For these groups, consider:

  • Pooling response categories (e.g., household size, age groups)
  • Pooling across geographic sub-sets (when conceptually appropriate)
  • Using model-based estimation techniques
  • Reporting confidence intervals where possible

For worked examples and additional guidance, see Section 8.8.6 in the Analyst Handbook.

5.7 Summary

Weighting for the 2025 Puget Sound Regional Council Household Travel Study dataset follows a structured and incremental process. Base weights correct for sample design. Round 1 adjustments correct for demographic nonresponse. Round 2 adjustments correct for day-pattern reporting bias. Round 3 adjustments correct for trip-type under-reporting. Together, these steps yield household, person, day, and trip weights for estimating population-level travel behavior.

6 Dataset Overview

6.1 Data Hierarchy

The PSRC data is hierarchical in nature. Data are disaggregated at the household-, person-, vehicle-, day-, trip-, and location-level. The delivered tables link to one another using the IDs present in this project: household_id, person_id, day_id, trip_id, and vehicle_id. Figure 25 shows the relationship among the delivered tables.

Figure 25: Data Linkages Across Delivered Tables

6.2 Summary of Data Tables

Table Name Record Unit Primary ID(s) What’s in the Table
Household (hh) One row per household household_id Household demographics, income, home location, vehicle ownership, sample segment, travel diary mode, household-level weights.
Person (person) One row per person person_id, household_id Age, gender, race/ethnicity, employment, student status, relationship to householder, travel diary participation, person-level weights.
Day (day) One row per person-day day_id, person_id, household_id Diary date, completion status, daily school attendance and telework records, deliveries, day-level weights.
Trip (trip) One row per person-trip trip_id, day_id, person_id, household_id Start/end time, locations, main mode, purpose, distance, duration, and trip weights.
Vehicle (vehicle) One row per household vehicle vehicle_id, household_id Vehicle attributes for household vehicles.
Location (location) One row per GPS ping Several rows per trip_id Accuracy, timestamp, latitude, longitude, altitude, speed, heading. Unweighted.

6.3 Trip Unit of Measure: Person-Trips

Understanding how individual travel events are represented in the dataset is essential for correctly interpreting the delivered trip table. This section describes the structure of person-trip records, how shared travel is captured, and the conceptual implications for analyses conducted elsewhere in the guide.

What is a Person-Trip?

The core trip-level dataset in the study is trip, constructed at the person-trip level. In this delivery, the trip table is a linked-trip product. This means:

  • Each row represents a single travel event made by a single person.
  • Transit trips have one record per origin-to-destination movement, even if they include multiple modes or transfers. Access and egress legs are consolidated into the main trip record.
  • If two or more household members traveled together, the shared movement appears in the data multiple times, once for each participating household member.
  • Because the unit of observation is the individual traveler, these tables intentionally do not represent vehicle trips, group trips, or household trips.

Replication of Shared Trips

When household members travel together—such as carpooling, walking together, or biking as a group—the dataset includes:

  • One person-trip record per traveler, and
  • A unique trip_id assigned to each person-trip, even if those trips share identical or near-identical characteristics.

This replication is expected and is a direct result of designing the data structure around person-based travel diaries. Shared travel behavior is therefore represented as multiple observations of the same physical movement, differentiated by the traveler.

It is common to observe:

  • Identical origin/destination coordinates
  • Nearly identical start/end times
  • Matching modes
  • Matching purposes

across members of the same household.

These patterns indicate shared travel, not data duplication errors.

6.4 Record Counts

The final unweighted dataset includes six distinct data tables. These tables include all user-input study variables, certain study metadata, and variables derived to support data analysis.

Data Table Number of Records Weighted Records Percent of Records Weighted
hh 2,772 2,772 100.0%
person 5,558 5,558 100.0%
day 10,868 7,652 70.4%
trip 38,980 26,119 67.0%
Table 28: Record Counts by Data Table

6.5 Data Types and Considerations

The dataset includes four main types of variables: categorical, continuous numeric, time, and location. Understanding these types is essential for proper analysis, visualization, and interpretation.

Categorical Variables

Categorical variables store labels rather than magnitudes. They include:

  • binary fields (e.g., can_drive yes/no),
  • nominal multi category fields with no inherent order (e.g., mode_class_5 = Drive, Transit, Bike, Walk)
  • ordinal fields with a natural order (e.g., hhincome_detailed, education).
  • count variables (e.g., vehicle_count, hhsize), which represent discrete quantities capped at a maximum (e.g., “13 or more people”).

Some sets of categorical variables are grouped. These are:

  • multiple response categorical variables (MRCVs) for “select all that apply” questions (e.g., delivery and race/ethnicity), which appear in the data as a group of Selected/Not selected indicators (e.g., deliver_*, race_*, and ethnicity_*).

Missing entries are represented as Missing Response.

NoteUse the Codebook to Order Categories

When you build a table, chart, or derived factor from a categorical variable, use Section 7.3 as the source of truth for both labels and ordering. The value-label rows preserve the intended category sequence from values.rds, so analysts do not need to hand-maintain factor levels or rely on alphabetical ordering from observed data.

Continuous Numeric Variables

Continuous, numeric variables represent numeric measures where arithmetic operations are meaningful. In travel surveys these include trip level metrics like:

  • distance (mi/km),
  • duration (minutes)
  • cost (currency),
  • derived quantities such as speed.

Values may be integers or decimals and usually carry explicit units. Ranges can be wide and may include extreme values (e.g., very long trips). Missing entries are represented as NULL.

NoteTop-Coded Count variables

Note that some variables including age, hhsize, vehicle_count, and hhincome_detailed, though numeric in nature, are categorical in practice with their discrete quantities capped at a maximum (treated as binned/bracketed values).

Outliers in the Dataset

Table 29 summarizes outlier diagnostics for all numeric variables in the HTS dataset. Outliers are defined using the interquartile range (IQR) method, where values below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR are flagged as outliers.

We recommend strategies for dealing with outliers in Section 8.

Code
## ----------------------------------------------------------
## 0. Prep: use a fixed list of variables
## ----------------------------------------------------------

# trip_del <- readRDS("data/Dataset_2025-10-21/ex_trip_unlinked.rds")

table_cols <- intersect(c("hh", "person", "day", "trip", "vehicle"), names(hts))
target_vars <- c(
  "num_trips",
  "age",
  "distance_meters",
  "distance_miles",
  # "distance_beeline_meters", <!--FIXME: PSRC: add to dataset and re-run diagnostics. -->
  "duration_minutes",
  "duration_seconds",
  "dwell_mins",
  "speed_mph",
  "travel_time"
)

## ----------------------------------------------------------
## 1. Find those variables in the core HTS tables
## ----------------------------------------------------------

numeric_vars_long <- rbindlist(
  lapply(table_cols, function(tbl_name) {
    dt <- hts[[tbl_name]]
    present_vars <- intersect(target_vars, names(dt))

    if (length(present_vars) == 0) {
      return(NULL)
    }

    data.table(
      variable = present_vars,
      table_name = tbl_name,
      storage_type = vapply(dt[, ..present_vars], function(x) class(x)[1], character(1))
    )
  }),
  use.names = TRUE,
  fill = TRUE
)

## ----------------------------------------------------------
## 2. Initialize outlier summary table
## ----------------------------------------------------------

summary_dt <- copy(numeric_vars_long)
summary_dt[, `:=`(
  min_value = NA_real_,
  max_value = NA_real_,
  p01 = NA_real_,
  p99 = NA_real_,
  iqr = NA_real_,
  lower_bound = NA_real_,
  upper_bound = NA_real_,
  n_outliers = NA_integer_,
  pct_outliers = NA_real_,
  has_outliers = NA,
  max_outlier_gap = NA_real_
)]

## ----------------------------------------------------------
## 3. Loop over (variable, table_name) and compute outlier diagnostics
## ----------------------------------------------------------

for (i in seq_len(nrow(summary_dt))) {
  var_name <- summary_dt$variable[i]
  tbl_name <- summary_dt$table_name[i]

  table_exists <- !is.null(hts[[tbl_name]])
  col_exists <- table_exists && var_name %in% names(hts[[tbl_name]])
  summary_dt[i, exists_in_table := col_exists]

  if (!col_exists) {
    next
  }

  this_vec <- hts[[tbl_name]][[var_name]]

  # Convert integer64 safely for summary statistics while leaving the
  # source object unchanged.
  all_values <- suppressWarnings(as.numeric(this_vec))
  all_values <- all_values[!is.na(all_values)]
  summary_dt[i, non_missing_n := length(all_values)]

  if (length(all_values) == 0) {
    next
  }

  ## --------------------------------------
  ## Compute basic stats and outliers
  ## --------------------------------------

  x_min <- min(all_values, na.rm = TRUE)
  x_max <- max(all_values, na.rm = TRUE)

  q01 <- as.numeric(quantile(all_values, 0.01, na.rm = TRUE, type = 7))
  q99 <- as.numeric(quantile(all_values, 0.99, na.rm = TRUE, type = 7))
  q1 <- as.numeric(quantile(all_values, 0.25, na.rm = TRUE, type = 7))
  q3 <- as.numeric(quantile(all_values, 0.75, na.rm = TRUE, type = 7))

  iqr_val <- q3 - q1

  lower <- q1 - 1.5 * iqr_val
  upper <- q3 + 1.5 * iqr_val

  is_outlier <- (all_values < lower) | (all_values > upper)
  n_out <- sum(is_outlier)
  pct_out <- n_out / length(all_values)

  below_vals <- all_values[all_values < lower]
  above_vals <- all_values[all_values > upper]

  if (length(below_vals) > 0) {
    max_below_gap <- lower - min(below_vals)
  } else {
    max_below_gap <- 0
  }

  if (length(above_vals) > 0) {
    max_above_gap <- max(above_vals) - upper
  } else {
    max_above_gap <- 0
  }

  max_gap <- max(max_below_gap, max_above_gap)

  summary_dt[
    i,
    `:=`(
      min_value = x_min,
      max_value = x_max,
      p01 = q01,
      p99 = q99,
      iqr = iqr_val,
      lower_bound = lower,
      upper_bound = upper,
      n_outliers = n_out,
      pct_outliers = pct_out,
      has_outliers = n_out > 0,
      max_outlier_gap = max_gap
    )
  ]
}

outlier_view <- copy(summary_dt)

outlier_view[
  ,
  severity := fcase(
    pct_outliers < 0.005 & max_outlier_gap < (5 * iqr),
    "Low",
    pct_outliers < 0.05 & max_outlier_gap < (25 * iqr),
    "Moderate",
    default = "High"
  )
]

outlier_view[
  ,
  recommended_action := fcase(
    severity == "Low",
    "No action needed",
    severity == "Moderate",
    "Consider trimming >= 99th pct.",
    severity == "High",
    "Trim or winsorize >= 95th pct."
  )
]

outlier_simple <- outlier_view[
  ,
  .(
    table = table_name,
    variable,
    storage_type,
    pct_outliers,
    max_outlier_gap,
    severity,
    recommended_action
  )
]

setorder(outlier_simple, table, variable)

library(gt)
library(scales)

outlier_simple_gt <-
  outlier_simple |>
  gt(rowname_col = "variable", groupname_col = "table") |>
  tab_header(
    title = "Outlier Diagnostics by Variable",
    subtitle = "Share and severity of IQR-based outliers"
  ) |>
  fmt_percent(
    columns = pct_outliers,
    decimals = 1
  ) |>
  fmt_number(
    columns = max_outlier_gap,
    decimals = 0,
    use_seps = TRUE
  ) |>
  data_color(
    columns = pct_outliers,
    method = "numeric",
    palette = sequential_teal_palette(4),
  ) %>%
  data_color(
    columns = severity,
    method = "factor",
    palette = c(
      "Low" = grDevices::adjustcolor(brand_colors[["green"]], alpha.f = 0.25),
      "Moderate" = grDevices::adjustcolor(brand_colors[["gold"]], alpha.f = 0.3),
      "High" = grDevices::adjustcolor(brand_colors[["orange"]], alpha.f = 0.25)
    )
  ) |>
  cols_label(
    storage_type = "Detected type",
    pct_outliers = "% of records flagged as outliers",
    max_outlier_gap = "Worst outlier beyond upper bound",
    severity = "Outlier severity",
    recommended_action = "Suggested handling"
  ) |>
  tab_options(
    table.font.size = px(12),
    data_row.padding = px(3)
  )

outlier_simple_gt
Outlier Diagnostics by Variable
Share and severity of IQR-based outliers
Detected type % of records flagged as outliers Worst outlier beyond upper bound Outlier severity Suggested handling
day
num_trips integer 4.3% 32 Moderate Consider trimming >= 99th pct.
hh
num_trips character 13.4% 272 High Trim or winsorize >= 95th pct.
person
age character NA NA High Trim or winsorize >= 95th pct.
num_trips integer 14.0% 128 High Trim or winsorize >= 95th pct.
trip
distance_meters numeric 9.7% 4,319,306 High Trim or winsorize >= 95th pct.
distance_miles numeric 9.7% 2,684 High Trim or winsorize >= 95th pct.
duration_minutes numeric 6.8% 1,234 High Trim or winsorize >= 95th pct.
duration_seconds numeric 6.8% 74,040 High Trim or winsorize >= 95th pct.
dwell_mins numeric 12.9% 7,997 High Trim or winsorize >= 95th pct.
speed_mph numeric 1.8% 102,892 High Trim or winsorize >= 95th pct.
travel_time numeric 6.8% 1,234 High Trim or winsorize >= 95th pct.
Table 29: Outlier Diagnostics by Variable

Missing Values

A study data table cell may be missing data for one of four reasons:

1. Value or response is missing due to survey logic, participant non-response, or error.

Example: Participants who traveled by bus were not asked if they were the driver or passenger on the trip.

Coded as: 995 (labeled as Missing Response) for categorical variables, NULL for continuous variables

2. A respondent indicated that the question was not applicable and skipped that question.

Example: Some participants did not share how they pay to park at work because they do not park at work (e.g., carpool).

Coded as: 996 (often labeled as “Not applicable”)

3. A respondent indicated that they didn’t know the answer and skipped that question.

Example: Some participants who took a taxi did not know the fare.

Coded as: “Don’t know”

4. A respondent indicated that they preferred not to answer a question and skipped that question.

Example: Some participants chose not to provide their household income.

Coded as: 999 (Prefer not to answer)

Other notes about missing study data:

  • Continuous variables (e.g., trip distance, trip duration) are not coded with missing value codes and are instead left empty (NULL) when missing to avoid interfering with statistical calculations.

  • Due to the large size of the location table, missing values were left exactly as they were collected. Speed, heading, and accuracy can all potentially contain missing values that are either stored as “-1”, NA, or 0. Analysis on those fields should filter to where the values are greater than zero.

7 Codebook

The codebook is the primary reference for understanding what each delivered PSRC variable means, which table it belongs to, how its values should be labeled, and how analysts should order categorical responses in reproducible work. The rest of this guide explains collection, processing, and weighting. This chapter explains the delivered data element by data element.

A strong codebook reduces guesswork. It lets analysts identify table membership quickly, distinguish derived fields from survey-response fields, verify valid values and skip logic, and build tables and plots in the intended category order without hand-maintained factor levels.

NoteUse the Codebook First

Start here whenever you need to answer any of these questions:

  • What table contains this variable?
  • Is this field categorical, numeric, or top-coded?
  • What order should categories appear in a plot or table?
  • Is this field part of a “Select-all-that-apply” (multiple-response categorical value, MRCV) family or controlled by survey logic?

7.1 What the Codebook Contains

Variable List

The variable list is the structural reference for the dataset. It shows:

  • variable name
  • tables that includes the variable across household, person, day, trip, and vehicle (0: table does not include the variable; 1: table includes the variable)
  • delivered data type (see Section 6.5)
  • description of the variable’s meaning, units, and (if applicable) how it was derived
  • survey logic that governs whether a respondent is asked the question or has a value for the variable, which is critical for understanding missing values and skip patterns in the data

Value Labels

The value-label table is the categorical reference for the dataset. It shows:

  • table name
  • variable name
  • human-readable label
  • category order (val_order)

7.2 Variable List

For display in this table, the delivered membership flags (hh, person, day, trip, vehicle) are combined into a single table_membership field. The underlying codebook data still retain the separate membership columns.

7.3 Value Labels

8 Analyst Handbook

This chapter provides practical, end-to-end guidance for analysts working with the 2025 Puget Sound Regional Council Household Travel Study dataset. It focuses on how to load, join, filter, weight, and analyze the data, with reproducible examples. It should serve as the primary resource for anyone conducting descriptive analyses, modeling, or statistical inference using these data.

8.1 Software Setup

This guide focuses on using R for data analysis. However, many of the principles and techniques discussed here can be applied in other statistical software packages such as Python, Stata, or SAS. To follow along with the examples, you will need to have R installed on your computer.

  • R: This guide was developed and tested using R version 4.5.3.
  • R GUI: Positron and RStudio both works well.

Next, install and load the necessary R packages. We use:

  • srvyr for survey-weighted analysis, a tidyverse-friendly wrapper around the survey package
  • tidyr and dplyr for data reshaping and data manipulation that pairs with srvyr
  • ggplot2 and gt for visualization and table generation, respectively
  • stringr for string manipulation and detection
  • lubridate for date/time handling
  • tidycensus and sf for fetching and preparing Census block-group density inputs

These will need to be installed if you don’t have them already. You can install them using install.packages().

Code
suppressPackageStartupMessages({
  library(dplyr)
  library(tidyr)
  library(srvyr)
  library(ggplot2)
  library(plotly)
  library(gt)
  library(stringr)
  library(lubridate)
})

brand_colors <- c(
  anchor = "#005753",
  teal = "#55B5B0",
  green = "#A3D063",
  gold = "#F0B356",
  orange = "#F05A28",
  purple = "#92278F",
  gray = "#7A8781",
  ink = "#25343D",
  surface = "#F6FAFB",
  line = "#D6E2E8"
)

single_bar_color <- brand_colors[["anchor"]]
subtitle_color <- brand_colors[["gray"]]

8.2 Getting Started with the Data

This section introduces the essential steps for preparing the 2025 Puget Sound Regional Council Household Travel Study dataset for analysis, from loading files to filtering data and linking tables.

Load Data

To load data from PSRC’s data catalog, use the arcgislayers package to read directly from the ArcGIS REST API endpoints. The code snippet below demonstrates how to load each of the key tables (households, persons, trips, vehicles, days) into R using their respective API URLs.

The HTS data tables and codebook are also available for download on the PSRC data portal.

Code
library(arcgislayers)

hts_urls <- list(
  households = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Households/FeatureServer/0",
  persons    = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Persons/FeatureServer/0",
  trips      = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Trips/FeatureServer/0",
  vehicles   = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Vehicles/FeatureServer/0",
  days       = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Days/FeatureServer/0"
)

# get 2025 data from psrc data portal
hts_data <- lapply(hts_urls, function(url) {
  layer_conn <- arc_open(url)
  arc_select(layer_conn, where = "survey_year = 2025")
})

cb_urls <- list(
  value_labels  = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Variables/FeatureServer/0",
  variable_list = "https://services6.arcgis.com/GWxg6t7KXELn1thE/arcgis/rest/services/Household_Travel_Survey_Values/FeatureServer/0"
)

cb <- lapply(cb_urls, arc_read)

# names(hts_data)

Inspect the Data

Once the data is loaded, you can inspect the key tables to understand their structure and contents.

Get List of Tables

Print the names of the loaded tables:

Code
names(hts)
## [1] "hh"      "person"  "day"     "trip"    "vehicle"

Glimpse Data

Each table has a mix of variable types, includes one or more ID columns, and a weight. View the structure of the person table:

Code
glimpse(hts$person)
## Rows: 5,558
## Columns: 144
## $ person_id                         <int64> 2500000601, 2500000602, 2500007101…
## $ survey_year                       <int> 2025, 2025, 2025, 2025, 2025, 2025, …
## $ ethnicity_other                   <chr> "Missing Response", "Missing Respons…
## $ household_id                      <int64> 25000006, 25000006, 25000071, 2500…
## $ industry_other                    <chr> "Missing Response", "Missing Respons…
## $ can_drive                         <chr> "Yes", "Yes", "Yes", "Yes", "Yes", "…
## $ num_trips                         <int> 0, 0, 2, 9, 0, 5, 0, 2, 2, 0, 3, 3, …
## $ numdayscomplete                   <chr> "1", "1", "1", "1", "1", "1", "1", "…
## $ pernum                            <int> 1, 2, 1, 1, 2, 3, 1, 2, 1, 2, 1, 2, …
## $ race_other_specify                <chr> "Missing Response", "1/2 Caucasian, …
## $ school_bg                         <chr> "Missing Response", "Missing Respons…
## $ school_loc_lat                    <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_loc_lng                    <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_puma10                     <chr> "Missing Response", "Missing Respons…
## $ school_rgcname                    <chr> "Missing Response", "Missing Respons…
## $ school_jurisdiction               <chr> "Missing Response", "Missing Respons…
## $ school_county                     <chr> "Missing Response", "Missing Respons…
## $ school_tract_2020                 <int64> NA, NA, NA, NA, NA, NA, NA, NA, NA…
## $ second_home_lat                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ second_home_lon                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ work_bg                           <chr> "Missing Response", "Missing Respons…
## $ work_lat                          <dbl> NA, NA, NA, 47.1646, NA, 47.1569, NA…
## $ work_lng                          <dbl> NA, NA, NA, -122.691, NA, -122.690, …
## $ work_puma10                       <chr> "Missing Response", "Missing Respons…
## $ work_rgcname                      <chr> "Missing Response", "Missing Respons…
## $ work_jurisdiction                 <chr> "Missing Response", "Missing Respons…
## $ work_county                       <chr> "Missing Response", "Missing Respons…
## $ work_tract_2020                   <int64> NA, NA, NA, 53053072603, NA, 53053…
## $ bike_freq                         <chr> "Missing Response", "Missing Respons…
## $ carshare_freq                     <chr> "Missing Response", "Missing Respons…
## $ commute_freq                      <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_1                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_2                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_3                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_4                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_5                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_6                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_7                 <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_996               <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_998               <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_1             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_2             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_3             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_4             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_5             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_6             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_7             <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_use_996           <chr> "Missing Response", "Missing Respons…
## $ disability_person                 <chr> "No", "No", "No", "No", "Yes", "Yes"…
## $ education                         <chr> "Some college", "Vocational/technica…
## $ employment                        <chr> "Not employed for pay (e.g., unpaid …
## $ ethnicity_1                       <chr> "Selected", "Selected", "Selected", …
## $ ethnicity_2                       <chr> "Not selected", "Not selected", "Not…
## $ ethnicity_3                       <chr> "Not selected", "Not selected", "Not…
## $ ethnicity_4                       <chr> "Not selected", "Not selected", "Not…
## $ ethnicity_997                     <chr> "Not selected", "Not selected", "Not…
## $ ethnicity_999                     <chr> "Not selected", "Not selected", "Not…
## $ race_hisp                         <chr> "Not Selected", "Not Selected", "Not…
## $ ev_typical_charge_1               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_2               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_3               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_4               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_5               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_6               <chr> "Missing Response", "Missing Respons…
## $ ev_typical_charge_997             <chr> "Missing Response", "Missing Respons…
## $ gender                            <chr> "Boy/Man (cisgender or transgender)"…
## $ hours_work                        <chr> "Missing Response", "Missing Respons…
## $ industry                          <chr> "Missing Response", "Missing Respons…
## $ is_participant                    <chr> "Yes", "Yes", "Yes", "Yes", "Yes", "…
## $ mobility_aides                    <chr> "Missing Response", "Missing Respons…
## $ office_available                  <chr> "Missing Response", "Missing Respons…
## $ participate                       <chr> "Yes", "No", "Yes", "Yes", "Yes", "Y…
## $ person_is_complete                <chr> "Yes", "Yes", "Yes", "Yes", "Yes", "…
## $ proxy                             <chr> "No", "No", "No", "No", "No", "No", …
## $ proxy_parent                      <chr> "No", "No", "No", "No", "No", "No", …
## $ race_afam                         <chr> "Not Selected", "Not Selected", "Not…
## $ race_aiak                         <chr> "Not Selected", "Not Selected", "Not…
## $ race_asian                        <chr> "Not Selected", "Not Selected", "Not…
## $ race_hapi                         <chr> "Not Selected", "Not Selected", "Not…
## $ race_noanswer                     <chr> "Not Selected", "Not Selected", "Not…
## $ race_other                        <chr> "Not Selected", "Selected", "Not Sel…
## $ race_white                        <chr> "Selected", "Not Selected", "Selecte…
## $ race_category                     <chr> "White non-Hispanic", "Some Other Ra…
## $ relationship                      <chr> "Self", "Spouse or partner", "Self",…
## $ remote_class_freq                 <chr> "Missing Response", "Missing Respons…
## $ school_freq                       <chr> "Missing Response", "Missing Respons…
## $ school_in_region                  <chr> "No", "No", "No", "No", "No", "No", …
## $ school_mode_typical               <chr> "Missing Response", "Missing Respons…
## $ schooltype                        <chr> "Missing Response", "Missing Respons…
## $ second_home                       <chr> "Does not regularly spend night at s…
## $ second_home_in_region             <chr> "Missing Response", "Missing Respons…
## $ sexuality                         <chr> "Heterosexual (straight)", "Heterose…
## $ share_1                           <chr> "Selected", "Selected", "Selected", …
## $ share_2                           <chr> "Not selected", "Not selected", "Not…
## $ share_3                           <chr> "Not selected", "Not selected", "Not…
## $ share_4                           <chr> "Not selected", "Not selected", "Not…
## $ share_5                           <chr> "Not selected", "Not selected", "Not…
## $ share_996                         <chr> "Not selected", "Not selected", "Not…
## $ smartphone_type                   <chr> "Has an Android phone", "Has an Andr…
## $ adult_student                     <chr> "No, not a student", "No, not a stud…
## $ surveyable                        <chr> "Yes", "Yes", "Yes", "Yes", "Yes", "…
## $ telecommute_freq                  <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_parking           <chr> "Missing Response", "Missing Respons…
## $ commute_subsidy_transit           <chr> "Missing Response", "Missing Respons…
## $ tnc_freq                          <chr> "Missing Response", "Missing Respons…
## $ transit_freq                      <chr> "Missing Response", "Missing Respons…
## $ walk_freq                         <chr> "2 days a week", "2 days a week", "6…
## $ work_in_region                    <chr> "No", "No", "No", "Yes", "No", "Yes"…
## $ work_mode                         <chr> "Missing Response", "Missing Respons…
## $ workplace                         <chr> "Missing Response", "Missing Respons…
## $ age                               <chr> "65-74 years", "65-74 years", "75-84…
## $ age_detailed                      <chr> "72 years old", "71 years old", "75 …
## $ drive_for_work                    <chr> "Missing Response", "Missing Respons…
## $ paid_work                         <chr> "No", "No", "No", "Yes", "No", "Yes"…
## $ return_from_leave                 <chr> "Missing Response", "Missing Respons…
## $ transportation_statement_car      <chr> "4-Agree", "Missing Response", "5-St…
## $ transportation_statement_environ  <chr> "1-Strongly disagree", "Missing Resp…
## $ transportation_statement_retail   <chr> "4-Agree", "Missing Response", "3-Ne…
## $ transportation_statement_telework <chr> "Missing Response", "Missing Respons…
## $ transportation_statement_transit  <chr> "4-Agree", "Missing Response", "2-Di…
## $ transportation_statement_travel   <chr> "4-Agree", "Missing Response", "2-Di…
## $ transportation_statement_walk     <chr> "2-Disagree", "Missing Response", "3…
## $ work_from_home                    <chr> "Missing Response", "Missing Respons…
## $ work_hh_1                         <dbl> NA, NA, NA, 103.69945, NA, 109.21502…
## $ work_hh_2                         <dbl> NA, NA, NA, 319.7172, NA, 320.4932, …
## $ work_emptot_1                     <dbl> NA, NA, NA, 1.396064, NA, 9.227058, …
## $ work_emptot_2                     <dbl> NA, NA, NA, 14.32801, NA, 18.41387, …
## $ school_hh_1                       <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_hh_2                       <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_emptot_1                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_emptot_2                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes1_1                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes3_1                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes4_1                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes1_2                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes3_2                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ school_nodes4_2                   <dbl> NA, NA, NA, NA, NA, NA, NA, NA, NA, …
## $ work_nodes1_1                     <dbl> NA, NA, NA, 0, NA, 0, NA, NA, 0, NA,…
## $ work_nodes3_1                     <dbl> NA, NA, NA, 0, NA, 0, NA, NA, 0, NA,…
## $ work_nodes4_1                     <dbl> NA, NA, NA, 20.146138, NA, 19.267898…
## $ work_nodes1_2                     <dbl> NA, NA, NA, 0, NA, 0, NA, NA, 0, NA,…
## $ work_nodes3_2                     <dbl> NA, NA, NA, 0, NA, 0, NA, NA, 0, NA,…
## $ work_nodes4_2                     <dbl> NA, NA, NA, 57.00128, NA, 55.18191, …
## $ person_weight                     <dbl> 22.80259, 22.80259, 75.94698, 25.277…

View Table Dimensions

Each table has a different number of rows and columns, reflecting the hierarchical nature of the data (households contain one or more persons, persons have one or more days, days have zero or more trips, etc.). You can check the dimensions of each table using the dim() function.

Code
sapply(hts, dim)
##        hh person   day  trip vehicle
## [1,] 2772   5558 10868 38980   29323
## [2,]   86    144    52   112       9

View Sample Records

You can view the first few records of each table using the head() function.

Code
lapply(hts, head)
## $hh
##    household_id survey_year diary_platform              pierce_uninc_bg
##           <i64>       <int>         <char>                       <char>
## 1:     25000006        2025        browser Pierce County unincorporated
## 2:     25000071        2025        browser Pierce County unincorporated
## 3:     25000096        2025        browser Pierce County unincorporated
## 4:     25000152        2025        browser Pierce County unincorporated
## 5:     25000200        2025        browser Pierce County unincorporated
## 6:     25000255        2025        browser Pierce County unincorporated
##    num_participants num_trips prev_home_lat prev_home_lng prev_home_rgcname
##               <int>    <char>         <num>         <num>            <char>
## 1:                2         0            NA            NA  Missing Response
## 2:                1         2            NA            NA  Missing Response
## 3:                3        14       47.6895     -122.5616  Missing Response
## 4:                2         2            NA            NA  Missing Response
## 5:                2         2            NA            NA  Missing Response
## 6:                3        11            NA            NA  Missing Response
##    prev_home_jurisdiction prev_home_county prev_home_tract_2020
##                    <char>           <char>                <i64>
## 1:       Missing Response Missing Response                 <NA>
## 2:       Missing Response Missing Response                 <NA>
## 3:      Bainbridge Island    Kitsap County          53035090800
## 4:       Missing Response Missing Response                 <NA>
## 5:       Missing Response Missing Response                 <NA>
## 6:       Missing Response Missing Response                 <NA>
##    prev_home_notwa_zip
##                 <char>
## 1:    Missing Response
## 2:    Missing Response
## 3:    Missing Response
## 4:    Missing Response
## 5:    Missing Response
## 6:               70648
##                                                                                                         prev_res_factors_specify
##                                                                                                                           <char>
## 1:                                                                                                              Missing Response
## 2:                                                                                                              Missing Response
## 3: no longer worked on bainbridge island so there was no need to live there anymore. the move was work related and life changes.
## 4:                                                                                                              Missing Response
## 5:                                                                                                              Missing Response
## 6:                                                                                                              Missing Response
##    reported_lat reported_lng sample_lat sample_lng signup_platform
##           <num>        <num>      <num>      <num>          <char>
## 1:     47.16610    -122.3689   47.16610  -122.3689         browser
## 2:     47.16448    -122.3452   47.16446  -122.3452         browser
## 3:     47.16462    -122.6908   47.16462  -122.6908         browser
## 4:     47.16329    -122.1382   47.16373  -122.1384         browser
## 5:     47.16280    -122.1396   47.16321  -122.1441         browser
## 6:     47.16292    -122.1387   47.16301  -122.1385         browser
##    traveldate_end traveldate_start home_state   home_county
##            <Date>           <Date>     <char>        <char>
## 1:     2025-03-24       2025-03-24 Washington Pierce County
## 2:     2025-03-17       2025-03-17 Washington Pierce County
## 3:     2025-03-06       2025-03-06 Washington Pierce County
## 4:     2025-03-18       2025-03-18 Washington Pierce County
## 5:     2025-03-13       2025-03-13 Washington Pierce County
## 6:     2025-03-18       2025-03-18 Washington Pierce County
##               home_jurisdiction home_rgcname home_bg_2010 home_bg_2020
##                          <char>       <char>       <char>       <char>
## 1: Unincorporated Pierce County      Not RGC 530530711002 530530711002
## 2: Unincorporated Pierce County      Not RGC 530530712062 530530712061
## 3: Unincorporated Pierce County      Not RGC 530530726033 530530726033
## 4:                  Bonney Lake      Not RGC 530530703113 530530703113
## 5:                  Bonney Lake      Not RGC 530530703113 530530703113
## 6:                  Bonney Lake      Not RGC 530530703113 530530703113
##    home_tract_2020 home_puma_2012 home_puma_2022 sample_home_bg hh_is_complete
##              <i64>          <int>          <int>         <char>         <char>
## 1:     53053071100        5311504        5325306   530530711002            Yes
## 2:     53053071206        5311506        5325305   530530712061            Yes
## 3:     53053072603        5311503        5325308   530530726033            Yes
## 4:     53053070311        5311507        5325304   530530703113            Yes
## 5:     53053070311        5311505        5325304   530530703113            Yes
## 6:     53053070311        5311507        5325304   530530703113            Yes
##                                                                     hhgroup
##                                                                      <char>
## 1: Signup survey completed via browserMove, Diary completed via browserMove
## 2: Signup survey completed via browserMove, Diary completed via browserMove
## 3: Signup survey completed via browserMove, Diary completed via browserMove
## 4: Signup survey completed via browserMove, Diary completed via browserMove
## 5: Signup survey completed via browserMove, Diary completed via browserMove
## 6: Signup survey completed via browserMove, Diary completed via browserMove
##          hhincome_broad    hhincome_detailed    hhincome_followup   hhsize
##                  <char>               <char>               <char>   <char>
## 1: Prefer not to answer Prefer not to answer Prefer not to answer 2 people
## 2: Prefer not to answer Prefer not to answer Prefer not to answer 1 person
## 3:    $100,000-$199,999    $100,000-$149,999     Missing Response 3 people
## 4:    $100,000-$199,999    $100,000-$149,999     Missing Response 2 people
## 5:      $75,000-$99,999      $75,000-$99,999     Missing Response 2 people
## 6:      $25,000-$49,999      $35,000-$49,999     Missing Response 4 people
##    home_in_region num_students       num_surveyable numadults numchildren
##            <char>       <char>               <char>    <char>      <char>
## 1:            Yes   0 students 2 surveyable persons  2 adults  0 children
## 2:            Yes   0 students  1 surveyable person   1 adult  0 children
## 3:            Yes   0 students 3 surveyable persons  3 adults  0 children
## 4:            Yes   0 students 2 surveyable persons  2 adults  0 children
## 5:            Yes   0 students 2 surveyable persons  2 adults  0 children
## 6:            Yes   0 students 4 surveyable persons  3 adults     1 child
##    numworkers prev_home_notwa_state                           prev_home_wa
##        <char>                <char>                                 <char>
## 1:  0 workers      Missing Response                       Missing Response
## 2:  0 workers      Missing Response                       Missing Response
## 3:  2 workers      Missing Response   Yes, previous home was in Washington
## 4:   1 worker      Missing Response                       Missing Response
## 5:   1 worker      Missing Response                       Missing Response
## 6:  0 workers      Missing Response No, previous home was in another state
##       prev_rent_own prev_res_factors_amenities
##              <char>                     <char>
## 1: Missing Response           Missing Response
## 2: Missing Response           Missing Response
## 3:             Rent               Not selected
## 4: Missing Response           Missing Response
## 5: Missing Response           Missing Response
## 6: Missing Response           Missing Response
##    prev_res_factors_community_change prev_res_factors_crime
##                               <char>                 <char>
## 1:                  Missing Response       Missing Response
## 2:                  Missing Response       Missing Response
## 3:                      Not Selected           Not Selected
## 4:                  Missing Response       Missing Response
## 5:                  Missing Response       Missing Response
## 6:                  Missing Response       Missing Response
##    prev_res_factors_employment prev_res_factors_forced prev_res_factors_hh_size
##                         <char>                  <char>                   <char>
## 1:            Missing Response        Missing Response         Missing Response
## 2:            Missing Response        Missing Response         Missing Response
## 3:                Not Selected                Selected             Not Selected
## 4:            Missing Response        Missing Response         Missing Response
## 5:            Missing Response        Missing Response         Missing Response
## 6:            Missing Response        Missing Response         Missing Response
##    prev_res_factors_housing_cost prev_res_factors_income_change
##                           <char>                         <char>
## 1:              Missing Response               Missing Response
## 2:              Missing Response               Missing Response
## 3:                  Not Selected                       Selected
## 4:              Missing Response               Missing Response
## 5:              Missing Response               Missing Response
## 6:              Missing Response               Missing Response
##    prev_res_factors_less_space prev_res_factors_more_space
##                         <char>                      <char>
## 1:            Missing Response            Missing Response
## 2:            Missing Response            Missing Response
## 3:                Not Selected                Not Selected
## 4:            Missing Response            Missing Response
## 5:            Missing Response            Missing Response
## 6:            Missing Response            Missing Response
##    prev_res_factors_no_answer prev_res_factors_other prev_res_factors_quality
##                        <char>                 <char>                   <char>
## 1:           Missing Response       Missing Response         Missing Response
## 2:           Missing Response       Missing Response         Missing Response
## 3:               Not Selected               Selected             Not Selected
## 4:           Missing Response       Missing Response         Missing Response
## 5:           Missing Response       Missing Response         Missing Response
## 6:           Missing Response       Missing Response         Missing Response
##    prev_res_factors_school prev_res_factors_telework     displaced
##                     <char>                    <char>        <char>
## 1:        Missing Response          Missing Response Not Displaced
## 2:        Missing Response          Missing Response Not Displaced
## 3:            Not Selected              Not selected     Displaced
## 4:        Missing Response          Missing Response Not Displaced
## 5:        Missing Response          Missing Response Not Displaced
## 6:        Missing Response          Missing Response Not Displaced
##                           prev_res_type            rent_own
##                                  <char>              <char>
## 1:                     Missing Response                Rent
## 2:                     Missing Response Own/paying mortgage
## 3: Single-family house (detached house) Own/paying mortgage
## 4:                     Missing Response Own/paying mortgage
## 5:                     Missing Response Own/paying mortgage
## 6:                     Missing Response Own/paying mortgage
##                    res_dur                             res_type
##                     <char>                               <char>
## 1:  Between 5 and 10 years Single-family house (detached house)
## 2:      More than 20 years Single-family house (detached house)
## 3:   Between 1 and 2 years Single-family house (detached house)
## 4:      More than 20 years                  Mobile home/trailer
## 5: Between 10 and 20 years                  Mobile home/trailer
## 6:   Between 3 and 5 years                  Mobile home/trailer
##                         sample_segment vehicle_count home_lat home_lng
##                                 <char>        <char>    <num>    <num>
## 1: Pierce County, Unincorporated Rural     1 vehicle  47.1661 -122.369
## 2: Pierce County, Unincorporated Rural     1 vehicle  47.1645 -122.345
## 3: Pierce County, Unincorporated Rural    2 vehicles  47.1646 -122.691
## 4: Pierce County, Unincorporated Rural     1 vehicle  47.1633 -122.138
## 5: Pierce County, Unincorporated Rural     1 vehicle  47.1628 -122.140
## 6: Pierce County, Unincorporated Rural     1 vehicle  47.1629 -122.139
##                 hh_race_category             lifecycle_class        broadband
##                           <char>                      <char>           <char>
## 1: Some Other Races non-Hispanic Household with older adults Missing Response
## 2:            White non-Hispanic Household with older adults Missing Response
## 3:            White non-Hispanic Household with older adults Missing Response
## 4:            White non-Hispanic Household with older adults Missing Response
## 5:              Missing Response Household with older adults Missing Response
## 6:            White non-Hispanic Household includes children Missing Response
##    home_hh_1 home_hh_2 home_emptot_1 home_emptot_2 home_auto_jobs_access
##        <num>     <num>         <num>         <num>                 <num>
## 1:  35.96706  167.4027      8.102395     167.88978                434495
## 2: 162.73276  484.0147      3.884260     115.87999                424872
## 3: 103.69945  319.7172      1.396064      14.32801                 10315
## 4:  87.81106  303.8064     13.519101    1051.89333                117497
## 5: 151.05243  551.0467    134.179668    1169.53745                117497
## 6:  65.64457  333.5057     15.916661     715.28750                117497
##    home_transit_jobs_access home_nodes1_1 home_nodes3_1 home_nodes4_1
##                       <num>         <num>         <num>         <num>
## 1:                     1617             0             0      4.736907
## 2:                     1843             0             0     13.692040
## 3:                        0             0             0     20.146138
## 4:                        0             0             0     10.312137
## 5:                        0             0             0      7.568961
## 6:                        0             0             0      4.538913
##    home_nodes1_2 home_nodes3_2 home_nodes4_2 hh_weight
##            <num>         <num>         <num>     <num>
## 1:             0             0      16.75891  22.80259
## 2:             0             0      38.97919  75.94698
## 3:             0             0      57.00128  25.27773
## 4:             0             0      26.33625  22.72812
## 5:             0             0      32.15610 124.77240
## 6:             0             0      29.46154  40.95438
## 
## $person
##     person_id survey_year  ethnicity_other household_id   industry_other
##         <i64>       <int>           <char>        <i64>           <char>
## 1: 2500000601        2025 Missing Response     25000006 Missing Response
## 2: 2500000602        2025 Missing Response     25000006 Missing Response
## 3: 2500007101        2025 Missing Response     25000071 Missing Response
## 4: 2500009601        2025 Missing Response     25000096 Missing Response
## 5: 2500009602        2025 Missing Response     25000096 Missing Response
## 6: 2500009603        2025 Missing Response     25000096 Missing Response
##    can_drive num_trips numdayscomplete pernum          race_other_specify
##       <char>     <int>          <char>  <int>                      <char>
## 1:       Yes         0               1      1            Missing Response
## 2:       Yes         0               1      2 1/2 Caucasian, 1/2 Japanese
## 3:       Yes         2               1      1            Missing Response
## 4:       Yes         9               1      1            Missing Response
## 5:       Yes         0               1      2            Missing Response
## 6:        No         5               1      3            Missing Response
##           school_bg school_loc_lat school_loc_lng    school_puma10
##              <char>          <num>          <num>           <char>
## 1: Missing Response             NA             NA Missing Response
## 2: Missing Response             NA             NA Missing Response
## 3: Missing Response             NA             NA Missing Response
## 4: Missing Response             NA             NA Missing Response
## 5: Missing Response             NA             NA Missing Response
## 6: Missing Response             NA             NA Missing Response
##      school_rgcname school_jurisdiction    school_county school_tract_2020
##              <char>              <char>           <char>             <i64>
## 1: Missing Response    Missing Response Missing Response              <NA>
## 2: Missing Response    Missing Response Missing Response              <NA>
## 3: Missing Response    Missing Response Missing Response              <NA>
## 4: Missing Response    Missing Response Missing Response              <NA>
## 5: Missing Response    Missing Response Missing Response              <NA>
## 6: Missing Response    Missing Response Missing Response              <NA>
##    second_home_lat second_home_lon          work_bg work_lat work_lng
##              <num>           <num>           <char>    <num>    <num>
## 1:              NA              NA Missing Response       NA       NA
## 2:              NA              NA Missing Response       NA       NA
## 3:              NA              NA Missing Response       NA       NA
## 4:              NA              NA     530530726033  47.1646 -122.691
## 5:              NA              NA Missing Response       NA       NA
## 6:              NA              NA     530530726033  47.1569 -122.690
##         work_puma10              work_rgcname            work_jurisdiction
##              <char>                    <char>                       <char>
## 1: Missing Response          Missing Response             Missing Response
## 2: Missing Response          Missing Response             Missing Response
## 3: Missing Response          Missing Response             Missing Response
## 4:            11502 No Regional Growth Center Unincorporated Pierce County
## 5: Missing Response          Missing Response             Missing Response
## 6:            11502 No Regional Growth Center Unincorporated Pierce County
##         work_county work_tract_2020        bike_freq    carshare_freq
##              <char>           <i64>           <char>           <char>
## 1: Missing Response            <NA> Missing Response Missing Response
## 2: Missing Response            <NA> Missing Response Missing Response
## 3: Missing Response            <NA> Missing Response Missing Response
## 4:    Pierce County     53053072603 Missing Response Missing Response
## 5: Missing Response            <NA> Missing Response Missing Response
## 6:    Pierce County     53053072603 Missing Response Missing Response
##        commute_freq commute_subsidy_1 commute_subsidy_2 commute_subsidy_3
##              <char>            <char>            <char>            <char>
## 1: Missing Response  Missing Response  Missing Response  Missing Response
## 2: Missing Response  Missing Response  Missing Response  Missing Response
## 3: Missing Response  Missing Response  Missing Response  Missing Response
## 4:    3 days a week  Missing Response  Missing Response  Missing Response
## 5: Missing Response  Missing Response  Missing Response  Missing Response
## 6:     1 day a week      Not selected      Not selected      Not selected
##    commute_subsidy_4 commute_subsidy_5 commute_subsidy_6 commute_subsidy_7
##               <char>            <char>            <char>            <char>
## 1:  Missing Response  Missing Response  Missing Response  Missing Response
## 2:  Missing Response  Missing Response  Missing Response  Missing Response
## 3:  Missing Response  Missing Response  Missing Response  Missing Response
## 4:  Missing Response  Missing Response  Missing Response  Missing Response
## 5:  Missing Response  Missing Response  Missing Response  Missing Response
## 6:      Not selected      Not selected      Not selected      Not selected
##    commute_subsidy_996 commute_subsidy_998 commute_subsidy_use_1
##                 <char>              <char>                <char>
## 1:    Missing Response    Missing Response      Missing Response
## 2:    Missing Response    Missing Response      Missing Response
## 3:    Missing Response    Missing Response      Missing Response
## 4:    Missing Response    Missing Response      Missing Response
## 5:    Missing Response    Missing Response      Missing Response
## 6:            Selected        Not selected      Missing Response
##    commute_subsidy_use_2 commute_subsidy_use_3 commute_subsidy_use_4
##                   <char>                <char>                <char>
## 1:      Missing Response      Missing Response      Missing Response
## 2:      Missing Response      Missing Response      Missing Response
## 3:      Missing Response      Missing Response      Missing Response
## 4:      Missing Response      Missing Response      Missing Response
## 5:      Missing Response      Missing Response      Missing Response
## 6:      Missing Response      Missing Response      Missing Response
##    commute_subsidy_use_5 commute_subsidy_use_6 commute_subsidy_use_7
##                   <char>                <char>                <char>
## 1:      Missing Response      Missing Response      Missing Response
## 2:      Missing Response      Missing Response      Missing Response
## 3:      Missing Response      Missing Response      Missing Response
## 4:      Missing Response      Missing Response      Missing Response
## 5:      Missing Response      Missing Response      Missing Response
## 6:      Missing Response      Missing Response      Missing Response
##    commute_subsidy_use_996 disability_person                     education
##                     <char>            <char>                        <char>
## 1:        Missing Response                No                  Some college
## 2:        Missing Response                No Vocational/technical training
## 3:        Missing Response                No Graduate/post-graduate degree
## 4:        Missing Response                No               Bachelor degree
## 5:        Missing Response               Yes               Bachelor degree
## 6:        Missing Response               Yes          High school graduate
##                                                                                               employment
##                                                                                                   <char>
## 1: Not employed for pay (e.g., unpaid furlough, looking for work, retired, stay-at-home parent, student)
## 2: Not employed for pay (e.g., unpaid furlough, looking for work, retired, stay-at-home parent, student)
## 3: Not employed for pay (e.g., unpaid furlough, looking for work, retired, stay-at-home parent, student)
## 4:                                                                                         Self-employed
## 5: Not employed for pay (e.g., unpaid furlough, looking for work, retired, stay-at-home parent, student)
## 6:                                                   Employed part time (fewer than 35 hours/week, paid)
##    ethnicity_1  ethnicity_2  ethnicity_3  ethnicity_4 ethnicity_997
##         <char>       <char>       <char>       <char>        <char>
## 1:    Selected Not selected Not selected Not selected  Not selected
## 2:    Selected Not selected Not selected Not selected  Not selected
## 3:    Selected Not selected Not selected Not selected  Not selected
## 4:    Selected Not selected Not selected Not selected  Not selected
## 5:    Selected Not selected Not selected Not selected  Not selected
## 6:    Selected Not selected Not selected Not selected  Not selected
##    ethnicity_999    race_hisp ev_typical_charge_1 ev_typical_charge_2
##           <char>       <char>              <char>              <char>
## 1:  Not selected Not Selected    Missing Response    Missing Response
## 2:  Not selected Not Selected    Missing Response    Missing Response
## 3:  Not selected Not Selected    Missing Response    Missing Response
## 4:  Not selected Not Selected    Missing Response    Missing Response
## 5:  Not selected Not Selected    Missing Response    Missing Response
## 6:  Not selected Not Selected    Missing Response    Missing Response
##    ev_typical_charge_3 ev_typical_charge_4 ev_typical_charge_5
##                 <char>              <char>              <char>
## 1:    Missing Response    Missing Response    Missing Response
## 2:    Missing Response    Missing Response    Missing Response
## 3:    Missing Response    Missing Response    Missing Response
## 4:    Missing Response    Missing Response    Missing Response
## 5:    Missing Response    Missing Response    Missing Response
## 6:    Missing Response    Missing Response    Missing Response
##    ev_typical_charge_6 ev_typical_charge_997
##                 <char>                <char>
## 1:    Missing Response      Missing Response
## 2:    Missing Response      Missing Response
## 3:    Missing Response      Missing Response
## 4:    Missing Response      Missing Response
## 5:    Missing Response      Missing Response
## 6:    Missing Response      Missing Response
##                                   gender         hours_work
##                                   <char>             <char>
## 1:    Boy/Man (cisgender or transgender)   Missing Response
## 2: Girl/Woman (cisgender or transgender)   Missing Response
## 3:    Boy/Man (cisgender or transgender)   Missing Response
## 4:    Boy/Man (cisgender or transgender) More than 50 hours
## 5: Girl/Woman (cisgender or transgender)   Missing Response
## 6:    Boy/Man (cisgender or transgender)  10 hours or fewer
##                                         industry is_participant
##                                           <char>         <char>
## 1:                              Missing Response            Yes
## 2:                              Missing Response            Yes
## 3:                              Missing Response            Yes
## 4:                                   Health care            Yes
## 5:                              Missing Response            Yes
## 6: Hospitality (e.g., restaurant, accommodation)            Yes
##      mobility_aides office_available participate person_is_complete  proxy
##              <char>           <char>      <char>             <char> <char>
## 1: Missing Response Missing Response         Yes                Yes     No
## 2: Missing Response Missing Response          No                Yes     No
## 3: Missing Response Missing Response         Yes                Yes     No
## 4: Missing Response Missing Response         Yes                Yes     No
## 5:              Yes Missing Response         Yes                Yes     No
## 6:              Yes Missing Response         Yes                Yes     No
##    proxy_parent    race_afam    race_aiak   race_asian    race_hapi
##          <char>       <char>       <char>       <char>       <char>
## 1:           No Not Selected Not Selected Not Selected Not Selected
## 2:           No Not Selected Not Selected Not Selected Not Selected
## 3:           No Not Selected Not Selected Not Selected Not Selected
## 4:           No Not Selected Not Selected Not Selected Not Selected
## 5:           No Not Selected Not Selected Not Selected Not Selected
## 6:           No Not Selected Not Selected Not Selected Not Selected
##    race_noanswer   race_other   race_white                race_category
##           <char>       <char>       <char>                       <char>
## 1:  Not Selected Not Selected     Selected           White non-Hispanic
## 2:  Not Selected     Selected Not Selected Some Other Race non-Hispanic
## 3:  Not Selected Not Selected     Selected           White non-Hispanic
## 4:  Not Selected Not Selected     Selected           White non-Hispanic
## 5:  Not Selected Not Selected     Selected           White non-Hispanic
## 6:  Not Selected Not Selected     Selected           White non-Hispanic
##                         relationship remote_class_freq      school_freq
##                               <char>            <char>           <char>
## 1:                              Self  Missing Response Missing Response
## 2:                 Spouse or partner  Missing Response Missing Response
## 3:                              Self  Missing Response Missing Response
## 4:                              Self  Missing Response Missing Response
## 5:                 Spouse or partner  Missing Response Missing Response
## 6: Son or daughter (or child in-law)  Missing Response Missing Response
##    school_in_region school_mode_typical       schooltype
##              <char>              <char>           <char>
## 1:               No    Missing Response Missing Response
## 2:               No    Missing Response Missing Response
## 3:               No    Missing Response Missing Response
## 4:               No    Missing Response Missing Response
## 5:               No    Missing Response Missing Response
## 6:               No    Missing Response Missing Response
##                                      second_home second_home_in_region
##                                           <char>                <char>
## 1: Does not regularly spend night at second home      Missing Response
## 2: Does not regularly spend night at second home      Missing Response
## 3: Does not regularly spend night at second home      Missing Response
## 4: Does not regularly spend night at second home      Missing Response
## 5: Does not regularly spend night at second home      Missing Response
## 6: Does not regularly spend night at second home      Missing Response
##                  sexuality      share_1      share_2      share_3      share_4
##                     <char>       <char>       <char>       <char>       <char>
## 1: Heterosexual (straight)     Selected Not selected Not selected Not selected
## 2: Heterosexual (straight)     Selected Not selected Not selected Not selected
## 3:    Prefer not to answer     Selected Not selected Not selected Not selected
## 4: Heterosexual (straight)     Selected Not selected     Selected Not selected
## 5: Heterosexual (straight) Not selected Not selected     Selected Not selected
## 6: Heterosexual (straight) Not selected Not selected     Selected Not selected
##         share_5    share_996            smartphone_type     adult_student
##          <char>       <char>                     <char>            <char>
## 1: Not selected Not selected       Has an Android phone No, not a student
## 2: Not selected Not selected       Has an Android phone No, not a student
## 3: Not selected Not selected       Has an Android phone No, not a student
## 4: Not selected Not selected       Has an Android phone No, not a student
## 5: Not selected Not selected       Has an Android phone No, not a student
## 6: Not selected Not selected Does not have a smartphone No, not a student
##    surveyable telecommute_freq commute_subsidy_parking commute_subsidy_transit
##        <char>           <char>                  <char>                  <char>
## 1:        Yes Missing Response        Missing Response        Missing Response
## 2:        Yes Missing Response        Missing Response        Missing Response
## 3:        Yes Missing Response        Missing Response        Missing Response
## 4:        Yes  6-7 days a week        Missing Response        Missing Response
## 5:        Yes Missing Response        Missing Response        Missing Response
## 6:        Yes            Never             Not offered             Not offered
##            tnc_freq     transit_freq        walk_freq work_in_region
##              <char>           <char>           <char>         <char>
## 1: Missing Response Missing Response    2 days a week             No
## 2: Missing Response Missing Response    2 days a week             No
## 3: Missing Response Missing Response  6-7 days a week             No
## 4: Missing Response    4 days a week     1 day a week            Yes
## 5: Missing Response    3 days a week Missing Response             No
## 6: Missing Response    3 days a week Missing Response            Yes
##                            work_mode
##                               <char>
## 1:                  Missing Response
## 2:                  Missing Response
## 3:                  Missing Response
## 4:                  Drive onto ferry
## 5:                  Missing Response
## 6: Household vehicle (or motorcycle)
##                                                                        workplace
##                                                                           <char>
## 1:                                                              Missing Response
## 2:                                                              Missing Response
## 3:                                                              Missing Response
## 4: Work location outside the home regularly varies (different offices/job sites)
## 5:                                                              Missing Response
## 6:                                        Only one work location outside of home
##            age age_detailed   drive_for_work paid_work return_from_leave
##         <char>       <char>           <char>    <char>            <char>
## 1: 65-74 years 72 years old Missing Response        No  Missing Response
## 2: 65-74 years 71 years old Missing Response        No  Missing Response
## 3: 75-84 years 75 years old Missing Response        No  Missing Response
## 4: 55-64 years 57 years old Missing Response       Yes  Missing Response
## 5: 75-84 years 76 years old Missing Response        No  Missing Response
## 6: 45-54 years 47 years old Missing Response       Yes  Missing Response
##    transportation_statement_car transportation_statement_environ
##                          <char>                           <char>
## 1:                      4-Agree              1-Strongly disagree
## 2:             Missing Response                 Missing Response
## 3:             5-Strongly Agree              1-Strongly disagree
## 4:             5-Strongly Agree                       2-Disagree
## 5:             Missing Response                 Missing Response
## 6:             Missing Response                 Missing Response
##    transportation_statement_retail transportation_statement_telework
##                             <char>                            <char>
## 1:                         4-Agree                  Missing Response
## 2:                Missing Response                  Missing Response
## 3:    3-Neither agree nor disagree                  Missing Response
## 4:    3-Neither agree nor disagree                           4-Agree
## 5:                Missing Response                  Missing Response
## 6:                Missing Response                  Missing Response
##    transportation_statement_transit transportation_statement_travel
##                              <char>                          <char>
## 1:                          4-Agree                         4-Agree
## 2:                 Missing Response                Missing Response
## 3:                       2-Disagree                      2-Disagree
## 4:                          4-Agree    3-Neither agree nor disagree
## 5:                 Missing Response                Missing Response
## 6:                 Missing Response                Missing Response
##    transportation_statement_walk
##                           <char>
## 1:                    2-Disagree
## 2:              Missing Response
## 3:  3-Neither agree nor disagree
## 4:                    2-Disagree
## 5:              Missing Response
## 6:              Missing Response
##                                                              work_from_home
##                                                                      <char>
## 1:                                                         Missing Response
## 2:                                                         Missing Response
## 3:                                                         Missing Response
## 4: Yes, most of the time (at least than 50% but less than 100% of the time)
## 5:                                                         Missing Response
## 6:                                                                       No
##    work_hh_1 work_hh_2 work_emptot_1 work_emptot_2 school_hh_1 school_hh_2
##        <num>     <num>         <num>         <num>       <num>       <num>
## 1:        NA        NA            NA            NA          NA          NA
## 2:        NA        NA            NA            NA          NA          NA
## 3:        NA        NA            NA            NA          NA          NA
## 4:  103.6994  319.7172      1.396064      14.32801          NA          NA
## 5:        NA        NA            NA            NA          NA          NA
## 6:  109.2150  320.4932      9.227058      18.41387          NA          NA
##    school_emptot_1 school_emptot_2 school_nodes1_1 school_nodes3_1
##              <num>           <num>           <num>           <num>
## 1:              NA              NA              NA              NA
## 2:              NA              NA              NA              NA
## 3:              NA              NA              NA              NA
## 4:              NA              NA              NA              NA
## 5:              NA              NA              NA              NA
## 6:              NA              NA              NA              NA
##    school_nodes4_1 school_nodes1_2 school_nodes3_2 school_nodes4_2
##              <num>           <num>           <num>           <num>
## 1:              NA              NA              NA              NA
## 2:              NA              NA              NA              NA
## 3:              NA              NA              NA              NA
## 4:              NA              NA              NA              NA
## 5:              NA              NA              NA              NA
## 6:              NA              NA              NA              NA
##    work_nodes1_1 work_nodes3_1 work_nodes4_1 work_nodes1_2 work_nodes3_2
##            <num>         <num>         <num>         <num>         <num>
## 1:            NA            NA            NA            NA            NA
## 2:            NA            NA            NA            NA            NA
## 3:            NA            NA            NA            NA            NA
## 4:             0             0      20.14614             0             0
## 5:            NA            NA            NA            NA            NA
## 6:             0             0      19.26790             0             0
##    work_nodes4_2 person_weight
##            <num>         <num>
## 1:            NA      22.80259
## 2:            NA      22.80259
## 3:            NA      75.94698
## 4:      57.00128      25.27773
## 5:            NA      25.27773
## 6:      55.18191      25.27773
## 
## $day
##          day_id survey_year daynum household_id num_trips pernum  person_id
##           <i64>       <int>  <int>        <i64>     <int>  <int>      <i64>
## 1: 250000060101        2025      1     25000006         0      1 2500000601
## 2: 250000060201        2025      1     25000006         0      2 2500000602
## 3: 250000710101        2025      1     25000071         2      1 2500007101
## 4: 250000960101        2025      1     25000096         9      1 2500009601
## 5: 250000960201        2025      1     25000096         0      2 2500009602
## 6: 250000960301        2025      1     25000096         5      3 2500009603
##    travel_date  attend_school_1  attend_school_2  attend_school_3
##         <POSc>           <char>           <char>           <char>
## 1:  2025-03-24 Missing Response Missing Response Missing Response
## 2:  2025-03-24 Missing Response Missing Response Missing Response
## 3:  2025-03-17 Missing Response Missing Response Missing Response
## 4:  2025-03-06 Missing Response Missing Response Missing Response
## 5:  2025-03-06 Missing Response Missing Response Missing Response
## 6:  2025-03-06 Missing Response Missing Response Missing Response
##    attend_school_998 attend_school_999 deliver_elsewhere     deliver_food
##               <char>            <char>            <char>           <char>
## 1:  Missing Response  Missing Response      Not selected               No
## 2:  Missing Response  Missing Response  Missing Response Missing Response
## 3:  Missing Response  Missing Response      Not selected               No
## 4:  Missing Response  Missing Response      Not selected               No
## 5:  Missing Response  Missing Response  Missing Response Missing Response
## 6:  Missing Response  Missing Response  Missing Response Missing Response
##     deliver_grocery     deliver_none   deliver_office    deliver_other
##              <char>           <char>           <char>           <char>
## 1:               No     Not selected     Not selected     Not selected
## 2: Missing Response Missing Response Missing Response Missing Response
## 3:               No         Selected     Not selected     Not selected
## 4:               No         Selected     Not selected     Not selected
## 5: Missing Response Missing Response Missing Response Missing Response
## 6: Missing Response Missing Response Missing Response Missing Response
##     deliver_package     deliver_work is_participant          loc_end
##              <char>           <char>         <char>           <char>
## 1:              Yes               No            Yes Missing Response
## 2: Missing Response Missing Response            Yes Missing Response
## 3:               No               No            Yes             Home
## 4:               No               No            Yes             Home
## 5: Missing Response Missing Response            Yes Missing Response
## 6: Missing Response Missing Response            Yes             Home
##           loc_start   proxy_complete summary_complete surveyable
##              <char>           <char>           <char>     <char>
## 1: Missing Response Missing Response              Yes        Yes
## 2: Missing Response Missing Response              Yes        Yes
## 3:             Home Missing Response              Yes        Yes
## 4:             Home Missing Response              Yes        Yes
## 5: Missing Response Missing Response              Yes        Yes
## 6:             Home Missing Response              Yes        Yes
##         telework_time travel_day travel_dow no_travel   no_school_sick
##                <char>     <char>     <char>    <char>           <char>
## 1:   Missing Response        Yes     Monday       Yes Missing Response
## 2:   Missing Response        Yes     Monday       Yes Missing Response
## 3:   Missing Response        Yes     Monday        No Missing Response
## 4: 3 hours 30 minutes        Yes   Thursday        No Missing Response
## 5:   Missing Response        Yes   Thursday       Yes Missing Response
## 6:          0 minutes        Yes   Thursday        No Missing Response
##    no_school_online_home no_school_online_other no_school_vacation
##                   <char>                 <char>             <char>
## 1:      Missing Response       Missing Response   Missing Response
## 2:      Missing Response       Missing Response   Missing Response
## 3:      Missing Response       Missing Response   Missing Response
## 4:      Missing Response       Missing Response   Missing Response
## 5:      Missing Response       Missing Response   Missing Response
## 6:      Missing Response       Missing Response   Missing Response
##    no_school_closed  no_school_other no_school_dont_know no_school_no_answer
##              <char>           <char>              <char>              <char>
## 1: Missing Response Missing Response    Missing Response    Missing Response
## 2: Missing Response Missing Response    Missing Response    Missing Response
## 3: Missing Response Missing Response    Missing Response    Missing Response
## 4: Missing Response Missing Response    Missing Response    Missing Response
## 5: Missing Response Missing Response    Missing Response    Missing Response
## 6: Missing Response Missing Response    Missing Response    Missing Response
##    notravel_madetrips notravel_vacation notravel_telecommute notravel_housework
##                <char>            <char>               <char>             <char>
## 1:   Missing Response  Missing Response     Missing Response   Missing Response
## 2:   Missing Response  Missing Response     Missing Response   Missing Response
## 3:   Missing Response  Missing Response     Missing Response   Missing Response
## 4:   Missing Response  Missing Response     Missing Response   Missing Response
## 5:   Missing Response  Missing Response     Missing Response   Missing Response
## 6:   Missing Response  Missing Response     Missing Response   Missing Response
##    notravel_kidsbreak notravel_notransport    notravel_sick notravel_delivery
##                <char>               <char>           <char>            <char>
## 1:   Missing Response     Missing Response Missing Response  Missing Response
## 2:   Missing Response     Missing Response Missing Response  Missing Response
## 3:   Missing Response     Missing Response Missing Response  Missing Response
## 4:   Missing Response     Missing Response Missing Response  Missing Response
## 5:   Missing Response     Missing Response Missing Response  Missing Response
## 6:   Missing Response     Missing Response Missing Response  Missing Response
##    notravel_kidshomeschool notravel_weather notravel_not_sure   notravel_other
##                     <char>           <char>            <char>           <char>
## 1:        Missing Response Missing Response  Missing Response Missing Response
## 2:        Missing Response Missing Response  Missing Response Missing Response
## 3:        Missing Response Missing Response  Missing Response Missing Response
## 4:        Missing Response Missing Response  Missing Response Missing Response
## 5:        Missing Response Missing Response  Missing Response Missing Response
## 6:        Missing Response Missing Response  Missing Response Missing Response
##    day_weight
##         <num>
## 1:   22.80259
## 2:   22.80259
## 3:   75.94698
## 4:   25.27773
## 5:   25.27773
## 6:   25.27773
## 
## $trip
##          trip_id survey_year household_id  person_id pernum tripnum traveldate
##            <i64>       <int>        <i64>      <i64>  <int>   <int>     <POSc>
## 1: 2500007101001        2025     25000071 2500007101      1       1 2025-03-17
## 2: 2500007101002        2025     25000071 2500007101      1       2 2025-03-17
## 3: 2500009601001        2025     25000096 2500009601      1       1 2025-03-06
## 4: 2500009601002        2025     25000096 2500009601      1       2 2025-03-06
## 5: 2500009601003        2025     25000096 2500009601      1       3 2025-03-06
## 6: 2500009601004        2025     25000096 2500009601      1       4 2025-03-06
##    daynum depart_time_timestamp arrival_time_timestamp origin_lat origin_lng
##     <int>                <POSc>                 <POSc>      <num>      <num>
## 1:      1   2025-03-17 11:00:00    2025-03-17 11:10:00   47.16448  -122.3452
## 2:      1   2025-03-17 11:40:00    2025-03-17 11:50:00   47.18004  -122.3185
## 3:      1   2025-03-06 09:32:00    2025-03-06 09:45:00   47.16462  -122.6908
## 4:      1   2025-03-06 10:30:00    2025-03-06 10:45:00   47.17742  -122.6795
## 5:      1   2025-03-06 10:47:00    2025-03-06 10:55:00   47.21412  -122.5366
## 6:      1   2025-03-06 10:58:00    2025-03-06 11:05:00   47.21982  -122.5367
##    dest_lat  dest_lng distance_miles travel_time          travelers_hh
##       <num>     <num>          <num>       <num>                <char>
## 1: 47.18004 -122.3185      4.0880112          10  1 household traveler
## 2: 47.16448 -122.3452      4.0451365          10  1 household traveler
## 3: 47.17742 -122.6795      2.0505300          13 2 household travelers
## 4: 47.21412 -122.5366      9.4715846          15 2 household travelers
## 5: 47.21982 -122.5367      0.4436601           8 2 household travelers
## 6: 47.22630 -122.5359      0.5834690           7 2 household travelers
##    travelers_nonhh travelers_total
##             <char>          <char>
## 1: No other people      1 traveler
## 2: No other people      1 traveler
## 3: No other people     2 travelers
## 4: No other people     2 travelers
## 5: No other people     2 travelers
## 6: No other people     2 travelers
##                                                      origin_purpose
##                                                              <char>
## 1:                                                        Went home
## 2:               Went to exercise (e.g., gym, walk, jog, bike ride)
## 3:                                                        Went home
## 4: Went to work-related place (e.g., meeting, second job, delivery)
## 5:                   Went to other shopping (e.g., mall, pet store)
## 6:            Conducted personal business (e.g., bank, post office)
##                      origin_purpose_cat origin_purpose_cat_5
##                                  <char>               <char>
## 1:                                 Home                 Home
## 2:                    Social/Recreation    Social/Recreation
## 3:                                 Home                 Home
## 4:                         Work-related          Work/School
## 5:                             Shopping      Errand/Shopping
## 6: Personal Business/Errand/Appointment      Errand/Shopping
##                                                         dest_purpose
##                                                               <char>
## 1:           Exercise or recreation (e.g., gym, jog, bike, walk dog)
## 2:                                                         Went home
## 3: Went to work-related activity (e.g., meeting, delivery, worksite)
## 4:                            Other shopping (e.g., mall, pet store)
## 5:                       Personal business (e.g., bank, post office)
## 6:                       Personal business (e.g., bank, post office)
##                        dest_purpose_cat dest_purpose_cat_5 dest_purpose_other
##                                  <char>             <char>             <char>
## 1:                    Social/Recreation  Social/Recreation               None
## 2:                                 Home               Home               None
## 3:                         Work-related        Work/School               None
## 4:                             Shopping    Errand/Shopping               None
## 5: Personal Business/Errand/Appointment    Errand/Shopping               None
## 6: Personal Business/Errand/Appointment    Errand/Shopping               None
##                               mode_1                   mode_2           mode_3
##                               <char>                   <char>           <char>
## 1: Household vehicle (or motorcycle) Walk (or jog/wheelchair) Missing Response
## 2: Household vehicle (or motorcycle) Walk (or jog/wheelchair) Missing Response
## 3: Household vehicle (or motorcycle)         Missing Response Missing Response
## 4: Household vehicle (or motorcycle)         Missing Response Missing Response
## 5: Household vehicle (or motorcycle)         Missing Response Missing Response
## 6: Household vehicle (or motorcycle)         Missing Response Missing Response
##              mode_4 mode_class mode_class_5 driver         mode_acc
##              <char>     <char>       <char> <char>           <char>
## 1: Missing Response  Drive SOV        Drive Driver Missing Response
## 2: Missing Response  Drive SOV        Drive Driver Missing Response
## 3: Missing Response Drive HOV2        Drive Driver Missing Response
## 4: Missing Response Drive HOV2        Drive Driver Missing Response
## 5: Missing Response Drive HOV2        Drive Driver Missing Response
## 6: Missing Response Drive HOV2        Drive Driver Missing Response
##            mode_egr speed_mph mode_other_specify origin_x_coord origin_y_coord
##              <char>     <num>             <char>          <num>          <num>
## 1: Missing Response 24.528166               None        1264306       63693.14
## 2: Missing Response 24.270916               None        1271069       69238.36
## 3: Missing Response  9.464014               None        1178372       65625.72
## 4: Missing Response 37.886338               None        1181287       70226.01
## 5: Missing Response  3.327459               None        1217116       82790.27
## 6: Missing Response  5.001191               None        1217137       84869.50
##    origin_county   origin_rgcname          origin_jurisdiction
##           <char>           <char>                       <char>
## 1: Pierce County          Not RGC Unincorporated Pierce County
## 2: Pierce County          Not RGC                     Puyallup
## 3: Pierce County          Not RGC Unincorporated Pierce County
## 4: Pierce County          Not RGC Unincorporated Pierce County
## 5: Pierce County University Place             University Place
## 6: Pierce County University Place             University Place
##    origin_tract_2010 origin_tract_2020 dest_x_coord dest_y_coord
##                <i64>             <i64>        <num>        <num>
## 1:       53053071206       53053071206      1271069     69238.36
## 2:       53053073404       53053073404      1264306     63693.14
## 3:       53053072603       53053072603      1181287     70226.01
## 4:       53053072603       53053072603      1217116     82790.27
## 5:       53053072313       53053072313      1217137     84869.50
## 6:       53053072311       53053072311      1217390     87228.18
##         dest_county     dest_rgcname            dest_jurisdiction
##              <char>           <char>                       <char>
## 1:    Pierce County          Not RGC                     Puyallup
## 2:    Pierce County          Not RGC Unincorporated Pierce County
## 3: Missing Response          Not RGC Unincorporated Pierce County
## 4:    Pierce County University Place             University Place
## 5:    Pierce County University Place             University Place
## 6:    Pierce County University Place             University Place
##    dest_tract_2010 dest_tract_2020     dest_is_home     dest_is_work
##              <i64>           <i64>           <char>           <char>
## 1:     53053073404     53053073404 Missing Response Missing Response
## 2:     53053071206     53053071206                1 Missing Response
## 3:     53053072603     53053072603 Missing Response Missing Response
## 4:     53053072313     53053072313 Missing Response Missing Response
## 5:     53053072311     53053072311 Missing Response Missing Response
## 6:     53053072311     53053072311 Missing Response Missing Response
##              modes arrival_time_hour arrival_time_minute arrival_time_second
##             <char>             <int>               <int>               <int>
## 1:   100,1,995,995                11                  10                   0
## 2:   100,1,995,995                11                  50                   0
## 3: 100,995,995,995                 9                  45                   0
## 4: 100,995,995,995                10                  45                   0
## 5: 100,995,995,995                10                  55                   0
## 6: 100,995,995,995                11                   5                   0
##    arrive_date arrive_dow copied_trip         d_bg d_in_region d_puma10
##         <Date>     <char>      <char>       <char>      <char>   <char>
## 1:  2025-03-17     Monday          No 530530734045         Yes  5311506
## 2:  2025-03-17     Monday          No 530530712061         Yes  5311506
## 3:  2025-03-06   Thursday          No 530530726033         Yes  5311502
## 4:  2025-03-06   Thursday          No 530530723132         Yes  5311503
## 5:  2025-03-06   Thursday          No 530530723113         Yes  5311502
## 6:  2025-03-06   Thursday          No 530530723113         Yes  5311502
##          day_id depart_date depart_dow depart_time_hour depart_time_minute
##           <i64>      <Date>     <char>            <int>              <int>
## 1: 250000710101  2025-03-17     Monday               11                  0
## 2: 250000710101  2025-03-17     Monday               11                 40
## 3: 250000960101  2025-03-06   Thursday                9                 32
## 4: 250000960101  2025-03-06   Thursday               10                 30
## 5: 250000960101  2025-03-06   Thursday               10                 47
## 6: 250000960101  2025-03-06   Thursday               10                 58
##    depart_time_second duration_minutes dwell_mins        flag_teleport
##                 <int>            <num>      <num>               <char>
## 1:                  0               10         30 No teleport detected
## 2:                  0               10         NA No teleport detected
## 3:                  0               13         45 No teleport detected
## 4:                  0               15          2 No teleport detected
## 5:                  0                8          3 No teleport detected
## 6:                  0                7         25 No teleport detected
##            o_bg o_in_region o_puma10 speed_flag transit_quality_flag
##          <char>      <char>   <char>     <char>               <char>
## 1: 530530712061         Yes  5311506         No                 None
## 2: 530530734045         Yes  5311506         No                 None
## 3: 530530726033         Yes  5311502         No                 None
## 4: 530530726033         Yes  5311502         No                 None
## 5: 530530723132         Yes  5311503         No                 None
## 6: 530530723113         Yes  5311502         No                 None
##    travel_date travel_dow traveldate_end traveldate_start      user_merged
##         <POSc>     <char>         <char>           <char>           <char>
## 1:  2025-03-17     Monday     2025-03-17       2025-03-17 Missing Response
## 2:  2025-03-17     Monday     2025-03-17       2025-03-17 Missing Response
## 3:  2025-03-06   Thursday     2025-03-06       2025-03-06 Missing Response
## 4:  2025-03-06   Thursday     2025-03-06       2025-03-06 Missing Response
## 5:  2025-03-06   Thursday     2025-03-06       2025-03-06 Missing Response
## 6:  2025-03-06   Thursday     2025-03-06       2025-03-06 Missing Response
##          user_split linked_trip_id linked_trip_num           n_legs
##              <char>         <char>          <char>           <char>
## 1: Missing Response  2500007101001               1 Missing Response
## 2: Missing Response  2500007101002               2 Missing Response
## 3: Missing Response  2500009601001               1 Missing Response
## 4: Missing Response  2500009601002               2 Missing Response
## 5: Missing Response  2500009601003               3 Missing Response
## 6: Missing Response  2500009601004               4 Missing Response
##    distance_meters distance_beeline_meters duration_seconds dest_zip
##              <num>                   <num>            <num>   <char>
## 1:       6579.0164                      NA              600    98371
## 2:       6510.0162                      NA              600    98373
## 3:       3300.0082                      NA              780    98303
## 4:      15243.0379                      NA              900    98466
## 5:        714.0018                      NA              480    98466
## 6:        939.0023                      NA              420    98466
##    initial_tripid origin_hh_1 origin_hh_2 origin_emptot_1 origin_emptot_2
##             <num>       <num>       <num>           <num>           <num>
## 1:   2.500007e+12   162.73276   484.01473        3.884260      115.879993
## 2:   2.500007e+12    74.33377   457.13254        2.224817       61.843372
## 3:   2.500010e+12   103.69945   319.71717        1.396064       14.328009
## 4:   2.500010e+12    23.79849    68.88144        1.390265        5.591072
## 5:   2.500010e+12  1091.18226  3483.57508      622.550043     2162.417861
## 6:   2.500010e+12  1708.93379  4088.14804     1561.232614     3117.555404
##     dest_hh_1  dest_hh_2 dest_emptot_1 dest_emptot_2 origin_nodes1_1
##         <num>      <num>         <num>         <num>           <num>
## 1:  162.73276  484.01473      3.884260    115.879993               0
## 2:   74.33377  457.13254      2.224817     61.843372               0
## 3:  103.69945  319.71717      1.396064     14.328009               0
## 4:   23.79849   68.88144      1.390265      5.591072               0
## 5: 1091.18226 3483.57508    622.550043   2162.417861               0
## 6: 1708.93379 4088.14804   1561.232614   3117.555404               0
##    origin_nodes3_1 origin_nodes4_1 origin_nodes1_2 origin_nodes3_2
##              <num>           <num>           <num>           <num>
## 1:               0       13.692040               0               0
## 2:               0        7.089761               0               0
## 3:               0       20.146138               0               0
## 4:               0        3.200234               0               0
## 5:               0       32.143501               0               0
## 6:               0       33.617271               0               0
##    origin_nodes4_2 dest_nodes1_1 dest_nodes3_1 dest_nodes4_1 dest_nodes1_2
##              <num>         <num>         <num>         <num>         <num>
## 1:        38.97919             0             0      7.089761             0
## 2:        33.06610             0             0     13.692040             0
## 3:        57.00128             0             0      3.200234             0
## 4:        17.45139             0             0     32.143501             0
## 5:       146.44651             0             0     33.617271             0
## 6:       145.31428             0             0     28.515082             0
##    dest_nodes3_2 dest_nodes4_2 trip_weight
##            <num>         <num>       <num>
## 1:             0      33.06610   102.00920
## 2:             0      38.97919   102.00920
## 3:             0      17.45139    31.96242
## 4:             0     146.44651    31.96242
## 5:             0     145.31428    33.95213
## 6:             0     109.34485    33.95213
## 
## $vehicle
##    vehicle_id survey_year household_id vehnum      make    model
##         <i64>       <int>        <i64>  <int>    <char>   <char>
## 1: 2500000601        2025     25000006      1    Subaru Forester
## 2: 2500007101        2025     25000071      1     Honda     CR-V
## 3: 2500009601        2025     25000096      1       Kia Carnival
## 4: 2500009602        2025     25000096      2       Kia     Soul
## 5: 2500015201        2025     25000152      1       Kia  Sorento
## 6: 2500020001        2025     25000200      1 Chevrolet   Blazer
##         model_other   year   fuel
##              <char> <char> <char>
## 1: Missing Response   2015    Gas
## 2: Missing Response   2017    Gas
## 3: Missing Response   2023    Gas
## 4: Missing Response   2024    Gas
## 5: Missing Response   2016    Gas
## 6: Missing Response   2021    Gas

Inspect the Weights

Weights are provided in each table to allow analysts to produce estimates that are representative of the population. The household table contains hh_weight, the person table contains person_weight, the day table contains day_weight, and the delivered trip table contains trip_weight.

Code
weight_summaries <- lapply(hts, function(tbl) {
  weight_cols <- grep("_weight$", names(tbl), value = TRUE)
  if (length(weight_cols) > 0) {
    sapply(tbl[, ..weight_cols, drop = FALSE], summary)
  }
})

weight_summaries
## $hh
##          hh_weight
## Min.      19.63636
## 1st Qu.  100.44458
## Median   209.72483
## Mean     631.52632
## 3rd Qu.  658.01126
## Max.    8061.99760
## 
## $person
##         person_weight
## Min.         19.63636
## 1st Qu.      99.01910
## Median      232.71716
## Mean        761.20234
## 3rd Qu.     831.54996
## Max.       8061.99760
## 
## $day
##         day_weight
## Min.       0.00000
## 1st Qu.    0.00000
## Median    62.00595
## Mean     389.28622
## 3rd Qu.  267.77545
## Max.    8061.99760
## 
## $trip
##          trip_weight
## Min.        4.909089
## 1st Qu.    55.780232
## Median    182.615680
## Mean      654.710998
## 3rd Qu.   566.743825
## Max.    10828.579900
## NA's    12861.000000
## 
## $vehicle
## NULL

Join Tables

Each table contains key ID columns that link records across tables (see Section 6.1 and Figure 25). The household ID (household_id) is the primary key for the household table and links to other tables. Person ID (person_id), day ID (day_id), and trip ID (trip_id) are additional keys in their respective tables.

Code
lapply(hts, function(tbl) {
  id_cols <- grep("_id$", names(tbl), value = TRUE)
  head(tbl[, ..id_cols, drop = FALSE])
})
## $hh
##    household_id
##           <i64>
## 1:     25000006
## 2:     25000071
## 3:     25000096
## 4:     25000152
## 5:     25000200
## 6:     25000255
## 
## $person
##     person_id household_id
##         <i64>        <i64>
## 1: 2500000601     25000006
## 2: 2500000602     25000006
## 3: 2500007101     25000071
## 4: 2500009601     25000096
## 5: 2500009602     25000096
## 6: 2500009603     25000096
## 
## $day
##          day_id household_id  person_id
##           <i64>        <i64>      <i64>
## 1: 250000060101     25000006 2500000601
## 2: 250000060201     25000006 2500000602
## 3: 250000710101     25000071 2500007101
## 4: 250000960101     25000096 2500009601
## 5: 250000960201     25000096 2500009602
## 6: 250000960301     25000096 2500009603
## 
## $trip
##          trip_id household_id  person_id       day_id linked_trip_id
##            <i64>        <i64>      <i64>        <i64>         <char>
## 1: 2500007101001     25000071 2500007101 250000710101  2500007101001
## 2: 2500007101002     25000071 2500007101 250000710101  2500007101002
## 3: 2500009601001     25000096 2500009601 250000960101  2500009601001
## 4: 2500009601002     25000096 2500009601 250000960101  2500009601002
## 5: 2500009601003     25000096 2500009601 250000960101  2500009601003
## 6: 2500009601004     25000096 2500009601 250000960101  2500009601004
## 
## $vehicle
##    vehicle_id household_id
##         <i64>        <i64>
## 1: 2500000601     25000006
## 2: 2500007101     25000071
## 3: 2500009601     25000096
## 4: 2500009602     25000096
## 5: 2500015201     25000152
## 6: 2500020001     25000200

For example, to get a summary of bicycle travel frequency (bike_freq, a person-level variable) by household income (hhincome_broad, a household-level variable), you would join the household and person tables on household_id:

Code
person_income_bike <- hts$person %>%
  select(person_id, household_id, bike_freq) %>%
  left_join(
    hts$hh %>%
      select(household_id, hhincome_broad),
    by = "household_id"
  )

glimpse(person_income_bike)
## Rows: 5,558
## Columns: 4
## $ person_id      <int64> 2500000601, 2500000602, 2500007101, 2500009601, 25000…
## $ household_id   <int64> 25000006, 25000006, 25000071, 25000096, 25000096, 250…
## $ bike_freq      <chr> "Missing Response", "Missing Response", "Missing Respon…
## $ hhincome_broad <chr> "Prefer not to answer", "Prefer not to answer", "Prefer…

8.3 Choosing the Right Analytic Unit

Section 6 describes the structure of households, persons, days, trips, and vehicles. This section shifts from structure to practice: how do you choose the correct analytic unit for the question you want to answer? Most analyses in an HTS fail not because of weighting errors but because the wrong table was chosen as the starting point.

Choosing the analytic unit is the first design decision in any analysis. The correct unit aligns with three things:

  1. Who or what is being measured? (a household, a person, a person-day, a trip…)
  2. What the variable conceptually describes (a household attribute, a person characteristic, a daily behavior, a movement, or a chain of movements)
  3. At what level the population is represented in sampling weights.

The guidance below connects each analytic unit to its best use cases. See also: definitions in Section 6.

Household-Level Analyses

Use households as the analytic unit when the phenomenon is shared or decided collectively-income, vehicle fleet, home location, delivery behavior, household makeup, or whether a household has zero vehicles. Even if a household variable is influenced by individual people (e.g., the number of workers in a household or the presence of children), the household is still the right level because sampling occurred at the household level.

Person-Level Analyses

Analyses about people-demographics, employment or student status, attitudinal data from Likert-scale questions-belong at the person level. Each person’s weight represents them in the regional population. Use day or trip tables only when the metric you want to measure exists at those levels.

Day-Level Analyses

Use person-days when studying daily behavior: trip rates, telework frequency, deliveries, or analyses that depend on “people who made zero trips.” The day table ensures all sampled days are included, not only days with trips.

Trip-Level Analyses

Most movement-based analyses start with trips. Mode share, trip purpose distributions, time-of-day patterns, and travel-time summaries are all trip-level analyses. Trip weights represent the population of trips, so use the trip table for any analysis where the trip is the unit of interest. If you want to analyze travel patterns across demographic groups or other attributes available in the day or person table, join that attribute to the trip table rather than switching your analytic unit.

Vehicle-Level Analyses

The vehicle table is the correct unit for vehicle fleet summaries-EV prevalence, MPG distributions (when using appended EPA fuel efficiency data), household fleet size, or daily mileage when combined with linked trip data. Vehicles belong to households, so vehicle analyses use household weights.

Research Questions and the Correct Analytic Unit

Here is a practical table mapping common questions to the proper analytic unit and weight. This is intended to serve as a quick diagnostic tool for analysts starting a new task.

Research Question Analytic Unit Starting Table Weight Variable
Share of households with zero vehicles Household hh hh_weight
Age distribution of self-identified drivers Person person person_weight
Average deliveries per day for remote workers Person-day day (joined to person) day_weight
Mode share of all trips Trip trip trip_weight
Mode differences by trip purpose Trip trip trip_weight
Proportion of electric vehicles in fleet Vehicle vehicle hh_weight
Distribution of commute times Trip trip trip_weight

8.4 Working with Categorical Response Data

The majority of data collected in the HTS are categorical variables, where respondents select from a predefined list of options. The data contains four types of categorical variables:

  1. Single-response categorical variables (SRCVs): Respondents select one option from a predefined list (e.g., gender, employment status, fuel type).
  2. Multiple-response categorical variables (MRCVs): Respondents can select multiple options from a predefined list (e.g., delivery types, previous-residence factors).
  3. Grouped categorical variables: Sets of binary indicator variables representing related categories stored across multiple columns (e.g., delivery_*, prev_res_factors_*, ethnicity_*).
  4. Count variables with top-coding: Variables representing counts with an open-ended top category (e.g., number of vehicles).

General Considerations

When working with categorical response data, keep the following best practices in mind:

  • Start with the codebook Confirm variable definitions, valid values, table membership, and skip logic there before opening the questionnaire for extra survey context.

  • Handle special codes explicitly Recode values like Missing or Prefer not to Answer before summarizing.

  • Be clear about denominators Specify whether counts are over all records, responding records, or an analytic universe.

  • Treat MRCVs appropriately Decide whether to count each selection separately (multi-count) or collapse to “Multiple.”

  • Reshape MRCV variables using long format Use pivot_longer() to analyze one response-option per row.

  • Respect question logic-or override it correctly Use the codebook logic field to determine who was in-universe, and override logic only when intentionally harmonizing denominators.

  • Don’t mistake missing for “No” Only recode missing to 0 when the respondent was logically not asked the question.

  • Clean and separate composite descriptions Split labels like "Category: Option" into Category and Response Option for readability.

  • Preserve ordering for grouped categorical scales Factor Likert or ordered scales using the codebook’s value order.

  • Avoid numeric summaries for top-coded categories Use frequencies or medians instead of means for “X or more” variables.

  • Document transformations clearly Mention recoding, filtering, collapsing, or logic overrides in both code and table notes.

The following sections provide detailed examples for each type of categorical variable, demonstrating how to implement these best practices in R using the dataset.

Single-Response Categorical Variables (SRCVs)

Single-response categorical variables are variables where respondents select one option from a predefined list. Examples include gender, employment status, and fuel type.

For example, the household income variable hhincome_broad is a single-response categorical variable with value labels that can be found in the codebook. The code snippet below retrieves those labels and their ordering from the codebook for use in analysis and presentation.

Code
# get the delivered labels and ordering for household income
income_value_labels <- cb$value_labels %>%
  filter(table == "hh", variable == "hhincome_broad") %>%
  arrange(val_order) %>%
  select(label)

gt(income_value_labels)
Table 30: Income Detailed Value Labels
label
Under $25,000
$25,000-$49,999
$50,000-$74,999
$75,000-$99,999
$100,000 or more
$100,000-$199,999
$200,000 or more
Prefer not to answer

Now we can summarize household counts by income category. Note that this summary is unweighted.

Code
# get count of households for each category of income
hh_income_counts <- hts$hh %>%
  mutate(
    hhincome_broad = factor(hhincome_broad, levels = income_value_labels$label, ordered = TRUE)
  ) %>%
  group_by(hhincome_broad) %>%
  summarize(n = n(), .groups = "drop")

gt(hh_income_counts)
hhincome_broad n
Under $25,000 201
$25,000-$49,999 337
$50,000-$74,999 356
$75,000-$99,999 379
$100,000-$199,999 844
$200,000 or more 421
Prefer not to answer 234
Table 31: Household Counts by Income Category

Note the “Prefer not to answer” category. Analysts often choose to filter out these responses before analysis. The table below shows household counts by income category, excluding “Prefer not to answer” responses.

WarningHandling Missing Data When Using Weights

If using weights and desiring confidence intervals, survey-aware methods (e.g., sryvr::survey_prop) will produce more accurate CIs compared to filtering before analysis. See Section 8.8.4 for more details.

Code
# get count of households for each category of income
hh_income_counts_nopnta <- hts$hh %>%
  mutate(
    hhincome_broad = factor(hhincome_broad, levels = income_value_labels$label, ordered = TRUE)
  ) %>%
  filter(hhincome_broad != "Prefer not to answer") %>%
  group_by(hhincome_broad) %>%
  summarize(
    n = n(),
    .groups = "drop"
  ) %>%
  ungroup() %>%
  mutate(
    pct = n / sum(n)
  )

gt(hh_income_counts_nopnta) %>%
  gt::grand_summary_rows(
    columns = "n",
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_number(., use_seps = TRUE, decimals = 0)
  ) %>%
  gt::grand_summary_rows(
    columns = "pct",
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_percent(., decimals = 1)
  ) %>%
  gt::fmt_percent(
    columns = pct,
    decimals = 1
  ) %>%
  gt::fmt_number(
    columns = n,
    decimals = 0,
    sep_mark = ","
  ) %>%
  gt::cols_label(
    hhincome_broad = "Household Income",
    n = "Count of Households",
    pct = "Weighted Percentage of Households"
  )
Household Income Count of Households Weighted Percentage of Households
Under $25,000 201 7.9%
$25,000-$49,999 337 13.3%
$50,000-$74,999 356 14.0%
$75,000-$99,999 379 14.9%
$100,000-$199,999 844 33.3%
$200,000 or more 421 16.6%
Total — 2,538 100.0%
Table 32: Household Counts by Income Category (Excluding ‘Prefer Not to Answer’)

Multiple-Response Categorical Variables (MRCVs)

Some questions in the survey allow respondents to check multiple options. Common examples include delivery types (delivery_* in the source questionnaire, delivered here as deliver_* fields), previous-residence factors (prev_res_factors_*), and race and ethnicity checkbox fields. These checkbox variables, also known as multiple response categorical variables (MRCVs), are represented differently in the dataset from SRCVs (Single Response Categorical Variables) and have to be handled differently. When available, the codebook is the best place to identify these fields before you start reshaping them. Table 33 shows a subset of those variables from the person table.

Code
# identify checkbox-style variables directly from the codebook metadata
checkbox_vars_from_codebook <- cb$variable_list %>%
  filter(person == 1, is_checkbox == 1) %>%
  pull(variable)

checkbox_vars <- checkbox_vars_from_codebook[checkbox_vars_from_codebook %in% names(hts$person)]

# build summary
checkbox_summary <- hts$person %>%
  select(all_of(checkbox_vars)) %>%
  pivot_longer(
    cols = everything(),
    names_to = "variable",
    values_to = "value"
  ) %>%
  mutate(value = ifelse(is.na(value), "NA", value)) %>%
  count(variable, value, name = "N") %>%
  pivot_wider(
    names_from = value,
    values_from = N,
    values_fill = 0
  ) %>%
  as.data.frame()

# order value columns
value_col_order <- c(
  "Selected",
  "Not Selected",
  "Not selected",
  "Missing Response",
  "Missing: Skip Logic",
  "NA"
)

checkbox_summary <- checkbox_summary %>%
  select(
    variable,
    any_of(value_col_order),
    everything()
  ) %>%
  as.data.frame()

# add grouping + cleanup
checkbox_summary <- checkbox_summary %>%
  mutate(
    group = str_match(variable, "^(.*)_(\\d+)$")[, 2],
    option = str_match(variable, "^(.*)_(\\d+)$")[, 3],
    group = ifelse(is.na(group), variable, group),
    option = ifelse(is.na(option), variable, option),
    not_selected_total = coalesce(`Not Selected`, 0L) + coalesce(`Not selected`, 0L),
    option = ifelse(
      str_detect(option, "^\\d+$"),
      paste0("Option ", option),
      option
    )
  ) %>%
  as.data.frame()

# final table
checkbox_label_map <- list(
  Selected = "Selected",
  not_selected_total = "Not selected",
  `Missing Response` = "Missing Response",
  `Missing: Skip Logic` = "Missing: Skip Logic"
)

checkbox_label_map <- checkbox_label_map[
  names(checkbox_label_map) %in% names(checkbox_summary)
]

if ("NA" %in% names(checkbox_summary)) {
  checkbox_label_map[["NA"]] <- html("NA")
}

checkbox_summary %>%
  select(
    group,
    option,
    variable,
    Selected,
    not_selected_total,
    any_of(c("Missing Response", "Missing: Skip Logic")),
    any_of("NA"),
    everything(),
    -any_of(c("Not Selected", "Not selected"))
  ) %>%
  gt(groupname_col = "group", rowname_col = "option") %>%
  cols_hide(columns = c(variable)) %>%
  cols_label(.list = checkbox_label_map) %>%
  fmt_number(columns = where(is.numeric), decimals = 0) %>%
  sub_missing(everything(), missing_text = "") %>%
  tab_options(
    table.font.size = px(12),
    data_row.padding = px(4)
  ) %>%
  opt_row_striping()
Selected Not selected Missing Response
commute_subsidy
Option 1 493 1,959 3,106
Option 2 286 2,166 3,106
Option 3 861 1,591 3,106
Option 4 129 2,323 3,106
Option 5 231 2,221 3,106
Option 6 495 1,957 3,106
Option 7 913 1,539 3,106
Option 996 571 1,881 3,106
Option 998 128 2,324 3,106
commute_subsidy_use
Option 1 330 1,422 3,806
Option 2 108 1,644 3,806
Option 3 738 1,014 3,806
Option 4 56 1,696 3,806
Option 5 138 1,614 3,806
Option 6 404 1,348 3,806
Option 7 867 885 3,806
Option 996 110 1,642 3,806
ethnicity
Option 1 3,986 867 705
Option 2 154 4,699 705
Option 3 25 4,828 705
Option 4 9 4,844 705
Option 997 76 4,777 705
Option 999 611 4,242 705
ev_typical_charge
Option 1 213 30 5,315
Option 2 30 213 5,315
Option 3 7 236 5,315
Option 4 61 182 5,315
Option 5 21 222 5,315
Option 6 7 236 5,315
Option 997 19 224 5,315
race_afam
race_afam 159 4,694 705
race_aiak
race_aiak 74 4,779 705
race_asian
race_asian 532 4,321 705
race_hapi
race_hapi 63 4,790 705
race_noanswer
race_noanswer 543 4,310 705
race_other
race_other 101 4,752 705
race_white
race_white 3,634 1,219 705
share
Option 1 2,951 1,901 706
Option 2 410 4,442 706
Option 3 1,360 3,492 706
Option 4 560 4,292 706
Option 5 30 4,822 706
Option 996 1,615 3,237 706
Table 33: Checkbox Variables in the Dataset

Example 1: Ethnicity Shares, Counting Multiple Selections

In this section, we demonstrate how to analyze multiple-response categorical variables (MRCVs) using the example of ethnicity shares. We will count the number of selections for each ethnicity category, allowing for multiple selections per person.

NoteUsing Regular Expressions to Identify MRCV Variables

In the code below, we identify all ethnicity-related variables by searching for variable names that start with ethnicity_ followed by a number. We do this with a regular expression:

For the regular expression ^ethnicity_[0-9]+$:

  • ^ asserts the start of the string
  • ethnicity_ matches the literal string
  • [0-9]+ matches one or more digits (0-9)
  • $ asserts the end of the string

This pattern is similar for many MRCVs in the dataset, where each option is represented by a separate binary variable (1 = selected, 0 = not selected, 995 = missing). Check out Mozilla’s Regular Expression Cheat Sheet for more on regular expressions.

Code
ethnicity_vars <- cb$variable_list %>%
  filter(person == 1, str_detect(variable, "^ethnicity_[0-9]+$")) %>%
  transmute(variable) %>%
  filter(variable %in% names(hts$person)) %>%
  left_join(
    cb$variable_list %>% select(variable, description),
    by = "variable"
  )

ethnicity_vars
##         variable                                              description
##           <char>                                                   <char>
## 1:   ethnicity_1  Ethnicity -- Not of Hispanic, Latino, or Spanish origin
## 2:   ethnicity_2          Ethnicity -- Mexican, Mexican American, Chicano
## 3:   ethnicity_3                                Ethnicity -- Puerto Rican
## 4:   ethnicity_4                                       Ethnicity -- Cuban
## 5: ethnicity_997 Ethnicity -- Another Hispanic, Latino, or Spanish origin
## 6: ethnicity_999                        Ethnicity -- Prefer not to answer

Next, we reshape the data from wide to long format to facilitate analysis. Each row in the resulting dataset represents a single person’s response to a specific ethnicity checkbox item. In the example below, person 5101 selected only ethnicity_1 and person 5102 did not provide ethnicity (as was the case for all children and nonrelatives).

Code
ethnicity_data_long <- hts$person %>%
  # keep only ethnicity_* columns
  select(person_id, all_of(ethnicity_vars$variable)) %>%
  # long format: variable = column name, value = 0/1/NA
  pivot_longer(
    cols = all_of(ethnicity_vars$variable),
    names_to = "variable",
    values_to = "value"
  )

ethnicity_data_long
## # A tibble: 33,348 × 3
##     person_id variable      value       
##       <int64> <chr>         <chr>       
##  1 2500000601 ethnicity_1   Selected    
##  2 2500000601 ethnicity_2   Not selected
##  3 2500000601 ethnicity_3   Not selected
##  4 2500000601 ethnicity_4   Not selected
##  5 2500000601 ethnicity_997 Not selected
##  6 2500000601 ethnicity_999 Not selected
##  7 2500000602 ethnicity_1   Selected    
##  8 2500000602 ethnicity_2   Not selected
##  9 2500000602 ethnicity_3   Not selected
## 10 2500000602 ethnicity_4   Not selected
## # ℹ 33,338 more rows

Note that the ethnicity columns contain labeled values such as Selected, Not selected, Prefer not to answer, and Missing Response.

As above, we remove both Prefer not to answer and “995” (Missing, not asked) responses from analysis (see Section 8.8.4 for additional guidance on this when using weights and generating confidence intervals). For checkbox variables like ethnicity, this needs to be done in two ways:

  • First, we remove persons who answer “Prefer not to Answer” (ethnicity_999 == "Selected").
  • Next, we remove missing-response records from the dataset.
  • Finally, we remove the ethnicity_999 variable itself, since it is no longer needed.

With checkbox variables like these, survey logic means that a person who checks “Prefer not to Answer” for their ethnicity cannot select other ethnicities; and missing responses for one ethnicity mean all ethnicity values are missing for that person.

Code
pnta_persons <- ethnicity_data_long %>%
  filter(variable == "ethnicity_999" & value == "Selected")

ethnicity_data_long <- ethnicity_data_long %>%
  filter(!person_id %in% pnta_persons$person_id) %>%
  filter(value != "Missing Response") %>%
  filter(variable != "ethnicity_999")
Table 34

Now we can summarize the data to get counts and proportions for each ethnicity option.

The summary below “double-counts” persons who selected multiple ethnicities; hence, the total of count_selected exceeds the number of persons who answered the question, and the sum of prop_selected exceeds 1 (100%).

Use the codebook to attach readable labels and preserve the intended checkbox order.

Code
ethnicity_summary <- ethnicity_data_long %>%
  group_by(variable) %>%
  summarise(
    count_persons = n(),
    count_selected = sum(value == "Selected", na.rm = TRUE),
    prop_selected = count_selected / count_persons,
    .groups = "drop"
  )
Table 35

Table 36 shows the resulting summary table with descriptions attached, formatted with the gt package.

Code
ethnicity_summary %>%
  left_join(ethnicity_vars, by = "variable") %>%
  mutate(
    description_short = dplyr::coalesce(description, stringr::str_remove(variable, "^ethnicity_")),
    variable = factor(variable, levels = ethnicity_vars$variable)
  ) %>%
  arrange(variable) %>%
  select(description_short, count_persons, count_selected, prop_selected) %>%
  gt() %>%
  gt::grand_summary_rows(
    columns = c(count_selected, count_persons),
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_number(., decimals = 0)
  ) %>%
  gt::grand_summary_rows(
    columns = prop_selected,
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_percent(., decimals = 1)
  ) %>%
  gt::fmt_percent(
    columns = prop_selected,
    decimals = 1
  ) %>%
  gt::fmt_number(
    columns = c(count_persons, count_selected),
    decimals = 0,
    sep_mark = ","
  ) %>%
  gt::cols_label(
    description_short = "Ethnicity",
    count_persons = "Count of Persons Responding",
    count_selected = "Count of Persons Selecting",
    prop_selected = "Percentage of Persons Selecting"
  )
Ethnicity Count of Persons Responding Count of Persons Selecting Percentage of Persons Selecting
Ethnicity -- Not of Hispanic, Latino, or Spanish origin 4,242 3,986 94.0%
Ethnicity -- Mexican, Mexican American, Chicano 4,242 154 3.6%
Ethnicity -- Puerto Rican 4,242 25 0.6%
Ethnicity -- Cuban 4,242 9 0.2%
Ethnicity -- Another Hispanic, Latino, or Spanish origin 4,242 76 1.8%
Total — 21,210 4,250 100.2%
Table 36: Ethnicity Summary

Example 2: Ethnicity Shares, Treating Multiple Selections as “Multiple”

Alternatively, we can code respondents who selected multiple ethnicities into a single “Multiple” category. With this approach, each person is counted only once in the summary.

We do this in the code below by first filtering to only the selected ethnicity variables (value == "Selected"), then grouping by person to count how many selections they made. If a person selected only one ethnicity, we keep that variable; if they selected more than one, we label them as “Multiple”.

Code
# For each person, find which ethnicity variables are selected
ethnicity_single_or_multiple <- ethnicity_data_long %>%
  filter(value == "Selected") %>%
  group_by(person_id) %>%
  summarise(
    count_selections = n(),
    ethnicity = if (count_selections == 1) variable else "Multiple",
    .groups = "drop"
  )

tail(ethnicity_single_or_multiple, 100)
## # A tibble: 100 × 3
##     person_id count_selections ethnicity  
##       <int64>            <int> <chr>      
##  1 2530748902                1 ethnicity_1
##  2 2530944701                1 ethnicity_1
##  3 2530944702                1 ethnicity_1
##  4 2530944703                1 ethnicity_1
##  5 2531023201                1 ethnicity_1
##  6 2531023202                1 ethnicity_1
##  7 2531023203                1 ethnicity_1
##  8 2531047501                1 ethnicity_1
##  9 2531294402                1 ethnicity_1
## 10 2531363801                1 ethnicity_2
## # ℹ 90 more rows

Table 37 shows the summary table with counts and proportions for each ethnicity. Now, the total of n equals the number of persons who answered the question, and the sum of pct equals 100%.

Code
ethnicity_vars <- dplyr::bind_rows(
  ethnicity_vars,
  data.frame(variable = "Multiple", description = "Multiple", stringsAsFactors = FALSE)
)

ethnicity_single_or_multiple %>%
  group_by(ethnicity) %>%
  summarise(
    count_persons = n(),
    .groups = "drop"
  ) %>%
  mutate(
    prop_persons = count_persons / sum(count_persons)
  ) %>%
  left_join(ethnicity_vars, by = c("ethnicity" = "variable")) %>%
  mutate(
    description_short = dplyr::case_when(
      ethnicity == "Multiple" ~ "Multiple",
      TRUE ~ description
    ),
    ethnicity = factor(ethnicity, levels = ethnicity_vars$variable)
  ) %>%
  arrange(ethnicity) %>%
  select(description_short, count_persons, prop_persons) %>%
  gt() %>%
  gt::grand_summary_rows(
    columns = c(count_persons),
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_number(., decimals = 0)
  ) %>%
  gt::grand_summary_rows(
    columns = prop_persons,
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_percent(., decimals = 1)
  ) %>%
  gt::fmt_percent(
    columns = prop_persons,
    decimals = 1
  ) %>%
  gt::fmt_number(
    columns = c(count_persons),
    decimals = 0,
    sep_mark = ","
  ) %>%
  gt::cols_label(
    description_short = "Ethnicity",
    count_persons = "Count of Persons",
    prop_persons = "Percentage of Persons"
  )
Ethnicity Count of Persons Percentage of Persons
Ethnicity -- Not of Hispanic, Latino, or Spanish origin 3,986 94.0%
Ethnicity -- Mexican, Mexican American, Chicano 148 3.5%
Ethnicity -- Puerto Rican 24 0.6%
Ethnicity -- Cuban 7 0.2%
Ethnicity -- Another Hispanic, Latino, or Spanish origin 69 1.6%
Multiple 8 0.2%
Total — 4,242 100.0%
Table 37: Ethnicity Summary (Multiple Selections as ‘Multiple’)

Missing Categorical Data

NoteWorking with Missing Categorical Data

When working with categorical variables (including MRCV fields), treat missingness and logic explicitly rather than silently recoding everything to “No.”

Always start from the codebook.
Check cb$variable_list and cb$value_labels first so you know the intended labels, ordering, table membership, and analytic universe before summarizing a field.

Handle special codes on purpose.

  • Treat “Don’t know / Prefer not to answer” as separate categories, not as 0s.
  • Be clear about your denominator: all sampled records, only in-universe records, or only respondents with valid answers.
  • For MRCVs, drop “Prefer not to answer” respondents at the person level first, then filter out 995s within the checkbox family.

Override question logic carefully.

  • Only recode out-of-universe cases to 0 when you can justify it from other variables (for example, if deliver_none indicates no deliveries, the other delivery checkbox fields should be treated as not selected).
  • Do not bulk-convert all “Missing Response” values to 0; that conflates “was not asked / did not answer” with a true “No.”

If you are using weights and need precise CIs, build the survey design on the full dataset and then filter the design object to your analytic universe. This preserves PSUs, strata, and correct variance estimation.

Count Variables with Top-Coding

Several variables in this dataset are capped with open-ended top bins, meaning the highest category is reported as “X or more” rather than an exact value. These include:

  • Household & people: hhsize, travelers_hh
  • Vehicles & equipment: vehicle_count, vehicle_year (capped at “1980 or earlier”)
  • Work & jobs: numworkers, num_jobs, tnc_work_hours

For example, the number of vehicles variable (vehicle_count) is coded as:

Code
vehicle_value_labels <- cb$value_labels %>%
  filter(table == "hh", variable == "vehicle_count") %>%
  arrange(val_order) %>%
  select(label)


gt(vehicle_value_labels)
label
0 (no vehicles)
1 vehicle
2 vehicles
3 vehicles
4 vehicles
5 vehicles
6 vehicles
7 vehicles
8 or more vehicles
8 vehicles
9 vehicles
10 or more vehicles
Table 38: Number of Vehicles Value Labels

How to analyze these variables:

  • Use frequency distributions or proportions (e.g., share of households with 3+ vehicles).
  • Report medians or category breakdowns rather than means.
  • If needed, combine categories into broader groups (e.g., 0, 1-2, 3+).
  • Avoid taking the mean or standard deviation, since the “or more” group has no defined upper bound.
NoteAnalytic Note on Household Size Variables

Household-level survey variables such as total household size and number of children are often highly skewed, with relatively few households reporting very large membership. Because precision decreases rapidly for categories with small unweighted sample sizes, modeling these variables as raw counts can lead to unstable estimates and inflated standard errors. To improve interpretability and statistical reliability, we recommend treating household size and number of children as binned categorical variables. For example, household size can be grouped into categories such as 1, 2, 3, 4, and 5+ people, and the number of children can be grouped into 0, 1, 2, and 3+ children. These binned versions better reflect the distribution of the survey data and reduce the impact of small-cell estimates, especially in multivariate or model-based analyses. Analysts examining household-level predictors of travel behavior should use these binned categories rather than raw counts to ensure more stable, design-consistent results.

Don’t calculate a mean:

Code
veh_year_wrong <-
  hts$vehicle %>%
  summarize(mean_model_year = mean(as.numeric(year), na.rm = TRUE))

veh_year_wrong
##   mean_model_year
## 1        2010.095

Create a frequency table instead:

Code
veh_summary <-
  hts$vehicle %>%
  group_by(year) %>%
  summarize(n = n()) %>%
  mutate(pct = 100 * n / sum(n))

gt(veh_summary)
year n pct
1980 41 0.139821983
1980 or earlier 342 1.166319954
1981 36 0.122770521
1982 28 0.095488183
1983 12 0.040923507
1984 31 0.105719060
1985 22 0.075026430
1986 55 0.187566074
1987 54 0.184155782
1988 64 0.218258705
1989 82 0.279643965
1990 113 0.385363026
1991 103 0.351260103
1992 91 0.310336596
1993 123 0.419465948
1994 217 0.740033421
1995 205 0.699109914
1996 215 0.733212836
1997 326 1.111755277
1998 378 1.289090475
1999 486 1.657402039
2000 602 2.052995942
2001 675 2.301947277
2002 702 2.394025168
2003 828 2.823721993
2004 959 3.270470279
2005 1088 3.710397981
2006 1209 4.123043345
2007 1314 4.481124032
2008 1240 4.228762405
2009 951 3.243187941
2010 1151 3.925246394
2011 1030 3.512601030
2012 1481 5.050642840
2013 1679 5.725880708
2014 1659 5.657674863
2015 1854 6.322681854
2016 1902 6.486375882
2017 1581 5.391672066
2018 1289 4.395866726
2019 895 3.052211575
2020 566 1.930225420
2021 527 1.797224022
2022 429 1.463015380
2023 375 1.278859598
2024 240 0.818470143
2025 72 0.245541043
2026 1 0.003410292

Or report the median model year:

Code
hts$vehicle %>%
  summarize(median_model_year = median(as.numeric(year), na.rm = TRUE))
##   median_model_year
## 1              2012

8.5 Working with Numeric (Integer, Continuous) Data

The HTS dataset contains several numeric variables, such as trip distances, durations, and speeds. These variables can sometimes include extreme values or outliers that may affect analysis results. When working with numeric data, follow these best practices to ensure accurate and meaningful analysis.

Check Delivered Variable Definitions First

Before analyzing any numeric variable, verify its meaning in the codebook first. Many errors stem from assuming what a variable represents rather than confirming it. Use the codebook to check table membership, variable descriptions, data types, and logic notes before turning to survey documentation. When reviewing metadata, check the following:

Units and Measurement

  • Confirm the unit (minutes vs. seconds, miles vs. kilometers, weekly vs. annual amounts).
  • Note any transformations (e.g., recodes, truncations, derived fields).
  • Convert or relabel units clearly in your summaries.

The data generally includes units in the variable names themselves (num_ for counts, pct_ for percentages, _seconds for seconds, _m for distance in meters, etc.). However, always double-check the codebook for clarity.

For example, the dataset includes several distance variables, each measured in different units and constructed in different ways:

Code
distance_fields <- names(hts$trip)[grepl("distance", names(hts$trip))]

data.frame(
  variable = distance_fields,
  stringsAsFactors = FALSE
)
##                  variable
## 1          distance_miles
## 2         distance_meters
## 3 distance_beeline_meters

Inspect the Data

Check for Missing Values

In the dataset, missing numeric data are encoded as NA. Identify any missing values in your numeric variables and decide how to handle them. Common strategies include:

  • Imputation (e.g., replacing with mean/median)
  • Exclusion (e.g., removing rows with missing values)
  • Flagging (e.g., creating a binary indicator for missingness)

For example, telework_time applies only to respondents age 16 or older who reported paid_work = "Yes", so some day records will still be missing that value even in a labeled delivery. By contrast, age is populated for essentially everyone in the person table.

Code
hts$day %>%
  summarize(
    total_records = n(),
    missing_num_trips = sum(is.na(num_trips)),
    missing_telework_time = sum(telework_time == "Missing Response"),
    pct_missing_num_trips = 100 * missing_num_trips / total_records,
    pct_missing_telework_time = 100 * missing_telework_time / total_records
  )
##   total_records missing_num_trips missing_telework_time pct_missing_num_trips
## 1         10868                 0                  5163                     0
##   pct_missing_telework_time
## 1                  47.50644

Visualize the Distribution

Before calculating any metric:

  • Generate a histogram, boxplot, or simple summary()
  • Check for skew, multimodality, or heavy tails
  • Identify minimum/maximum and question plausibility

This prevents misinterpretation and ensures the selected summary is appropriate.

Figure 26 shows the distribution of num_trips per day. Note the high number of responses with zero trips, as well as a long right tail with some days having many trips.

Code
plot_num_trips <- hts$day %>%
  ggplot(aes(
    x = num_trips,
    text = paste0("Trips: ", num_trips)
  )) +
  geom_histogram(binwidth = 1, fill = single_bar_color, color = "white") +
  labs(
    x = "Number of Trips",
    y = "Count of Days"
  ) +
  theme_minimal(base_family = "Inter") +
  theme(
    panel.grid.minor = element_blank()
  )

ggplotly(plot_num_trips, tooltip = "text") %>%
  config(displayModeBar = FALSE)
Figure 26: Distribution of Number of Trips per Day

Handle Outliers

Outliers can represent:

  • Real, meaningful extreme behavior
  • Data entry or GPS errors
  • Artifacts of device measurements

Decide how to handle outliers based on your analysis goals. Options include:

Binning Binning groups values into ranges can reduce the impact of extreme values. For example, we could create bins for num_trips such as 0-1, 2-3, 4-5, etc.

Code
day_binned <- hts$day %>%
  mutate(
    num_trips_binned = case_when(
      num_trips >= 0 & num_trips <= 1 ~ "0-1",
      num_trips >= 2 & num_trips <= 3 ~ "2-3",
      num_trips >= 4 & num_trips <= 5 ~ "4-5",
      TRUE ~ "6+"
    )
  )

Binning essentially transforms the variable from numeric to categorical. A number of techniques exist to address outliers while still preserving the continuous nature of the data, depending on the context and goals of the analysis. We present these below from most to least conservative.

Domain-Based Thresholds Domain-informed limits (e.g., physical speed limits, hours of the day) can be used when quantiles alone don’t reflect behavioral constraints.

For example, if analyzing trip speeds, we could exclude trips with an unusually high speed given their mode, made possible through the speed_flag variable in the trip table (see Section 4.2.1).

Table 39 shows the counts of trips by speed_flag value. Trips with speed_flag of 1 indicate implausible speeds given the mode.

Code
hts$trip %>%
  mutate(
    mode_label = dplyr::if_else(is.na(mode_class_5), "Missing Response", mode_class_5)
  ) %>%
  group_by(mode_label, speed_flag) %>%
  tally() %>%
  ungroup() %>%
  group_by(mode_label) %>%
  mutate(prop = n / sum(n)) %>%
  gt() %>%
  cols_label(
    mode_label = "Mode",
    speed_flag = "Speed Flag",
    n = "Trip Count",
    prop = "Proportion"
  ) %>%
  fmt_percent(
    columns = prop,
    decimals = 1
  ) %>%
  fmt_number(
    columns = n,
    decimals = 0,
    sep_mark = ","
  )
Speed Flag Trip Count Proportion
Bike/Micromobility
Missing Response 11 1.5%
No 696 97.1%
Yes 10 1.4%
Drive
Missing Response 925 3.1%
No 28,418 95.3%
Yes 464 1.6%
Missing Response
No 13 100.0%
Other
Missing Response 358 28.3%
No 888 70.1%
Yes 21 1.7%
Transit
Missing Response 38 2.3%
No 1,577 94.0%
Yes 63 3.8%
Walk
Missing Response 78 1.4%
No 4,921 89.5%
Yes 499 9.1%
Table 39: Trip Counts by Speed Flag

Winsorization Winsorization replaces values beyond a percentile with the cutoff value instead of removing them.

The example below identifies the 99.5th percentile of num_trips per day as an outlier threshold.

Code
outlier_value <- hts$day %>%
  summarize(p99_5 = quantile(num_trips, 0.995, na.rm = TRUE)) %>%
  pull(p99_5)

outlier_value
## [1] 16

To winsorize num_trips, we can cap values at the 99.5th percentile:

Code
day_winsorized <- hts$day %>%
  mutate(
    num_trips_winsorized = ifelse(
      num_trips > outlier_value,
      outlier_value,
      num_trips
    )
  )

Trimming

Trimming removes extreme values beyond a certain percentile. For example, we could remove the top 0.5% of num_trips values.

Code
day_trimmed <- hts$day %>%
  filter(num_trips <= outlier_value)

Other Approaches

Depending on the analytic context, additional techniques can help address outliers:

  • Transformations (log or square root) can reduce skew without removing or capping values, especially in modeling contexts.

  • Robust estimators (medians, MAD, quantile regression) provide summaries or models less influenced by extreme values and may eliminate the need for trimming or winsorization.

  • Longitudinal or within-household checks can detect outliers relative to a person’s typical behavior, not just the population as a whole.

These methods complement trimming, winsorization, and exclusion and are often used in combination.

NoteTakeaway: Outliers Require Intentional Strategies

There is no single correct way to handle outliers-your approach should match the analytic goal. Combine trimming, winsorization, transformations, robust statistics, and domain-informed checks to produce results that are both defensible and meaningful.

Example: Trip Speeds by Mode

The section below demonstrates several numeric-analysis principles using trip speeds as an example.

How do average trip speeds vary across travel modes, and how do outlier-handling strategies (trimming, winsorization, and mode-specific plausibility flags) affect the results?

Trip speed is a useful diagnostic variable, but it is also highly sensitive to outliers-both genuine extreme cases (e.g., long-distance rail) and implausibly high values caused by GPS noise or very short trips. This example illustrates how different approaches to handling extreme values can change summary statistics.

Because the goal is to demonstrate numeric-analysis principles rather than produce weighted population estimates, we use unweighted data.

First, we calculate raw means and medians by mode.

Code
trip_speed <- hts$trip %>%
  mutate(
    mode_label = dplyr::if_else(is.na(mode_class_5), "Missing Response", mode_class_5)
  ) %>%
  filter(
    !is.na(speed_mph),
    !is.na(mode_label)
  )
Code
speed_raw <- trip_speed %>%
  group_by(mode_label) %>%
  summarize(
    mean_raw = mean(speed_mph, na.rm = TRUE),
    median_raw = median(speed_mph, na.rm = TRUE)
  )

Here we remove the top 0.5% of observed speeds within each mode.

Code
speed_trimmed <- trip_speed %>%
  group_by(mode_label) %>%
  mutate(
    cutoff_trim = quantile(speed_mph, 0.995, na.rm = TRUE)
  ) %>%
  filter(speed_mph <= cutoff_trim) %>%
  summarize(
    mean_trimmed = mean(speed_mph, na.rm = TRUE),
    median_trimmed = median(speed_mph, na.rm = TRUE)
  )

The alternative approach below applies winsorization (capping values at upper 99.5th percentile).

Code
speed_winsorized <- trip_speed %>%
  group_by(mode_label) %>%
  mutate(
    cutoff_win = quantile(speed_mph, 0.995, na.rm = TRUE),
    speed_mph_win = pmin(speed_mph, cutoff_win)
  ) %>%
  summarize(
    mean_winsorized = mean(speed_mph_win, na.rm = TRUE),
    median_winsorized = median(speed_mph_win, na.rm = TRUE)
  )

We could otherwise remove implausible speeds (use the speed_flag variable) based on mode-specific thresholds.

Code
speed_flagged <- trip_speed %>%
  filter(speed_flag == "No") %>%
  group_by(mode_label) %>%
  summarize(
    mean_flagged = mean(speed_mph, na.rm = TRUE),
    median_flagged = median(speed_mph, na.rm = TRUE)
  )

Table 40 shows a side-by-side comparison of mean and median trip speeds by mode across the four outlier-handling methods: raw, trimmed, winsorized, and speed-flag filtered.

Across all modes, mean trip speeds are far more sensitive to extreme values than medians. Raw means tend to be inflated for nearly every mode, while medians remain stable across all outlier-handling methods, making them much more reliable for descriptive summaries.

Trimming-removing the extreme upper tail-produces the largest downward shift in means, especially for modes prone to occasional erroneous spikes. Winsorization has a more moderate effect, capping extremes while retaining all observations, which makes it a useful compromise when analysts want to preserve dataset size but still dampen outliers.

Applying mode-specific plausibility flags often yields results that align closely with real-world expectations, particularly for slower modes such as walking or biking. For many modes, the flagged means approach the median values, indicating that implausible high speeds were driving much of the inflation in the raw data.

Code
speed_compare <- speed_raw %>%
  left_join(speed_trimmed, by = "mode_label") %>%
  left_join(speed_winsorized, by = "mode_label") %>%
  left_join(speed_flagged, by = "mode_label")

speed_compare %>%
  filter(mode_label != "Missing Response") %>%
  gt(rowname_col = "mode_label") %>%
  # Now format with percent/number as needed
  fmt_number(
    columns = where(is.numeric),
    decimals = 1
  ) %>%
  # Group Mean columns
  tab_spanner(
    label = "Means",
    columns = c(mean_raw, mean_trimmed, mean_winsorized, mean_flagged)
  ) %>%
  # Group Median columns
  tab_spanner(
    label = "Medians",
    columns = c(median_raw, median_trimmed, median_winsorized, median_flagged)
  ) %>%
  # Rename columns for readability
  cols_label(
    mean_raw = "Raw",
    mean_trimmed = "Trimmed",
    mean_winsorized = "Winsorized",
    mean_flagged = "Flag-Filtered",
    median_raw = "Raw",
    median_trimmed = "Trimmed",
    median_winsorized = "Winsorized",
    median_flagged = "Flag-Filtered"
  ) %>%
  cols_label(
    mode_label = "Mode"
  )
Means
Medians
Raw Trimmed Winsorized Flag-Filtered Raw Trimmed Winsorized Flag-Filtered
Bike/Micromobility 8.7 8.6 8.7 8.6 8.1 8.1 8.1 8.1
Drive 20.3 20.0 20.3 20.3 18.7 18.7 18.7 18.8
Other 237.4 39.4 73.5 36.0 12.0 12.0 12.0 12.0
Transit 10.3 10.0 10.2 10.1 7.5 7.4 7.5 7.4
Walk 2.9 2.9 2.9 2.8 2.5 2.5 2.5 2.4
Table 40: Comparison of Trip Speed Summaries by Mode and Outlier-Handling Method

Overall, this comparison highlights that:

  • Means can be misleading in the presence of extreme values, especially for modes affected by GPS noise or very short trip durations.

  • Medians provide a more robust measure of central tendency and remain consistent across cleaning strategies.

  • Trimming and winsorization each offer systematic ways to reduce outlier influence, with trimming being more aggressive.

  • Behaviorally informed rules, such as mode-specific speed limits, often produce the most interpretable and defensible cleaned dataset.

The take-home message is that conclusions about relative speeds across travel modes can change substantially depending on how outliers are handled, so analysts should examine both means and medians and be explicit about their chosen outlier strategy.

8.6 Working with Date and Time Data

In this dataset, timestamps are stored in two different ways depending on the table:

  • Location table: All timestamps are recorded in Coordinated Universal Time (UTC), regardless of where the trip took place.
  • Trip table: Timestamps are recorded in local time (America/Los_Angeles Time).

To make the dataset easier to store and transfer across systems, timestamps have been split into separate fields for hour, minute, and second (e.g., depart_hour, depart_minute, depart_second in the trip table). This avoids problems that sometimes occur when software like Microsoft Excel interprets or reformats full datetime objects differently during import/export.

We use this approach because:

  • It prevents unexpected shifts caused by daylight savings adjustments or timezone conversions applied automatically by some software (e.g., Microsoft Excel).

  • Analysts can reconstruct a full timestamp if needed, or work with the components separately (e.g., for modeling departure-hour effects).

  • The UTC standard in the location table allows comparison across different regions (e.g., if a participant recorded a trip that crossed time zones) and avoids ambiguity.

Reconstructing Timestamps in R

Split fields can be recombined into a full time object using lubridate::make_datetime() or hms::hms(). For example:

Code
hts$trip$depart_datetime <- make_datetime(
  year = year(hts$trip$depart_date),
  month = month(hts$trip$depart_date),
  day = day(hts$trip$depart_date),
  hour = hts$trip$depart_time_hour,
  min = hts$trip$depart_time_minute,
  sec = hts$trip$depart_time_second,
  tz = "America/Los_Angeles" # Pacific time for trip table
)

head(hts$trip$depart_datetime)
## [1] "2025-03-17 11:00:00 PDT" "2025-03-17 11:40:00 PDT"
## [3] "2025-03-06 09:32:00 PST" "2025-03-06 10:30:00 PST"
## [5] "2025-03-06 10:47:00 PST" "2025-03-06 10:58:00 PST"

For the location table, use the same approach but set tz = "UTC".

Things to watch out for:

  • Time zone alignment: If location data (UTC) is merged with trip data (Pacific), convert one set so they are in the same time zone before analysis.
  • Daylight savings time: When reconstructing local times, be careful about DST changes–lubridate handles this if the correct tz is set.
  • Aggregation by time of day: If the hour of travel (e.g., peak vs. off-peak) is the focus, full timestamps likely do not need to be recombined.

8.7 Getting Started with Weights

Analyses designed to draw conclusions about travel behavior in the region (as opposed to just the survey respondents) should use weighted data.

When applied, the weights make the dataset representative of travel for residents within the study region for the time period studied.

NoteSurvey Coverage Limitations

Just a reminder that the dataset does not include:

  • Commercial vehicle travel
  • Travel for persons residing in group quarters outside of the address-based sample frame (e.g., college dorms, institutional housing)
  • Travel from non-residents (i.e., visitors to the region)
  • Seasonal/holiday travel outside of the survey fielding period.

Data users should also keep in mind the following when creating weighted statistics and summaries from HTS data:

  • Filter to the data relevant to your analysis. Note that not all people are asked every question, so understanding the ‘missing value’ codes and ‘survey logic’ in the data dictionary are important.
  • Remember the survey design when using and interpreting weighted values. For example, the study included both one-day online and call center participation and seven-day smartphone participation. Therefore, it is best to avoid filtering by day of week since not all participants traveled on all days.

Choosing the Right Weight

When deciding which weight to use for your analysis, consider the following:

  • Level of Analysis: Choose the weight that corresponds to the level of analysis you are conducting (e.g., household, person, trip).
  • Data Structure: Be mindful of the data structure and the relationships between different tables. Use the weight from the lowest level of the data hierarchy when combining variables from multiple tables.
  • Research Questions: Align the weight with your specific research questions and the population you aim to represent.

In general:

  • household weights should be used for household- and vehicle-level analyses
  • person weights for person-level analyses
  • day weights for day-level analyses, and
  • trip weights for trip-level analyses.

To calculate weighted summaries or descriptive statistics, sum the weights for that table.

In the case of an analysis that requires variables from two or more tables, the weight from the lowest level of the data hierarchy should be used (trip weights, then day weights, person weights, and finally household weights). For example, a weighted summary of race (a person-level variable) by household income should use person-level weights, unless constructing a household-level variable that summarizes race (e.g., presence of any non-white members in the household).

Trip rates require special consideration; see Section 8.9 for details.

Calculating Simple Weighted Estimates

In the simplest terms, calculating weighted estimates can be done by multiplying each observation by its corresponding weight and then summing the results, or by totaling the weights for proportions.

Calculating weighted estimates in R can be done using the {dplyr} package along with base R functions. Here’s a simple example of calculating the estimated number and proportion of households by the number of vehicles owned using household weights.

Code
hh_vehicle_weighted_summary <- hts$hh %>%
  mutate(
    vehicle_count = factor(vehicle_count, levels = vehicle_value_labels$label, ordered = TRUE)
  ) %>%
  group_by(vehicle_count) %>%
  summarize(
    unweighted_count = n(),
    unweighted_proportion = unweighted_count / nrow(hts$hh),
    weighted_count = sum(hh_weight),
    weighted_proportion = weighted_count / sum(hts$hh$hh_weight)
  )

Table 41 shows the resulting weighted and unweighted estimates using the labeled vehicle_count values provided in this delivery.

Code
hh_vehicle_weighted_summary %>%
  select(
    vehicle_count,
    unweighted_count,
    unweighted_proportion,
    weighted_proportion,
    weighted_count
  ) %>%
  gt() %>%
  gt::grand_summary_rows(
    columns = c(
      "unweighted_count",
      "weighted_count"
    ),
    label = "Total",
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_number(., use_seps = TRUE, decimals = 0)
  ) %>%
  gt::grand_summary_rows(
    columns = c(
      "unweighted_proportion",
      "weighted_proportion"
    ),
    label = "Total",
    fns = list(label = "Total", fn = "sum"),
    fmt = ~ fmt_percent(., decimals = 1)
  ) %>%
  gt::fmt_percent(
    columns = c(unweighted_proportion, weighted_proportion),
    decimals = 1
  ) %>%
  gt::fmt_number(
    columns = c(unweighted_count, weighted_count),
    decimals = 0,
    sep_mark = ","
  ) %>%
  gt::cols_label(
    vehicle_count = "Number of Vehicles",
    unweighted_count = "Unweighted Count",
    unweighted_proportion = "Unweighted Percent",
    weighted_proportion = "Weighted Percent",
    weighted_count = "Weighted Count"
  )
Number of Vehicles Unweighted Count Unweighted Percent Weighted Percent Weighted Count
0 (no vehicles) 226 8.2% 8.7% 151,671
1 vehicle 1,055 38.1% 32.0% 560,536
2 vehicles 1,015 36.6% 39.8% 696,103
3 vehicles 334 12.0% 14.1% 247,566
4 vehicles 93 3.4% 3.2% 56,387
5 vehicles 31 1.1% 1.4% 23,990
6 vehicles 7 0.3% 0.1% 2,184
7 vehicles 8 0.3% 0.6% 9,933
8 or more vehicles 3 0.1% 0.1% 2,221
Total — 2,772 100.0% 100.0% 1,750,591
Table 41: Household Counts and Proportions by vehicle Category (Weighted and Unweighted)

8.8 Using Survey-Aware Methods for Inference

In many cases, simple weighted calculations-multiplying each observation by its weight and then summing or averaging-are sufficient for producing basic totals or proportions. However, these direct calculations do not account for the full structure of the survey’s sampling approach. Although the study is not as complex as some multi-stage, clustered national surveys, it is still a complex survey design because it incorporates unequal probabilities of selection, geography-based oversampling, and multi-step nonresponse adjustments. These features affect both the point estimates (means, proportions) and the accuracy of statistical inference (confidence intervals, model results).

When Do You Need Survey-Aware Methods?

When Filtering/Subsetting and You Care About Uncertainty Measures (CIs, SEs). Households in different geographic and oversample strata had different probabilities of being selected, which required correction through base weights. Additional nonresponse adjustments at the household, person, and trip levels further changed the distribution of weights and the effective sample size. When analysts filter the data-especially for fine-scaled geographic or detailed demographic subgroups-they alter the implicit design structure. Simple weighted calculations do not preserve the relationship between the filtered records and the full designed sample, which can lead to biased estimates or misleading conclusions.

When You Want to Compare Groups or Estimate Uncertainty. Any analysis requiring uncertainty estimation-such as standard errors, confidence intervals, relative standard errors (RSE), comparisons between MPO regions, or demographic comparisons-must use survey-aware methods. RSG uses Taylor series linearization for variance estimation, implemented in the {samplics} and {srvyr} packages. These methods correctly incorporate the effects of weighting and nonresponse adjustments into the standard errors, something that direct weighted calculations cannot do.

When Using Models for Inference. Survey-aware modeling is also essential. Functions like svyglm() ensure that regression coefficients, interaction terms (e.g., race-mode or MPO-income), hypothesis tests, and confidence intervals all reflect the survey design. This is particularly important when comparing population groups or estimating policy-relevant differences.

In short, use simple weighted calculations for basic descriptive summaries, but rely on statistical packages like {srvyr} or {samplics} whenever analyses involve subsetting, comparisons across groups, inference, uncertainty estimation, or modeling. These tools preserve the integrity of the sample design and ensure results that are statistically defensible and generalizable.

NoteWhat is the Taylor Series Linearization Method?

There are two common approaches for variance estimation using weights. The first, replicate weights, requires the data user to leverages multiple sets of weights for the same set of observations. Analyses using each set of weights are “averaged” to create variance estimates. This procedure has major drawbacks, because the replicates increase dataset complexity and add require complex weighting procedures.

The preferred approach is an approximation using Taylor-series linearization. This procedure approximates the variance using a simpler to implement formula, simplifying the dataset and weight generation. Most survey statistics-like weighted means, proportions, totals, or regression coefficients-can be thought of as functions of the underlying weighted data. Linearization asks a simple question:

If the data changed just a little bit, how much would the estimate change?

Rather than repeatedly re-estimating our statistic (as replicate weights do), linearization uses this sensitivity idea to approximate the amount of variation we would see across many possible samples.

This method is widely used in survey analysis software, including R packages like {survey} and {srvyr}, Stata’s svy commands, and SAS’s survey procedures. It provides a practical way to account for complex survey designs while estimating variances, standard errors, and confidence intervals.

Specifying the Survey Design

The {srvyr} package requires a survey design object, which defines how the data were generated and how observations relate to one another. This structure is used to correctly estimate totals, proportions, and standard errors.

A well-specified survey design includes the following components:

  • ids (Primary Sampling Unit / PSU) This specifies the unit at which observations are assumed to be independent. In this survey, the household (household_id) is the PSU, because households were the unit of recruitment and sampling.

    Even when analyzing person-, day-, or trip-level data, observations remain clustered within households. Therefore, household_id should always be used as the PSU. Specifying a lower-level unit (e.g., person_id or trip_id) would incorrectly treat correlated observations as independent and may underestimate standard errors.

    It is possible to specify nested IDs (e.g., ids = c(household_id, person_id, day_id, trip_id)), but in practice this typically has little effect on estimates or standard errors, since variance is primarily driven by clustering at the household level.

  • weights (Survey weights) This specifies the weight variable appropriate for the unit of analysis. Weights reflect the inverse probability of selection, adjusted for nonresponse and calibrated to match population control totals.

    The correct weight depends on the analysis level (e.g., household, person, day, or trip). See Section 8.7.1 for guidance on selecting the appropriate weight, particularly when combining data across levels.

  • strata (Stratification variable) This defines groups within which sampling or weighting was structured. In an ideal design-based framework, strata correspond to the original sampling strata. In PSRC’s case, this is the sample_segment variable on the household table. This is the default stratification variable used in this guide.

    Strata can be skipped entirely, especially to speed up point processing time on estimate calculations, but this may lead to slightly less accurate standard errors.

For example, the household-level survey design can be defined in {srvyr} as:

Code
hh_design <- hts$hh %>%
  as_survey_design(
    ids = household_id,
    strata = sample_segment,
    weights = hh_weight
  )

Analyst Tip: Always specify household_id as the PSU, even for trip-level analysis. This ensures that standard errors account for clustering of observations within households.

This ensures that the variance estimation correctly reflects the effective sample size (and avoids {survey} warnings).

You may also need to join sample_segment from the household table to the day or trip table before defining the design, since these tables do not include it by default. If you also want county summaries, join home_county alongside it and use county only as the reporting domain.

Code
# Define the design
trip_design <- hts$trip %>%
  left_join(
    hts$hh %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  filter(trip_weight > 0) %>%
  as_survey_design(
    ids = household_id,
    weights = trip_weight,
    strata = sample_segment
  )

Finally, we set the options of srvyr to use lonely.psu = "adjust" by default. This ensures that sample-design strata with only one PSU are handled correctly during variance estimation after filtering or crosstabulating the data.

Code
options(srvyr.lonely.psu = "adjust")

In this guide, sample_segment is the default stratification variable because it reflects the sample design. This is separate from the weighting zone groups described in Section 5, which are grouped PUMAs used to build weighting controls rather than the default variance-estimation strata for analyst workflows.

If you want to see how much the strata choice matters in practice, you can compare the same estimate with and without sample_segment. The point estimate will usually be identical or nearly identical, while the standard error and confidence interval may shift modestly because the stratified design better preserves the original sample structure.

Code
zero_vehicle_compare <- dplyr::bind_rows(
  hts$hh %>%
    filter(hh_weight > 0) %>%
    as_survey_design(
      ids = household_id,
      weights = hh_weight
    ) %>%
    summarize(
      estimate = survey_mean(vehicle_count == "0 (no vehicles)", vartype = c("se", "ci"))
    ) %>%
    mutate(design = "Without sample_segment strata"),
  hts$hh %>%
    filter(hh_weight > 0) %>%
    as_survey_design(
      ids = household_id,
      weights = hh_weight,
      strata = sample_segment
    ) %>%
    summarize(
      estimate = survey_mean(vehicle_count == "0 (no vehicles)", vartype = c("se", "ci"))
    ) %>%
    mutate(design = "With sample_segment strata")
) %>%
  select(design, estimate, estimate_se, estimate_low, estimate_upp)

zero_vehicle_compare %>%
  gt() %>%
  fmt_percent(columns = c(estimate, estimate_low, estimate_upp), decimals = 4) %>%
  fmt_number(columns = estimate_se, decimals = 4) %>%
  cols_label(
    design = "Design",
    estimate = "Share of Zero-Vehicle Households",
    estimate_se = "Standard error (SE)",
    estimate_low = "Confidence interval (CI) lower bound",
    estimate_upp = "Confidence interval (CI) upper bound"
  )
Design Share of Zero-Vehicle Households Standard error (SE) Confidence interval (CI) lower bound Confidence interval (CI) upper bound
Without sample_segment strata 8.6640% 0.0091 6.8816% 10.4463%
With sample_segment strata 8.6640% 0.0090 6.9075% 10.4205%

Using the Survey Design for Weighted Estimates

In the code below, we use the household survey design defined above to re-calculate weighted and unweighted counts and proportions of households by number of vehicles owned. This time, we use the {srvyr} functions survey_total() and survey_prop() to calculate weighted totals and proportions along with standard errors (vartype = se) and 95% confidence intervals (vartype = ci). In survey_prop, we specify proportion = TRUE to indicate that our data are proportions rather than counts; this is optional (and can be time-intensive) but recommended for the most accurate CI estimation.

Code
num_vehicles_summary_srvyr <- hh_design %>%
  mutate(
    vehicle_count = factor(vehicle_count, levels = vehicle_value_labels$label, ordered = TRUE)
  ) %>%
  group_by(vehicle_count) %>%
  summarize(
    unweighted_count = n(),
    unweighted_proportion = unweighted_count / nrow(hts$hh),
    weighted_count = survey_total(vartype = c("ci", "se"), level = 0.95),
    weighted_proportion = survey_prop(
      vartype = c("ci", "se"),
      level = 0.95,
      proportion = TRUE
    )
  )

Table 42 shows the resulting weighted and unweighted estimates using the delivered vehicle_count labels. The weighted counts and proportions include 95% confidence intervals.

Code
num_vehicles_summary_pretty <- num_vehicles_summary_srvyr %>%
  mutate(
    # Compute RSEs
    rse_total = weighted_count_se / weighted_count,
    rse_prop = weighted_proportion_se / weighted_proportion,

    # Asterisk flags
    total_ast = case_when(
      rse_total > 0.50 ~ "<strong>**</strong>",
      rse_total > 0.30 ~ "<strong>*</strong>",
      TRUE ~ ""
    ),
    prop_ast = case_when(
      rse_prop > 0.50 ~ "<strong>**</strong>",
      rse_prop > 0.30 ~ "<strong>*</strong>",
      TRUE ~ ""
    ),

    # Display strings
    weighted_count_display = sprintf(
      "%s <span style='color:%s;'>(%s - %s)</span>%s",
      formatC(weighted_count, format = "f", digits = 0, big.mark = ","),
      subtitle_color,
      formatC(weighted_count_low, format = "f", digits = 0, big.mark = ","),
      formatC(weighted_count_upp, format = "f", digits = 0, big.mark = ","),
      total_ast
    ),
    weighted_prop_display = sprintf(
      "%s%% <span style='color:%s;'>(%.1f%% - %.1f%%)</span>%s",
      sprintf("%.1f", weighted_proportion * 100),
      subtitle_color,
      weighted_proportion_low * 100,
      weighted_proportion_upp * 100,
      prop_ast
    )
  )

final_table <- num_vehicles_summary_pretty %>%
  select(
    vehicle_count,
    unweighted_count,
    weighted_count_display,
    weighted_prop_display
  )

final_table %>%
  gt() %>%
  fmt_number(columns = unweighted_count, decimals = 0) %>%
  cols_label(
    vehicle_count = "Number of Vehicles",
    unweighted_count = "Unweighted Households",
    weighted_count_display = md(
      sprintf("Weighted Households <span style='color:%s;'>(confidence interval, CI)</span>", subtitle_color)
    ),
    weighted_prop_display = md(
      sprintf("Weighted Percent <span style='color:%s;'>(confidence interval, CI)</span>", subtitle_color)
    )
  ) %>%
  fmt_markdown(
    columns = c("weighted_count_display", "weighted_prop_display")
  )
Number of Vehicles Unweighted Households Weighted Households (confidence interval, CI) Weighted Percent (confidence interval, CI)
0 (no vehicles) 226 151,671 (120,902 - 182,440) 8.7% (7.1% - 10.6%)
1 vehicle 1,055 560,536 (500,306 - 620,766) 32.0% (28.8% - 35.4%)
2 vehicles 1,015 696,103 (613,677 - 778,529) 39.8% (36.1% - 43.6%)
3 vehicles 334 247,566 (191,639 - 303,493) 14.1% (11.4% - 17.4%)
4 vehicles 93 56,387 (33,496 - 79,278) 3.2% (2.2% - 4.8%)
5 vehicles 31 23,990 (4,814 - 43,166)* 1.4% (0.6% - 3.0%)*
6 vehicles 7 2,184 (-748 - 5,116)** 0.1% (0.0% - 0.5%)**
7 vehicles 8 9,933 (-3,463 - 23,330)** 0.6% (0.1% - 2.2%)**
8 or more vehicles 3 2,221 (-1,954 - 6,395)** 0.1% (0.0% - 0.8%)**
Table 42: Household Counts and Proportions by Vehicle Category (Weighted and Unweighted, with Reliability Flags)

Filtering Data vs. Filtering the Survey Design

NoteTLDR

If you want to analyze a subpopulation, and you want correct standard errors / confidence intervals, always create the survey design object first, then filter the design object to the subpopulation of interest.

When working with weighted survey data, analysts frequently want to examine a specific subpopulation, such as trips that depart after 12 PM or households within a particular region. Although this type of filtering seems straightforward, the timing of the filter-whether it happens before or after the survey design object is created-has important consequences for variance estimation. The overall point estimates usually remain the same, but the standard errors and confidence intervals can shift depending on how much information from the original sample design is preserved.

The preferred approach is to define the survey design on the full set of valid observations and then apply subpopulation filters to the survey design object. When the design is created first, the variance estimator continues to use the full structure of the sample: all households remain included as primary sampling units (PSUs), all strata are represented, and the degrees of freedom reflect the original sampling geometry. Even if only a small fraction of the data is ultimately analyzed, the design object still “knows” about the complete set of PSUs and strata, which is essential for accurate standard errors under Taylor linearization.

Filtering the raw data before defining the survey design produces a very different outcome. In this case, any PSU that does not meet the filter condition is dropped entirely, even if it was part of the sampling frame. The same happens for entire strata: if no records from a stratum remain after filtering, the variance estimator treats that stratum as if it never existed. The result is a survey design that is artificially narrow, with fewer PSUs, fewer strata, and reduced degrees of freedom. The weighted means or proportions are often unchanged-because they depend only on weights and values within the domain-but the confidence intervals can be misleadingly wide or narrow because the underlying design information has been lost.

This distinction matters most when the filter removes whole households or entire regions. For example, filtering to only border-crossing trips before defining the design may remove many PSUs altogether, because many households never produce such a trip. Filtering by region before building the design can also remove entire strata. In both cases, the variance estimator constructed on the restricted data is no longer aligned with the true sampling design.

The recommended workflow is therefore to construct the design object first and then filter. For example:

Code
trip_design_all <- hts$trip %>%
  mutate(
    long_trip = ifelse(
      distance_miles >= 10,
      1,
      0
    )
  ) %>%
  left_join(hts$hh %>% select(household_id, sample_segment), by = "household_id") %>%
  filter(trip_weight > 0) %>%
  mutate(
    mode_label = dplyr::if_else(is.na(mode_class_5), "Missing Response", mode_class_5)
  ) %>%
  as_survey_design(
    ids = household_id,
    strata = sample_segment,
    weights = trip_weight
  )

mode_share_correct <- trip_design_all %>%
  filter(long_trip == 1) %>%
  group_by(mode_label) %>%
  summarize(
    n_hh = n_distinct(household_id),
    n_days = n_distinct(day_id),
    n_trips = n(),
    weighted_count = survey_total(vartype = c("ci", "se"), level = 0.95),
    weighted_prop = survey_prop(
      vartype = c("ci", "se"),
      level = 0.95,
      proportion = TRUE
    ),
    .groups = "drop"
  )
Code
mode_share_correct %>%
  mutate(mode_type_fct = mode_label) %>%
  mutate(
    # Compute RSEs
    rse_total = weighted_count_se / weighted_count,
    rse_prop = weighted_prop_se / weighted_prop,

    # Asterisk flags
    total_ast = case_when(
      rse_total > 0.50 ~ "<strong>**</strong>",
      rse_total > 0.30 ~ "<strong>*</strong>",
      TRUE ~ ""
    ),
    prop_ast = case_when(
      rse_prop > 0.50 ~ "<strong>**</strong>",
      rse_prop > 0.30 ~ "<strong>*</strong>",
      TRUE ~ ""
    ),

    # Display strings
    weighted_count_display = sprintf(
      "%s <span style='color:%s;'>(%s - %s)</span>%s",
      formatC(weighted_count, format = "f", digits = 0, big.mark = ","),
      subtitle_color,
      formatC(weighted_count_low, format = "f", digits = 0, big.mark = ","),
      formatC(weighted_count_upp, format = "f", digits = 0, big.mark = ","),
      total_ast
    ),
    weighted_prop_display = sprintf(
      "%s%% <span style='color:%s;'>(%.1f%% - %.1f%%)</span>%s",
      sprintf("%.1f", weighted_prop * 100),
      subtitle_color,
      weighted_prop_low * 100,
      weighted_prop_upp * 100,
      prop_ast
    )
  ) %>%
  select(
    mode_type_fct,
    n_hh, n_days, n_trips,
    weighted_count_display,
    weighted_prop_display
  ) %>%
  gt() %>%
  fmt_number(
    columns = c(n_hh, n_days, n_trips),
    decimals = 0,
    use_seps = TRUE
  ) %>%
  cols_label(
    mode_type_fct = "Mode",
    n_hh = "Unweighted Households",
    n_days = "Unweighted Days",
    n_trips = "Unweighted Trips",
    weighted_count_display = md(sprintf("Weighted Trips <span style='color:%s;'>(confidence interval, CI)</span>", subtitle_color)),
    weighted_prop_display = md(sprintf("Weighted Share <span style='color:%s;'>(confidence interval, CI)</span>", subtitle_color))
  ) %>%
  fmt_markdown(columns = ends_with("_display")) %>%
  tab_source_note(
    md(
      "**Notes:** <br>
      `*` indicates estimates with high relative standard error (RSE; 30-50%); <br>
      `**` indicates very high relative standard error (RSE > 50%). <br>
      Estimates with one or two asterisks should be interpreted with caution, <br>
      and those with two asterisks are generally not suitable for fine-grained comparisons."
    )
  )
Mode Unweighted Households Unweighted Days Unweighted Trips Weighted Trips (confidence interval, CI) Weighted Share (confidence interval, CI)
Bike/Micromobility 10 13 21 2,944 (164 - 5,723)* 0.1% (0.0% - 0.2%)*
Drive 1,212 2,147 4,443 2,988,159 (2,598,274 - 3,378,044) 91.5% (88.9% - 93.6%)
Other 79 97 159 123,022 (66,230 - 179,814) 3.8% (2.4% - 5.9%)
Transit 130 187 311 139,129 (93,816 - 184,442) 4.3% (3.0% - 5.9%)
Walk 7 8 8 11,244 (-3,885 - 26,374)** 0.3% (0.1% - 1.3%)**
Notes:
* indicates estimates with high relative standard error (RSE; 30-50%);
** indicates very high relative standard error (RSE > 50%).
Estimates with one or two asterisks should be interpreted with caution,
and those with two asterisks are generally not suitable for fine-grained comparisons.
Table 43: Mode Share for Trips of 10 Miles or More (Weighted and Unweighted, with Reliability Flags)

Calculating Estimate Reliability

In addition to reporting standard errors and confidence intervals, survey analysts often use Relative Standard Errors (RSEs) to assess the reliability of survey estimates. The RSE expresses the uncertainty of an estimate as a proportion of the estimate itself, making it easy to compare precision across variables measured on very different scales. An estimate with a small RSE has relatively little uncertainty, while a large RSE indicates that sampling variability is high relative to the estimate’s magnitude. This metric is especially useful for identifying unstable or sparse categories-such as rare modes, uncommon trip purposes, or small demographic groups-where point estimates may look reasonable but are supported by few observations.

RSEs provide a practical framework for deciding whether an estimate is reliable enough to report, whether it requires a cautionary note, or whether it is too imprecise to publish. Many transportation surveys and federal statistical agencies adopt simple RSE-based quality tiers to help analysts interpret results consistently. These guidelines are not prescriptive rules, but they help ensure that outputs reflect the inherent uncertainty in the sample rather than overstating precision.

Common approaches include:

  • RSE < 30%: Estimate is generally considered reliable and may be reported without qualification.
  • 30% <= RSE < 50%: Estimate is usable but should be accompanied by a caution that sampling error is substantial.
  • RSE >= 50%: Estimate is typically considered unreliable; analysts may suppress it or combine categories to improve stability.
  • Extreme RSEs often arise for small subgroups, rare behaviors, or estimates based on few PSUs, and may indicate the need to revisit the domain definition or weighting approach.

When presenting results, it is good practice to compute and review RSEs alongside confidence intervals, especially for modes or subpopulations with low sample sizes. RSEs help ensure that key findings reflect genuine patterns in the data rather than noise from sampling variability.

Example: Relative Standard Errors for “Mode Share by County”

Relative Standard Errors are especially recommended when conducting analysis for geographic subgroups such as counties, where sample size can be small. In the case of mode share specifically, even if each region has a large number of trips, the distribution of modes may differ substantially, and some counties may contribute very few observations for certain modes. In these cases, RSEs help determine whether the weighted mode share estimates are stable enough to interpret and report.

The example below computes mode share by county using the full trip-level survey design. For each county-mode combination, we estimate the weighted proportion of trips and its standard error using Taylor linearization. The RSE quantifies the amount of sampling uncertainty relative to the estimate itself. Higher RSEs typically occur for modes with low usage in smaller counties, and those estimates should be treated cautiously.

NoteWhy is {srvyr} Slow?

The {srvyr} package provides a user-friendly interface for survey analysis in R, but it can be slower than other methods, especially for large datasets or complex survey designs - and for categorical summaries with many levels, like mode type.

Python users can see speed improvements by using the {samplics} package, which is optimized for performance with large survey datasets. It provides similar (though simplified) functionality to {srvyr} but is designed to handle larger data volumes more efficiently.

Code
# 1. Build trip-level survey design using the FULL dataset ----
trip_design_all <- hts$trip %>%
  left_join(hts$hh %>% select(household_id, home_county, sample_segment), by = "household_id") %>%
  filter(trip_weight > 0) %>%
  mutate(
    mode_label = dplyr::if_else(is.na(mode_class_5), "Missing Response", mode_class_5)
  ) %>%
  as_survey_design(
    ids     = household_id, # PSU: household
    strata  = sample_segment, # sample-design strata
    weights = trip_weight
  )

# 2. Compute weighted mode share by county with SEs and RSEs ----
mode_share_rse <- trip_design_all %>%
  filter(mode_label != "Missing Response") %>%
  group_by(home_county, mode_label) %>%
  summarize(
    n = n(),
    n_hh = n_distinct(household_id),
    mode_share = survey_prop(
      vartype    = "se",
      proportion = TRUE
    )
  ) %>%
  mutate(
    rse = (mode_share_se / mode_share)
  ) %>%
  arrange(desc(rse))

Table 44 formats these results for presentation. It includes the weighted mode share, standard error, and RSE for each county-mode combination, while preserving sample_segment as the survey-design stratification variable.

Code

mode_share_rse %>%
  select(
    home_county,
    mode_label,
    n,
    n_hh,
    mode_share,
    mode_share_se,
    rse
  ) %>%
  gt() %>%
  # Format proportions as percents
  fmt_percent(
    columns = c(mode_share, mode_share_se),
    decimals = 2
  ) %>%
  fmt_percent(
    columns = c(rse),
    decimals = 1
  ) %>%
  fmt_number(
    columns = c(n, n_hh),
    decimals = 0
  ) %>%
  # Apply text color formatting based on RSE flag
  tab_style(
    style = list(
      cell_text(color = "red")
    ),
    locations = cells_body(
      rows = rse >= 0.5,
      columns = everything()
    )
  ) %>%
  tab_style(
    style = list(
      cell_text(color = "orange") # "yellow" is often too light; orange reads clearly
    ),
    locations = cells_body(
      rows = rse >= 0.3 & rse < 0.5,
      columns = everything()
    )
  ) %>%
  cols_label(
    home_county = "County",
    mode_label = "Mode",
    n = "Unweighted Trip Count",
    n_hh = "Unweighted Household Count",
    mode_share = "Weighted Share",
    mode_share_se = "Standard error (SE)",
    rse = "Relative standard error (RSE, %)"
  ) %>%
  tab_header(
    title = "Mode Share by County",
    subtitle = "Color-coded Reliability Based on Relative Standard Error"
  )
Mode Share by County
Color-coded Reliability Based on Relative Standard Error
Mode Unweighted Trip Count Unweighted Household Count Weighted Share Standard error (SE) Relative standard error (RSE, %)
Kitsap County
Bike/Micromobility 21 2 1.27% 1.16% 90.9%
Transit 34 13 0.69% 0.38% 54.7%
Other 13 9 3.24% 1.53% 47.2%
Walk 104 22 3.41% 1.54% 45.3%
Drive 976 116 91.38% 2.71% 3.0%
Snohomish County
Bike/Micromobility 21 7 0.50% 0.41% 81.9%
Transit 74 27 1.29% 0.55% 43.1%
Other 75 27 4.25% 1.52% 35.7%
Walk 290 60 5.51% 1.57% 28.4%
Drive 2,292 276 88.45% 2.44% 2.8%
Pierce County
Bike/Micromobility 79 22 0.88% 0.45% 51.6%
Transit 168 59 2.56% 0.78% 30.4%
Other 384 139 3.44% 0.77% 22.4%
Walk 574 157 4.50% 0.87% 19.4%
Drive 9,690 1,031 88.62% 1.54% 1.7%
King County
Bike/Micromobility 356 100 2.21% 0.52% 23.7%
Other 305 116 4.35% 0.90% 20.8%
Transit 983 282 4.93% 0.52% 10.6%
Walk 2,552 459 14.37% 1.18% 8.2%
Drive 7,127 842 74.14% 1.63% 2.2%
Table 44: Mode Share by County with Relative Standard Errors (RSEs). Red text indicates estimates that should be suppressed (RSE >= 50%), orange text indicates estimates that should be reported with caution (30-50%).

Across counties in the central Puget Sound region, mode share estimates vary widely in their reliability because many modes have very small sample sizes in some areas. Modes such as car travel and walking generally show low Relative Standard Errors (RSEs), reflecting both large unweighted counts and much higher representation in the weighted population. In contrast, modes with very low usage-car-share, scooter-share, shuttle/vanpool, long-distance passenger modes, and various micro-mobility services-often have extremely high RSEs, sometimes exceeding 50 percent. These high RSEs do not necessarily mean the estimates are incorrect; instead they signal that the survey provides limited information about those modes within that region, and any reported estimates should be interpreted cautiously.

RSE patterns also naturally differ across counties. Larger regions with more households and more diverse mode usage tend to produce more stable estimates for most modes, while smaller counties often have very few observations for anything beyond car and walk trips. For these smaller regions, even modes that appear plausible may have RSEs in the cautionary (30-50 percent) or unreliable (50+ percent) ranges simply because the underlying sample is sparse. When RSEs are high, analysts may consider collapsing categories, suppressing rare modes, or focusing on broader geographic or behavioral groupings.

Because the table includes both unweighted and weighted counts, SEs, and RSEs, it provides a clear view of which mode-by-county combinations are well supported by the sample and which require extra care in reporting. This is an important step in ensuring that the narrative and visualizations in an analysis reflect the strength of the underlying data, rather than overstating precision for rare or low-incidence travel behaviors.

When Sample Sizes Are Small: Practical Steps for Working With Unreliable Estimates

Even when estimates are produced correctly using the survey design object, some results will inevitably be unstable. Small sample sizes, sparse categories, and highly variable weights can all inflate standard errors, widen confidence intervals, and produce large Relative Standard Errors (RSEs). These issues do not mean the dataset is flawed; rather, they reflect the limits of what the sample can reliably support. Analysts need strategies for handling these situations in a way that preserves interpretability and avoids overstating precision.

The first step is simply recognizing when estimates are weak. RSEs, confidence intervals, and unweighted counts all provide signals about reliability. Very wide confidence intervals or extremely large RSEs indicate that the data provide little information about the true population value. Often this arises because only a handful of households or trips fall into a given category-rare modes in small counties, infrequent trip purposes, or subgroups defined by multiple demographic variables. In these cases, the survey is still functioning correctly, but it cannot produce stable estimates for all possible cross-tabulations.

When faced with unreliable estimates, analysts have several practical options:

  • Combine categories. Adjacent or conceptually similar groups can be collapsed to increase sample size and stabilize estimates-for example, combining rare micro-mobility modes, grouping similar trip purposes, or collapsing detailed race categories to a higher level.
  • Report results with caution rather than suppressing them outright. Some agencies add footnotes for estimates with RSEs between 30-50%, indicating that results should be interpreted carefully.
    • Suppress estimates entirely when RSEs exceed a chosen threshold. Many statistical programs suppress or shade estimates with RSE >= 50%, reflecting extremely limited precision.
  • Use broader geographic or demographic domains. If a subgroup is too small, analyzing the same metric for an entire MPO, county, or demographic group may yield stable estimates even when finer cuts do not.
  • Focus on direction rather than magnitude. In some cases, the presence or absence of a behavior (e.g., a rare commute mode) may be more meaningful than the exact percentage.
  • Check unweighted counts. A cell with only a handful of households or trips is unlikely to produce stable results, no matter how the weights behave. RSG’s rule of thumb is to avoid reporting estimates based on fewer than 30 unweighted PSUs (households).
  • Consider alternative modeling approaches. Weighted regression or multilevel models can sometimes borrow strength across groups, though these require more technical care.

These methods help ensure that the stories told with the data remain grounded in what the survey can reliably support. Analysts should not feel obligated to publish every estimate the dataset can technically produce. Instead, the goal is to choose representations-summary tables, visualizations, and analytic breakdowns-that reflect the strength of the underlying evidence and communicate uncertainty transparently.

8.9 Analysis of Trip Rates

Understanding trip rates-and the behaviors underlying them-requires aligning the unit of analysis with the survey’s hierarchical structure and weighting design. This section introduces the recommended analytic units for trip records and person-days, outlines how to calculate weighted trip rates correctly, and highlights key pitfalls to avoid.

Calculating Weighted Trip Rates

To calculate a weighted trip rate, divide the weighted count of trips by the weighted count of person-days. This ensures that both travelers and non-travelers are represented correctly.

For example:

If there are 300,000 weighted person-trips across 75,000 weighted person-days, the average weekday trip rate per person is 4.0.

If 225,000 of those trips are made by auto, then the auto trip rate is 3.0.

In the code below, we calculate the weighted household-level trip rate:

Code
weighted_hh_trip_rate <- sum(hts$trip$trip_weight, na.rm = T) /
  sum(hts$hh$hh_weight)

round(weighted_hh_trip_rate, 1)
## [1] 9.8

Why the Denominator (Household, Person, Day) Weights Matter

Trip rates depend on both the number of trips recorded and the number of diary days those trips came from. In the 2025 Puget Sound Regional Council Household Travel Study, neither quantity is observed uniformly. The Household, Person and Day weights created during weighting adjust for:

  • nonresponse bias (some types of days are more likely to be unobserved),
  • differing diary modes (rMove vs. web-only),
  • differing numbers of diary days per person (some persons provide 1 day, others provide up to 7).

Without day weights, persons who provided many diary days exert disproportionate influence, and persons who missed diary days look artificially low-travel. Even if zero-travel days are included, the unweighted mean is biased toward respondents with complete diaries.

Why Trip Weights Matter

Trip weights expand recorded trips to population-level trip totals. During weighting, we use a model-based approach to adjust these weights complements day-pattern weighting, that accounts for:

  • reporting errors and omissions,
  • differences in completeness across modes and purposes,
  • within-person differences in reporting between rMove and web modes,
  • calibration to regional/segment trip-rate targets.

A correct trip rate must therefore use:

  • trip weights in the numerator, and
  • day/household/person weights in the denominator.

Together, they reflect the true population of trips and the true population of person-days.

Why Zero-Travel Days Matter

Even after correcting for nonresponse and trip underreporting, persons who did not travel on a given day comprise a significant portion of the population. Excluding zero-travel days leads to overstated trip rates because the denominator omits days when no trips were made (or people, households, who make no trips).

Correct trip-rate estimation requires including days with:

  • no recorded trips, and
  • positive day weight.

To do this, analysts must:

  1. Count weighted trips per person-day using the trip table.
  2. Join that summary back to the day table.
  3. Replace missing trip counts with zero for that person-day.
  4. Divide weighted trips by day weights when day weights are positive.

This ensures that persons who did not travel-and those who missed diary days-are represented correctly in the denominator.

Constructing a Person-Day Trip Rate Dataset

A typical workflow is illustrated below. The code begins by aggregating trips to the day level, joining that to the day table, and filling in zeros for days without travel. Once each person-day carries a trip count and a day weight, you can group by any person- or household-level variable to estimate trip rates:

Code
#------------------------------------------------------------
# 1. Compute weighted trips for each person-day
#    (sum of trip weights for all trips on that day)
#------------------------------------------------------------
weighted_trips <- hts$trip %>%
  group_by(day_id) %>%
  summarize(
    weighted_trips = sum(trip_weight),
    .groups = "drop"
  )

#------------------------------------------------------------
# 2. Join weighted trips back to the day table
#    (ensuring zero-trip days are explicitly represented)
#------------------------------------------------------------
day_trips <- hts$day %>%
  left_join(weighted_trips, by = "day_id") %>%
  mutate(
    # Replace NA with 0 for days with no trips
    weighted_trips = ifelse(is.na(weighted_trips), 0, weighted_trips)
  )

#------------------------------------------------------------
# 3. Compute weighted trips per population-day (trip rate)
#    wtd_trips_on_day = (expanded trips) / (expanded days)
#------------------------------------------------------------
day_trips <- day_trips %>%
  mutate(
    wtd_trips_on_day = ifelse(
      day_weight > 0,
      weighted_trips / day_weight,
      0 # fallback safe value (should not occur in analysis)
    )
  )

#------------------------------------------------------------
# 4. Inspect resulting structure
#------------------------------------------------------------
day_trips %>%
  select(
    household_id,
    person_id,
    day_weight,
    day_id,
    wtd_trips_on_day
  ) %>%
  glimpse()
## Rows: 10,868
## Columns: 5
## $ household_id     <int64> 25000006, 25000006, 25000071, 25000096, 25000096, 2…
## $ person_id        <int64> 2500000601, 2500000602, 2500007101, 2500009601, 250…
## $ day_weight       <dbl> 22.80259, 22.80259, 75.94698, 25.27773, 25.27773, 25.…
## $ day_id           <int64> 250000060101, 250000060201, 250000710101, 250000960…
## $ wtd_trips_on_day <dbl> 0.000000, 0.000000, 2.686327, 11.931042, 0.000000, 6.…

This creates a person-day file with the number of weighted trips made that day. From here, you can summarize trip rates by any variable attached to days, persons or households.

WarningCommon Pitfalls in Trip Rate Estimation

1. Using unweighted trip counts or day counts

This implicitly assumes all persons, all days, and all trips are equally likely to be observed. In practice, diary nonresponse and mode effects make this incorrect.

2. Using mean(num_trips) instead of weighted means

This computes the average number of observed trips across observed person-days-a very different quantity than the average trips per population day.

3. Weighting only the numerator (trip weights) or only the denominator (day weights)

Both sides must be weighted to reflect the survey’s two-stage expansion:

  1. day-level expansion, and
  2. trip-level expansion.

Using only one results in biased estimates.

4. Ignoring zero-travel days

Trip tables contain no records for zero-travel days. Trip rates that rely only on the trip table automatically exclude these days unless joined to the day table. This usually overstates trip rates.

5. Aggregating trips by person without normalizing for person-days

A person with one diary day and a person with five diary days cannot be treated as equivalent “units” without applying day weights.

The example below illustrates several common mistakes in trip rate estimation. Using a person-day-level dataset (day_trips), we compute trip rates by gender using four different methods:

Code
trip_rates_gender <- day_trips %>%
  left_join(
    hts$person %>% select(person_id, gender),
    by = "person_id"
  ) %>%
  group_by(gender) %>%
  summarize(
    # INCORRECT: simple mean of observed trips
    trip_rate_wrong_1 = mean(num_trips),

    # INCORRECT: weighting only the day side
    trip_rate_wrong_2 = weighted.mean(num_trips, w = day_weight),

    # INCORRECT: weighting only the trip side
    trip_rate_wrong_3 = mean(wtd_trips_on_day),

    # CORRECT: weighted trips / weighted days
    trip_rate_correct =
      weighted.mean(wtd_trips_on_day, w = day_weight)
  )

Also incorrect would be calculating a trip rate from the trip table:

Code
trip_rates_gender_wrong_4 <- hts$trip %>%
  group_by(person_id, day_id) %>%
  summarize(
    trip_count_unwtd = n(),
    trip_count_wtd = sum(trip_weight, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  left_join(
    hts$person %>% select(person_id, gender),
    by = "person_id"
  ) %>%
  group_by(gender) %>%
  summarize(
    trip_rate_wrong_4 = mean(trip_count_unwtd),
    trip_rate_wrong_5 = mean(trip_count_wtd)
  )

Interpretation:

  • trip_rate_wrong_1 ignores all weighting.
  • trip_rate_wrong_2 adjusts for the person-day universe but not for trip expansion.
  • trip_rate_wrong_3 counts weighted trips but ignores how many population-days those trips represent.
  • trip_rate_wrong_4 and trip_rate_wrong_5 both ignore zero-travel days entirely by relying only on the trip table.
  • trip_rate_correct is the only estimator that uses both components of the weighting design.

Using srvyr to Calculate Trip Rates With Confidence Intervals

The {srvyr} package provides a convenient way to calculate trip rates along with standard errors and confidence intervals that account for the complex survey design. Below is an example of how to set up a survey design object and compute average trip rates with confidence intervals.

Code
day_trips %>%
  left_join(
    hts$hh %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  as_survey_design(
    ids = household_id,
    strata = sample_segment,
    weights = day_weight
  ) %>%
  summarize(
    trip_rate = survey_mean(
      wtd_trips_on_day,
      vartype = "ci"
    )
  )
## # A tibble: 1 × 3
##   trip_rate trip_rate_low trip_rate_upp
##       <dbl>         <dbl>         <dbl>
## 1      4.04          3.84          4.24

Person-Level Trip Rates by Household, Person or Day-Level Variables

NoteWhere We Are Starting From

The examples below start with the person-day-level trip counts (day_trips)created in Section 8.9.1.4. From there, we define survey design objects and compute trip rates disaggregated by variables of interest.

Calculating trip rates by person- or household-level variables (e.g., county, disability status, age) requires joining those variables to the person-day dataset. Day-level variables (e.g., telework_time) can be used directly since they are already present in the person-day file.

Example 1: Trip Rates by county

In this example, we compute trip rates by household county. We join both home_county and sample_segment from the household table to the person-day trip counts, define the survey design using the sample-design strata, and then calculate weighted trip rates with confidence intervals for each county.

Code
#------------------------------------------------------------
# 1. Add county and sample-segment information to each person-day
#------------------------------------------------------------
trip_rates_by_strata <- day_trips %>%
  left_join(
    hts$hh %>% select(household_id, home_county, sample_segment),
    by = "household_id"
  ) %>%
  # Keep valid weighted person-days
  filter(day_weight > 0)

#------------------------------------------------------------
# 2. Define the survey design on ALL valid person-days
#------------------------------------------------------------
trip_rates_by_strata <- trip_rates_by_strata %>%
  as_survey_design(
    ids     = household_id, # PSU: household
    strata  = sample_segment, # sample-design strata
    weights = day_weight
  )

#------------------------------------------------------------
# 3. Estimate weighted trip rates by county
#------------------------------------------------------------
trip_rates_by_strata <- trip_rates_by_strata %>%
  group_by(home_county) %>%
  summarize(
    trip_rate = survey_mean(
      wtd_trips_on_day,
      vartype = c("ci", "se") # include CIs for reporting & reliability checks
    ),
    .groups = "drop"
  )
Code
trip_rates_by_strata %>%
  mutate(rse = trip_rate_se / trip_rate) %>%
  gt() %>%
  fmt_number(
    columns = c(trip_rate, trip_rate_low, trip_rate_upp),
    decimals = 2
  ) %>%
  cols_label(
    home_county = "County",
    trip_rate = "Weighted Trip Rate",
    trip_rate_low = "95% confidence interval (CI) lower bound",
    trip_rate_upp = "95% confidence interval (CI) upper bound"
  ) %>%
  tab_header(
    title    = "Weighted Trip Rates by county",
    subtitle = "With 95% Confidence Intervals"
  )
Weighted Trip Rates by county
With 95% Confidence Intervals
County Weighted Trip Rate 95% confidence interval (CI) lower bound 95% confidence interval (CI) upper bound trip_rate_se rse
King County 4.01 3.79 4.24 0.1139012 0.02838973
Kitsap County 3.78 3.15 4.41 0.3213877 0.08496289
Pierce County 4.25 3.79 4.70 0.2306997 0.05434286
Snohomish County 3.98 3.35 4.60 0.3167407 0.07967371
Table 45: Weighted Trip Rates by County with 95% Confidence Intervals

Example 2: Trip Rates by Disability Status and Age

Next, we compute trip rates by person disability status and age group. This requires joining person-level variables to the person-day file. We also create labeled and binned age categories for easier interpretation and larger sample sizes. Because we’re working with a subpopulation (persons with and without disabilities), we filter inside the survey design object to ensure correct variance estimation. We calculate RSE to assess reliability, important here because disability subgroups may have small sample sizes.

Code
#------------------------------------------------------------
# 2. Prepare person-day file:
#    - join county, sample segment, disability, age
#    - add binned age categories
#    - keep only valid (nonzero) day weights
#------------------------------------------------------------
trip_rates_by_disability_age <- day_trips %>%
  # Add household county and sample-design strata
  left_join(
    hts$hh %>% select(household_id, home_county, sample_segment),
    by = "household_id"
  ) %>%
  # Add disability + age from person table
  left_join(
    hts$person %>%
      select(person_id, age, disability_person),
    by = "person_id"
  ) %>%
  mutate(age = age) %>%
  mutate(disability = disability_person) %>%
  # Create broader age bins directly from the delivered age field
  mutate(
    age_binned = case_when(
      suppressWarnings(as.numeric(age)) < 18 ~ "<18",
      suppressWarnings(as.numeric(age)) >= 18 & suppressWarnings(as.numeric(age)) < 65 ~ "18-64",
      suppressWarnings(as.numeric(age)) >= 65 ~ "65+",
      TRUE ~ "Missing"
    )
  ) %>%
  # Keep weighted (valid) person-days
  filter(day_weight > 0)

#------------------------------------------------------------
# 3. Define the survey design on the full person-day file
#------------------------------------------------------------
trip_rates_by_disability_age <- trip_rates_by_disability_age %>%
  as_survey_design(
    ids     = household_id, # PSU: household
    strata  = sample_segment, # sample-design strata
    weights = day_weight
  )

#------------------------------------------------------------
# 4. Filter subpopulation of interest INSIDE the design object
#------------------------------------------------------------
trip_rates_by_disability_age <- trip_rates_by_disability_age %>%
  filter(disability %in% c("Yes", "No"))

#------------------------------------------------------------
# 5. Compute weighted trip rates and supporting counts
#------------------------------------------------------------
trip_rates_by_disability_age <- trip_rates_by_disability_age %>%
  group_by(disability, age_binned) %>%
  summarize(
    n_days    = n(),
    n_persons = n_distinct(person_id),
    n_hh      = n_distinct(household_id),
    trip_rate = survey_mean(wtd_trips_on_day, vartype = "se"),
    .groups   = "drop"
  ) %>%
  # Add RSE for reliability checking
  mutate(
    rse = trip_rate_se / trip_rate
  )

Trip Rates by Day-Level Telecommute Time

Here, we compute trip rates by day-level telework_time categories. We bin the raw telework-time responses into meaningful groups for interpretation. Because telework_time is already at the day level, we can use it directly from the person-day file without additional joins.

Code
#------------------------------------------------------------
# 1. Join day-level data to sample-design strata
#------------------------------------------------------------
trip_rates_by_telework_time <-
  day_trips %>%
  left_join(
    hts$hh %>% select(household_id, sample_segment),
    by = "household_id"
  ) %>%
  #----------------------------------------------------------
  # 2. Create a binned version of telework_time
  #    (converts delivered telework categories into readable groups)
  #----------------------------------------------------------
  mutate(
    telework_time_binned = case_when(
      telework_time == "0 minutes" ~ "No Telecommute",
      telework_time %in% c("30 minutes", "1 hour", "1 hour 30 minutes", "2 hours") ~ "Up to 2 hours",
      telework_time %in% c("2 hours 30 minutes", "3 hours", "3 hours 30 minutes", "4 hours") ~ "2 to 4 hours",
      telework_time %in% c("4 hours 30 minutes", "5 hours", "5 hours 30 minutes", "6 hours") ~ "4 to 6 hours",
      telework_time %in% c("6 hours 30 minutes", "7 hours", "7 hours 30 minutes", "8 hours", "8 hours 30 minutes", "9 hours", "9 hours 30 minutes", "10+ hours") ~ "More than 6 hours",
      TRUE ~ "Missing"
    )
  ) %>%
  #----------------------------------------------------------
  # 3. Keep only population-representative person-days
  #    (day_weight = 0 indicates non-representative rows)
  #----------------------------------------------------------
  filter(day_weight > 0) %>%
  #----------------------------------------------------------
  # 4. Define the survey design using day-level weights
  #----------------------------------------------------------
  as_survey_design(
    ids = household_id, # PSU
    strata = sample_segment, # sample-design strata
    weights = day_weight
  ) %>%
  #----------------------------------------------------------
  # 5. Estimate trip rate by telecommute-time category
  #----------------------------------------------------------
  group_by(telework_time_binned) %>%
  summarize(
    trip_rate = survey_mean(
      wtd_trips_on_day,
      vartype = "ci" # return standard error + CI
    ),
    .groups = "drop"
  )

Person-Level Trip Rates by Trip-Level Variables

Most trip-rate analysis focuses on person-, day-, or household-level characteristics-such as gender, age, disability status, employment, or household income. In those cases, each row in the person-day table represents one unit in the denominator, and weighted trip rates are calculated as weighted trips / weighted days. The structure is straightforward because the variables describing the subgroups (e.g., gender, county, telework_time) come from the same or a higher level of aggregation than the trip rate itself.

When summarizing trip rates by trip-level variables, such as mode, purpose, or access/egress type, the data structure changes. Trip attributes live in the trip table, not the day table, which means:

  • Each person-day may contribute trips in multiple categories (e.g., a person may walk, take transit, and drive on the same day).
  • The denominator must still be the number of person-days, not the number of trips or trip-days.
  • Categories should be defined for all person-days-including those with zero trips of a given type-to ensure that trip rates remain interpretable as “average number of X-type trips per person per day.”

The workflow therefore differs slightly from person- or day-level grouping:

  1. Aggregate weighted trips per day - trip category (e.g., per day - mode).
  2. Construct a complete grid of all person-days - all categories, explicitly including days with zero trips of that type.
  3. Attach day weights, compute weighted trips per weighted day, and then
  4. Summarize to obtain the mean trip rate for each category.

This ensures that a category such as “transit trips” is interpreted as:

The average number of transit trips made per person on a typical weekday, including all person-days where no transit trip occurred.

The method is consistent with the broader guidance in this handbook: trip rates should always incorporate both weighted trips and weighted days, regardless of whether the grouping variable comes from the day table, the person table, or the trip table.

Example: Trip Rates by Mode

Code
#------------------------------------------------------------

# 1. Count weighted trips per person-day by mode

#------------------------------------------------------------
weighted_trips_by_day_mode <- hts$trip %>%
  mutate(
    mode_label = dplyr::if_else(is.na(mode_class_5), "Missing Response", mode_class_5)
  ) %>%
  group_by(day_id, mode_label) %>%
  summarise(
    weighted_trips = sum(trip_weight, na.rm = TRUE),
    .groups = "drop"
  )

#------------------------------------------------------------

# 3. Create the full day - mode grid and attach day weights

#------------------------------------------------------------
day_mode_trips <- hts$day %>%
  select(day_id, day_weight) %>%
  # keep only days that are actually in the weighted universe
  filter(day_weight > 0) %>%
  crossing(
    mode_label = sort(unique(dplyr::coalesce(
      hts$trip$mode_class_5,
      hts$trip$mode_class,
      hts$trip$mode_1,
      "Missing Response"
    )))
  ) %>%
  left_join(
    weighted_trips_by_day_mode,
    by = c("day_id", "mode_label")
  ) %>%
  mutate(
    # Fill in days with no trips of this mode
    weighted_trips = if_else(is.na(weighted_trips), 0, weighted_trips),
    # Compute weighted trips per person-day for this mode
    wtd_trips_on_day_mode = weighted_trips / day_weight
  )


#------------------------------------------------------------

# 4. Summarize person-level trip rates by mode

#------------------------------------------------------------
trip_rates_by_mode <- day_mode_trips %>%
  group_by(mode_label) %>%
  summarise(
    trip_rate = mean(wtd_trips_on_day_mode, na.rm = TRUE),
    .groups = "drop"
  )

trip_rates_by_mode
## # A tibble: 6 × 2
##   mode_label         trip_rate
##   <chr>                  <dbl>
## 1 Bike/Micromobility  0.0686  
## 2 Drive               3.03    
## 3 Missing Response    0.000131
## 4 Other               0.114   
## 5 Transit             0.184   
## 6 Walk                0.503

Household-Level Trip Rates

At the household level, trip rates can be calculated by dividing the total weighted trips by the total household weight:

Code
sum(hts$trip$trip_weight) /
  sum(hts$hh$hh_weight)
## [1] NA

This approach works because each household’s weight reflects its representation in the population, and all trips made by household members are included in the numerator.

To calculate trip rates by household characteristics (e.g., household size), we can aggregate trips and households by that characteristic and then compute the trip rate as the ratio of weighted trips to weighted households. For example, to compute trip rates by household size:

Code
trips_by_hh_size <- hts$trip %>%
  left_join(
    hts$hh %>% select(household_id, hhsize),
    by = "household_id"
  ) %>%
  group_by(hhsize) %>%
  summarise(
    weighted_trips_hh = sum(trip_weight, na.rm = TRUE),
    .groups = "drop"
  )

hhs_by_hh_size <- hts$hh %>%
  group_by(hhsize) %>%
  summarize(
    weighted_hhs = sum(hh_weight, na.rm = TRUE),
    .groups = "drop"
  )

hh_trip_rates_by_size <- trips_by_hh_size %>%
  left_join(
    hhs_by_hh_size,
    by = "hhsize"
  ) %>%
  mutate(
    trip_rate = weighted_trips_hh / weighted_hhs
  )

Table 46 shows the resulting trip rates by household size:

Code
hh_trip_rates_by_size %>%
  gt() %>%
  fmt_number(
    columns = trip_rate,
    decimals = 2
  ) %>%
  fmt_number(
    columns = c(weighted_trips_hh, weighted_hhs),
    decimals = 0
  ) %>%
  cols_label(
    hhsize            = "Household Size",
    weighted_trips_hh = "Weighted Trips",
    weighted_hhs      = "Weighted Households",
    trip_rate         = "Trip Rate (Trips per Household)"
  )
Household Size Weighted Trips Weighted Households Trip Rate (Trips per Household)
1 person 2,073,378 480,662 4.31
2 people 4,593,737 608,417 7.55
3 people 3,637,266 298,266 12.19
4 people 4,194,167 246,375 17.02
5 people 1,508,928 69,948 21.57
6 people 628,864 25,543 24.62
7 people 449,360 21,066 21.33
8 people 14,696 314 46.87
Table 46: Average Household Trip Rates by Household Size.

These simple calculations work as long as:

  • CIs or SEs are not required
  • Disaggregation is only by household-level variables (or variables that can be rolled up to the household level).

If either condition is not met, we use a modified approach that constructs a household-day-level dataset, then uses {srvyr} to calculate trip rates from a survey design object.

First, we create a household-day-level dataset that includes day weights adjusted for the number of completed weekdays per household. This ensures that each household-day represents its share of the household’s total weight appropriately:

Code
hh_weighted_weekday_counts <-
  hts$day %>%
  distinct(household_id, travel_date, .keep_all = TRUE) %>%
  group_by(household_id) %>%
  summarise(
    num_days_complete_weighted_weekday = sum(
      summary_complete == "Yes" &
        travel_dow %in% c("Monday", "Tuesday", "Wednesday", "Thursday"),
      na.rm = TRUE
    ),
    .groups = "drop"
  )

hh_day <-
  hts$day %>%
  select(household_id, day_id, travel_date, day_weight, travel_dow, summary_complete) %>%
  distinct(household_id, travel_date, .keep_all = TRUE) %>%
  left_join(
    hh_weighted_weekday_counts,
    by = "household_id"
  ) %>%
  left_join(
    hts$hh %>%
      select(household_id, hh_weight),
    by = "household_id"
  ) %>%
  mutate(
    hh_day_weight = dplyr::if_else(
      summary_complete == "Yes" & travel_dow %in% c("Monday", "Tuesday", "Wednesday", "Thursday"),
      hh_weight / num_days_complete_weighted_weekday,
      0,
      # Some rows have missing completeness or weekday info; treat them as zero-weight days
      missing = 0
    )
  )

glimpse(hh_day)
## Rows: 5,730
## Columns: 9
## $ household_id                       <int64> 25000006, 25000071, 25000096, 250…
## $ day_id                             <int64> 250000060101, 250000710101, 25000…
## $ travel_date                        <dttm> 2025-03-24, 2025-03-17, 2025-03-06…
## $ day_weight                         <dbl> 22.80259, 75.94698, 25.27773, 22.72…
## $ travel_dow                         <chr> "Monday", "Monday", "Thursday", "Tu…
## $ summary_complete                   <chr> "Yes", "Yes", "Yes", "Yes", "Yes", …
## $ num_days_complete_weighted_weekday <int> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,…
## $ hh_weight                          <dbl> 22.80259, 75.94698, 25.27773, 22.72…
## $ hh_day_weight                      <dbl> 22.80259, 75.94698, 25.27773, 22.72…

The process above essentially spreads each household’s weight evenly across its completed weighted weekdays (Monday-Thursday in this delivery), assigning zero weight to incomplete or unweighted days.

The check below ensures that the total household-day weights equal the total household weights:

Code
stopifnot(
  all.equal(
    sum(hh_day$hh_day_weight, na.rm = TRUE),
    sum(hts$hh$hh_weight, na.rm = TRUE),
    tolerance = 1e-6,
    check.attributes = FALSE
  ) == TRUE
)

Next, we aggregate trips to the household-day level, summing trip weights for all trips made by household members on each day:

Code
hh_day_trips <- hts$trip %>%
  group_by(household_id, day_id) %>%
  summarise(
    weighted_trips_hh_day = sum(trip_weight, na.rm = TRUE),
    .groups = "drop"
  )


glimpse(hh_day_trips)
## Rows: 8,966
## Columns: 3
## $ household_id          <int64> 25000071, 25000096, 25000096, 25000152, 250002…
## $ day_id                <int64> 250000710101, 250000960101, 250000960301, 2500…
## $ weighted_trips_hh_day <dbl> 204.01840, 301.58973, 169.76064, 61.05516, 315.5…

We then join the household-day trip counts back to the household-day table, filling in zeros for days without trips:

Code
hh_day_trips <- hh_day %>%
  left_join(
    hh_day_trips,
    by = c("household_id", "day_id")
  ) %>%
  mutate(
    # Fill in days with no trips
    weighted_trips_hh_day = if_else(
      is.na(weighted_trips_hh_day),
      0,
      weighted_trips_hh_day
    )
  )

glimpse(hh_day_trips)
## Rows: 5,730
## Columns: 10
## $ household_id                       <int64> 25000006, 25000071, 25000096, 250…
## $ day_id                             <int64> 250000060101, 250000710101, 25000…
## $ travel_date                        <dttm> 2025-03-24, 2025-03-17, 2025-03-06…
## $ day_weight                         <dbl> 22.80259, 75.94698, 25.27773, 22.72…
## $ travel_dow                         <chr> "Monday", "Monday", "Thursday", "Tu…
## $ summary_complete                   <chr> "Yes", "Yes", "Yes", "Yes", "Yes", …
## $ num_days_complete_weighted_weekday <int> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,…
## $ hh_weight                          <dbl> 22.80259, 75.94698, 25.27773, 22.72…
## $ hh_day_weight                      <dbl> 22.80259, 75.94698, 25.27773, 22.72…
## $ weighted_trips_hh_day              <dbl> 0.00000, 204.01840, 301.58973, 0.00…

The check below confirms that the household-day calculation is internally consistent when recomputed from the same weighted household-day totals:

Code
hhday_trip_rate <-
  sum(hh_day_trips$weighted_trips_hh_day) /
    sum(hh_day_trips$hh_day_weight)


direct_trip_rate <-
  sum(hh_day_trips$weighted_trips_hh_day) /
    sum(hh_day$hh_day_weight)


all.equal(hhday_trip_rate, direct_trip_rate)
## [1] TRUE

Finally, we can use {srvyr} to compute average household trip rates by household size, along with standard errors and RSEs for reliability assessment:

Code
hh_day_trips_by_size <- hh_day_trips %>%
  left_join(hts$hh %>% select(household_id, hhsize, sample_segment), by = "household_id") %>%
  filter(hh_day_weight > 0) %>%
  as_survey_design(
    ids = household_id,
    weights = hh_day_weight,
    strata = sample_segment
  ) %>%
  group_by(hhsize) %>%
  summarize(
    n_hh_days = n(),
    n_hhs = n_distinct(household_id),
    trip_rate = survey_mean(
      weighted_trips_hh_day / hh_day_weight,
      vartype = "se"
    ),
    .groups = "drop"
  ) %>%
  mutate(
    rse = trip_rate_se / trip_rate
  )

Table 47 presents the resulting trip rates by household size, along with standard errors and RSEs for reliability assessment. Cells with small sample sizes or high uncertainty are highlighted for reader awareness.

Code
hh_day_trips_by_size %>%
  mutate(
    # Flag cases where SE / RSE are not meaningful
    se_unreliable = (n_hhs < 2 | trip_rate_se == 0 | is.na(trip_rate_se)),

    # Build Trip Rate (SE) display string
    trip_rate_display = if_else(
      se_unreliable,
      sprintf("%.2f (--)", trip_rate),
      sprintf("%.2f (%.2f)", trip_rate, trip_rate_se)
    ),

    # Suppress RSE when SE is unreliable
    rse_for_display = if_else(se_unreliable, NA_real_, rse)
  ) %>%
  select(
    hhsize,
    n_hh_days,
    n_hhs,
    trip_rate_display,
    rse_for_display
  ) %>%
  gt() %>%
  # Format counts
  fmt_number(
    columns = c(n_hh_days, n_hhs),
    decimals = 0
  ) %>%
  # Format RSE as percent where available
  fmt_percent(
    columns = rse_for_display,
    decimals = 1
  ) %>%
  # Replace missing RSE with "--"
  sub_missing(
    columns = rse_for_display,
    missing_text = "--"
  ) %>%
  # Column labels
  cols_label(
    trip_rate_display = "Trip Rate (standard error, SE)",
    rse_for_display   = "Relative standard error (RSE, %)"
  ) %>%
  # RED: very small samples or very high uncertainty
  tab_style(
    style = list(
      cell_text(color = "red")
    ),
    locations = cells_body(
      rows = (n_hhs < 10) | (rse_for_display > 0.50),
      columns = c(n_hhs, rse_for_display)
    )
  ) %>%
  # ORANGE: small samples or high uncertainty
  tab_style(
    style = list(
      cell_text(color = "orange")
    ),
    locations = cells_body(
      rows = (n_hhs >= 10 & n_hhs < 30) |
        (rse_for_display > 0.30 & rse_for_display <= 0.50),
      columns = c(n_hhs, rse_for_display)
    )
  ) %>%
  cols_label(
    hhsize     = "Household Size",
    n_hh_days  = "Unweighted Household-Days",
    n_hhs      = "Unweighted Households"
  ) %>%
  # Legend / explanation for readers
  tab_source_note(
    source_note = md(
      "**Reliability flags:**
      Values in <span style='color: orange;'>*orange*</span> indicate small samples (10-29 households) or high uncertainty (relative standard error, RSE, of 30-50%).
      Values in <span style='color: red;'>*red*</span> indicate very small samples (&lt; 10 households) or very high uncertainty (RSE &gt; 50%)."
    )
  )
Household Size Unweighted Household-Days Unweighted Households Trip Rate (standard error, SE) Relative standard error (RSE, %)
1 person 1,721 989 4.28 (0.16) 3.8%
2 people 1,624 1,179 4.34 (0.17) 3.9%
3 people 435 318 5.30 (0.39) 7.3%
4 people 326 205 5.75 (0.40) 6.9%
5 people 86 59 6.80 (1.27) 18.7%
6 people 17 14 6.74 (0.97) 14.4%
7 people 6 6 5.15 (1.97) 38.3%
8 people 2 2 13.87 (1.18) 8.5%
Reliability flags: Values in orange indicate small samples (10-29 households) or high uncertainty (relative standard error, RSE, of 30-50%). Values in red indicate very small samples (< 10 households) or very high uncertainty (RSE > 50%).
Table 47

The same approach can be used to compute household-level trip rates by any number of person-, day-, or household-level variables, while ensuring correct weighting and variance estimation. To compute trip rates disaggregated by trip-level variables (e.g., mode, purpose), follow the guidance in Section 8.9.3 to construct a complete household-day - trip-category dataset before applying the survey design and summarization steps.

8.10 Analysis of Person-Miles Traveled (PMT) and Vehicle-Miles Traveled (VMT)

Analysis of person-miles and vehicle-miles traveled proceeds similarly to the analysis of trip rates, with some additional considerations for occupancy and non-survey passengers.

Because the unit of the trip table is person-trips, total person-miles traveled (PMT) can be calculated by summing the product of trip distance and trip weight across all trips (see Section 4.2.2 for how distance is computed and cleaned):

Code
total_pmt <- hts$trip %>%
  summarize(
    total_pmt = sum(distance_miles * trip_weight, na.rm = TRUE)
  ) %>%
  pull(total_pmt)

format(total_pmt, big.mark = ",", scientific = FALSE)
## [1] "138,605,010"

Total vehicle-miles traveled (VMT) requires accounting for vehicle occupancy. To calculate VMT, we first divide distance traveled by the number of travelers in the vehicle. We then sum the product of trip distance and vehicle trip weight across all trips:

Code
total_vmt <- hts$trip %>%
  filter(mode_class_5 == "Drive") %>%
  mutate(
    total_occupancy = case_when(
      travelers_total == "1 traveler" ~ 1,
      travelers_total == "2 travelers" ~ 2,
      travelers_total == "3 travelers" ~ 3,
      travelers_total == "4 travelers" ~ 4,
      travelers_total == "5+ travelers" ~ 5,
      TRUE ~ NA_real_
    ),
    vmt = if_else(
      total_occupancy > 0,
      distance_miles / total_occupancy,
      NA_real_
    )
  ) %>%
  summarize(
    total_vmt = sum(vmt * trip_weight, na.rm = TRUE)
  ) %>%
  pull(total_vmt)

format(total_vmt, big.mark = ",", scientific = FALSE)
## [1] "74,811,233"

Disaggregating PMT and VMT by Population Subgroups

To disaggregate PMT or VMT by population subgroups (e.g., household residence type from hts$hh$res_type), we follow a similar approach to trip rates, constructing a day-level dataset that includes weighted PMT or VMT per person-day.

We start by calculating total weighted VMT per day:

Code
day_trip_vmt <- hts$trip %>%
  filter(mode_class_5 == "Drive") %>%
  mutate(
    total_occupancy = case_when(
      travelers_total == "1 traveler" ~ 1,
      travelers_total == "2 travelers" ~ 2,
      travelers_total == "3 travelers" ~ 3,
      travelers_total == "4 travelers" ~ 4,
      travelers_total == "5+ travelers" ~ 5,
      TRUE ~ NA_real_
    ),
    vmt = if_else(
      total_occupancy > 0,
      distance_miles / total_occupancy,
      NA_real_
    )
  ) %>%
  group_by(day_id) %>%
  summarise(
    total_wtd_vmt_on_day = sum(vmt * trip_weight, na.rm = TRUE),
    .groups = "drop"
  )

glimpse(day_trip_vmt)
## Rows: 7,459
## Columns: 2
## $ day_id               <int64> 250000710101, 250000960101, 250000960301, 25000…
## $ total_wtd_vmt_on_day <dbl> 829.655911, 473.682088, 249.787611, 5.501012, 148…

Next, we join this to the person-day table to add in zeros for days with no driving trips, then compute weighted VMT per person-day by dividing total weighted VMT on the day by the day weight:

Code
# Join to person-day table to add in zeros for days with no drivng trips

day_trip_vmt <- hts$day %>%
  select(day_id, day_weight, household_id) %>%
  left_join(
    day_trip_vmt,
    by = "day_id"
  ) %>%
  mutate(
    total_wtd_vmt_on_day = if_else(is.na(total_wtd_vmt_on_day), 0, total_wtd_vmt_on_day),
    wtd_vmt_per_day = total_wtd_vmt_on_day / day_weight
  )

glimpse(day_trip_vmt)
## Rows: 10,868
## Columns: 5
## $ day_id               <int64> 250000060101, 250000060201, 250000710101, 25000…
## $ day_weight           <dbl> 22.80259, 22.80259, 75.94698, 25.27773, 25.27773,…
## $ household_id         <int64> 25000006, 25000006, 25000071, 25000096, 2500009…
## $ total_wtd_vmt_on_day <dbl> 0.000000, 0.000000, 829.655911, 473.682088, 0.000…
## $ wtd_vmt_per_day      <dbl> 0.0000000, 0.0000000, 10.9241462, 18.7391033, 0.0…

Now we can disaggregate average VMT per person by household residence type.

We first create a dataset that joins the day-level VMT data to household data to get the household res_type and survey strata.

Next, we use {srvyr} to compute average VMT per person by residence type, along with confidence intervals:

The table chunk below shows the resulting average VMT per person by residence type, along with confidence intervals.

Annualizing VMT/PMT from Survey Data

Annualizing PMT and VMT is possible, but it requires several assumptions because the HTS is designed to represent typical weekdays, not full-year travel. If producing annual estimates, apply these methods transparently and document assumptions clearly.

A practical workflow includes:

1. Compute average weekday VMT or PMT using weighted trip data.

Use day-level or trip-level weights to obtain a representative weekday estimate:

  • PMT: sum(distance * trip_weight) / sum(day_weight)
  • VMT: sum(vehicle-miles * trip_weight) / sum(day_weight)

2. Estimate a weekday-weekend adjustment factor.
Because weekend travel differs systematically from weekday travel, adjust using one of the following: - Traffic counts or continuous count stations (preferred).
- Survey data, restricted to respondents with complete rMove reporting across both weekdays and weekends.
- Model-based estimates, e.g., fitting a GLM predicting VMT as a function of day-of-week.

This produces a ratio such as:
VMT_weekend = VMT_weekday * adjustment_factor.

3. Annualize using the calendar distribution of weekdays/weekends.
A typical year includes 260 weekdays and 105 weekend days. Thus:

Code
Annual_VMT <- (Avg_Weekday_VMT * 260) +
  (Avg_Weekend_VMT * 105)

Annual_PMT <- (Avg_Weekday_PMT * 260) +
  (Avg_Weekend_PMT * 105)

If only a weekday estimate is available:

Code
Annual_VMT <- Avg_Weekday_VMT * (260 + weekend_adjustment * 105)

Worked Example: Annualizing PMT and VMT from Day-Weighted Data

The code below demonstrates a simple approach to annualizing PMT and VMT using day-weighted survey data.

Code
#------------------------------------------------------------
# Step 1: Compute trip-level PMT and VMT
#   - PMT: distance_miles per person-trip
#   - VMT: distance_miles divided by total occupancy
#------------------------------------------------------------
trip_pmt_vmt <- hts$trip %>%
  mutate(
    pmt = distance_miles,
    total_occupancy = case_when(
      travelers_total == "1 traveler" ~ 1,
      travelers_total == "2 travelers" ~ 2,
      travelers_total == "3 travelers" ~ 3,
      travelers_total == "4 travelers" ~ 4,
      travelers_total == "5+ travelers" ~ 5,
      TRUE ~ NA_real_
    ),
    vmt = if_else(
      total_occupancy > 0,
      distance_miles / total_occupancy,
      NA_real_
    )
  )

#------------------------------------------------------------
# Step 2: Aggregate to person-day level and join weekday weights
#   Each day_id gets total PMT and VMT, then is linked to day_weight.
#   day_weight > 0 indicates an in-sample *weekday*.
#------------------------------------------------------------
day_pmt_vmt <- trip_pmt_vmt %>%
  group_by(day_id) %>%
  summarise(
    pmt_weighted = sum(pmt * trip_weight, na.rm = TRUE),
    vmt_weighted = sum(vmt * trip_weight, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  right_join(
    hts$day %>% select(day_id, day_weight),
    by = "day_id"
  ) %>%
  mutate(
    # Fill in days with no travel
    pmt_weighted = if_else(is.na(pmt_weighted), 0, pmt_weighted),
    vmt_weighted = if_else(is.na(vmt_weighted), 0, vmt_weighted)
  )

#------------------------------------------------------------
# Step 3: Design-based average *weekday* PMT and VMT
#   Use only rows with day_weight > 0 (i.e., weighted weekdays).
#------------------------------------------------------------
avg_weekday_pmt_vmt <- day_pmt_vmt %>%
  filter(day_weight > 0) %>%
  summarise(
    avg_weekday_pmt = sum(pmt_weighted) / sum(day_weight),
    avg_weekday_vmt = sum(vmt_weighted) / sum(day_weight),
    .groups = "drop"
  )

avg_weekday_pmt <- avg_weekday_pmt_vmt$avg_weekday_pmt
avg_weekday_vmt <- avg_weekday_pmt_vmt$avg_weekday_vmt

#------------------------------------------------------------
# Step 4: Specify weekend : weekday ratios (from external info)
#   Set these based on traffic counts, modeling, or other evidence.
#   Example placeholders:
#     weekend_pmt is 80% of weekday
#     weekend_vmt is 90% of weekday
#------------------------------------------------------------
weekend_factor_pmt <- 0.80 # replace with study-specific ratio
weekend_factor_vmt <- 0.90 # replace with study-specific ratio

avg_weekend_pmt <- avg_weekday_pmt * weekend_factor_pmt
avg_weekend_vmt <- avg_weekday_vmt * weekend_factor_vmt

#------------------------------------------------------------
# Step 5: Annualize using assumed weekday/weekend counts
#   Example: 260 weekdays and 105 weekend days in a year.
#------------------------------------------------------------
num_weekdays <- 260
num_weekends <- 105

avg_annual_pmt <- avg_weekday_pmt * num_weekdays +
  avg_weekend_pmt * num_weekends

avg_annual_vmt <- avg_weekday_vmt * num_weekdays +
  avg_weekend_vmt * num_weekends

total_annual_pmt <- avg_annual_pmt * sum(hts$person$person_weight)
total_annual_vmt <- avg_annual_vmt * sum(hts$person$person_weight)

annual_vmt_table <-
  data.frame(
    metric = c(
      "Avg Weekday PMT", "Avg Weekend PMT (assumed)",
      "Avg Annual PMT (assumed weekends)", "Total Annual PMT (assumed weekends)",
      "Avg Weekday VMT", "Avg Weekend VMT (assumed)",
      "Avg Annual VMT (assumed weekends)", "Total Annual VMT (assumed weekends)"
    ),
    miles_value = c(
      avg_weekday_pmt,
      avg_weekend_pmt,
      avg_annual_pmt,
      total_annual_pmt,
      avg_weekday_vmt,
      avg_weekend_vmt,
      avg_annual_vmt,
      total_annual_vmt
    )
  )

gt(annual_vmt_table) %>%
  fmt_number(
    columns = miles_value,
    decimals = 0
  ) %>%
  cols_label(
    metric     = "Metric",
    miles_value = "Miles"
  )
Metric Miles
Avg Weekday PMT 33
Avg Weekend PMT (assumed) 26
Avg Annual PMT (assumed weekends) 11,270
Total Annual PMT (assumed weekends) 47,680,123,323
Avg Weekday VMT 22
Avg Weekend VMT (assumed) 20
Avg Annual VMT (assumed weekends) 7,849
Total Annual VMT (assumed weekends) 33,206,101,966

This example is intentionally simple. For formal reporting, consider:

  • adding design-based standard errors (e.g., via srvyr), and
  • testing how sensitive annual totals are to different weekday/weekend assumptions.

8.11 From Description to Inference: Using Weighted Models

Simple weighted proportions, with accompanying standard errors or confidence intervals, are an excellent first tool for describing population patterns. However, there are many situations where weighted proportions alone are not sufficient for reliable inference. When key demographic groups are small, or when design effects are large in some weighting groups, estimates for those subgroups become unstable: sampling variance increases, effective sample size shrinks, and confidence intervals widen substantially. In these cases, analysts should use weighted multivariate models rather than relying solely on subgroup proportions. Modeling keeps the full sample intact, improves statistical precision, and allows analysts to estimate the unique contribution of each factor while holding others constant. This approach avoids the instability that arises from slicing the data into many small subpopulations and provides a principled way to test differences across groups through main effects, interactions, or model-based comparisons.

NoteUsing Survey Weights in Regression Models

Most analysts will work in R, Stata, SPSS, or SAS. Each platform provides dedicated tools for fitting regression models that correctly incorporate survey weights, clustering, and stratification. In R, the {survey} package supports design-corrected estimation through svydesign() and model fitting using functions such as svyglm(), with pseudo-likelihood methods available through {survey} and {srvyr}. In Stata, the svyset command defines the design, and survey-weighted models are fit by adding the svy: prefix to commands such as regress, logit, glm, and mixed. SPSS Complex Samples provides specialized procedures such as CSLOGISTIC and CSGLM that automatically incorporate design variables and produce linearized standard errors. In SAS, the SURVEY procedures such as PROC SURVEYREG, PROC SURVEYLOGISTIC, and PROC SURVEYMEANS support stratification, clustering, and weights, with options for domain estimation. For multilevel structures, these tools allow analysts to define PSUs such as households and incorporate stratification variables while estimating the effects of predictors across hierarchical units without manually subsetting the data.

Across all platforms, the key principle is the same: define the survey design once, then fit models using functions that respect the sampling structure to obtain valid, population-representative inferences.

Does Telework Reduce VMT?

To illustrate the use of weighted models for inference, we consider the question: Does teleworking reduce vehicle-miles traveled (VMT)? We use a linear regression model to estimate average VMT per person-day as a function of telework status, while controlling for household vehicles, presence of children, household income, age group, and neighborhood population density.

Before fitting the model, we enrich the household table with block-group population density keyed to home_bg_2020. The chunk below uses {tidycensus} to fetch ACS 2023 5-year population counts for block groups in King, Kitsap, Pierce, and Snohomish counties, computes land area from geometry, and immediately joins the result to the household table. Analysts running this chunk will need a Census API key and internet access.

Next, we build a person-day modeling dataset that includes daily VMT, telework status, household covariates, and the new density measure. The chunk below uses the delivered household and person fields directly, derives telework and children categories from their labeled values, and appends inline block-group population density.

Next, we create a survey design object using {srvyr} to account for the complex survey design, including day weights, household PSUs, and sample-segment strata.

Finally, we fit a survey-weighted linear regression model predicting daily VMT as a function of telework status, neighborhood population density, vehicles, kids, income, and age group. The density term enters as log1p(population_density) to reduce skew while preserving a straightforward public density measure in the modeling dataset.

Table 48 presents the results of the survey-weighted linear regression model of daily VMT.

Survey-weighted Linear Model of Daily VMT1
Outcome: Weighted Vehicle-Miles Traveled per Diary Day
Term Estimate Standard error t-value p-value CI 2.5% CI 97.5%
Intercept 39.45 11.07 3.56 <0.001 *** 17.73 61.17
Telecommuting
Partial telework (vs None) -0.70 3.20 -0.22 0.827 -6.98 5.58
Full telework (vs None) -6.71 2.15 -3.11 0.002 ** -10.93 -2.48
Built environment
Log population density -4.55 1.07 -4.26 <0.001 *** -6.65 -2.46
Vehicles
Vehicles: 1 16.02 2.95 5.43 <0.001 *** 10.23 21.80
Vehicles: 2 18.31 2.52 7.26 <0.001 *** 13.37 23.26
Vehicles: 3 24.83 3.78 6.57 <0.001 *** 17.41 32.24
Vehicles: 4+ 20.48 4.85 4.22 <0.001 *** 10.96 30.00
Household composition
Kids: 1 kid -2.71 3.09 -0.88 0.380 -8.78 3.35
Kids: 2+ kids -4.47 2.83 -1.58 0.114 -10.03 1.08
Household income
Income: $25,000-$49,999 4.07 3.76 1.08 0.279 -3.30 11.44
Income: $50,000-$74,999 12.23 5.44 2.25 0.025 * 1.56 22.90
Income: $75,000-$99,999 7.69 4.65 1.65 0.099 -1.44 16.81
Income: $100,000-$199,999 5.89 3.41 1.73 0.085 -0.80 12.57
Income: $200,000 or more 3.69 3.64 1.01 0.311 -3.45 10.83
Age group
Age: 35-54 3.06 2.53 1.21 0.227 -1.90 8.02
Age: 55+ 1.00 2.54 0.39 0.694 -3.98 5.98
1 Notes:

  • Estimates reflect Taylor-series linearized standard errors using day weights, household PSUs, and sample-segment strata.
  • population_density is measured in persons per square mile and enters the model as log1p(population_density).
  • The reference categories are No telework, 0 vehicles, 0 kids, the lowest observed household-income level, and Age 18-34.
  • p < 0.05 (), p < 0.01 (), p < 0.001 ()
Table 48: Survey-weighted Linear Model of Daily VMT

Wrapping Up: Why Use Weighted Models?

This example highlights why weighted models are essential for understanding travel behavior and why simple summary statistics or subgroup tabulations are often insufficient. Descriptive tables can show how much travel occurs across different households, but they cannot adjust for confounding variables or account for the structure of the survey sample. Weighted regression models, by contrast, allow us to estimate the independent effects of telework status, neighborhood density, vehicle ownership, household composition, income, and age while incorporating sampling weights, clustering, and stratification to produce valid population-level inferences and realistic measures of uncertainty.

This model-based approach becomes especially important when working with small or unevenly represented subgroups. For example, zero-vehicle households, transit-oriented households, or specific telework groups can appear unstable in simple weighted tabulations because their effective sample sizes are smaller than their raw counts suggest. In a multivariate model, the information loss is often far more modest because estimation draws strength from related predictors rather than relying on a single subgroup cell count alone. In other words, modeling stabilizes groups that would otherwise look noisy when examined in isolation.

Weighted models also help analysts avoid a common pitfall: subsetting data to compare groups directly. Because design effects in this survey can be substantial, splitting the data into small racial, income, or geographic categories can dramatically inflate variance and lead to misleading conclusions. Instead of slicing the dataset into ever-smaller groups, weighted models allow analysts to examine differences through interaction terms, contrast statements, or model-based predictions, all while respecting the survey design.

Taken together, these features make weighted regression a more reliable and generalizable way to study travel behavior. It allows analysts to control for multiple factors at once, correctly propagate design-based uncertainty, and draw conclusions that extend beyond the sample to the full regional population, something simple summaries cannot achieve on their own.