Spatial Regression Analysis of Frequent Mental Distress in Washington State
by: Nicole Vicenti
by: Nicole Vicenti
Frequent mental distress, defined as experiencing poor mental health for at least 14 days within a 30-day period, represents an important public health concern (CDC, 2025). Rates of mental distress vary across census tracts in Washington, suggesting that underlying behavioral and social factors—such as sleep patterns, depression, and cognitive disability—may contribute to these spatial differences. Understanding how these factors influence mental distress can help identify high-risk communities and support more targeted public health interventions.
Research Question
To what extent do cognitive disability, depression, and short sleep duration explain variation in frequent mental distress across Washington census tracts, and how do these relationships vary spatially?
Study Area
The study area for this project is the state of Washington, which has among the highest reported rates of mental health conditions in the United States (NAMI, 2023). In 2020, more than half of individuals with a mental illness in Washington did not receive treatment, and over 2.8 million residents lived in areas with a shortage of mental health professionals. Additionally, approximately 1 in 20 adults and 1 in 6 youth (ages 6–17) experience a mental disorder each year, highlighting the need for improved access to services. Census tracts are used as the unit of analysis to capture localized spatial variation in frequent mental distress across the state.
Dataset
This analysis uses the PLACES: Local Data for Better Health (2022) dataset, developed by the Centers for Disease Control and Prevention in collaboration with the Robert Wood Johnson Foundation. The dataset provides modeled estimates of health outcomes, behaviors, and social determinants at the census tract level across the United States, including variables such as frequent mental distress, smoking, sleep duration, depression, and social isolation. A subset of the national polygon feature layer was extracted for Washington and used for spatial regression analysis (Figure 1).
Metadata:
Name of Dataset: PLACES: Local Data for Better Health 2022
Publication/Update Date: February 5, 2026
Geometry Type: Polygon Feature Layer
Owner: Centers for Disease Control and Prevention (CDC)
Description: This dataset includes chronic disease and other health outcomes (in %) at the census tract level across the U.S.
Figure 1. Spatial distribution of frequent mental distress (crude prevalence, %) across census tracts in Washington. Darker shades represent higher prevalence, illustrating variation in mental distress across the study area.
The dataset was projected in NAD 1983 USFS R6 Albers (Meters) (WKID 9674), a projected coordinate system designed for the Pacific Northwest. This projection preserves area and distance relationships, making it appropriate for spatial analysis involving census tract polygons. Using meters ensures consistency in distance-based calculations required for regression models such as GWR.
The dependent variable in this study is frequent mental distress (crude prevalence %). An initial set of 12 explanatory variables was selected for exploratory analysis, including lack of health insurance, binge drinking, high blood pressure, current smoking, obesity, depression, fair or poor health, self-care disability, cognitive disability, mobility disability, social isolation, and population density. All variables are expressed as crude prevalence (%) unless otherwise noted.
Exploratory Data Analysis
Exploratory Data Analysis was conducted using correlation tables and charts to evaluate relationships between the dependent and explanatory variables. Variables were selected based on two criteria: a strong correlation with the dependent variable (R² ≥ 0.5) and low correlation among explanatory variables (R² < 0.5) to reduce multicollinearity. Selected variables were further evaluated using the Exploratory Regression tool to confirm statistical significance.
Generalized Linear Regression (GLR) Analysis
Generalized Linear Regression (GLR) is a global regression model that assumes relationships between variables are constant across the study area. A GLR model was developed using the selected explanatory variables with a Continuous (Gaussian) model type, appropriate for percentage data. Model performance was evaluated using multiple and adjusted R² values, while explanatory variables were assessed using robust probability values and variance inflation factors (VIF). Model significance was determined using the Joint F-statistic and Joint Wald statistic. The Koenker (BP) statistic was used to test for nonstationarity, and the Jarque-Bera statistic assessed residual normality.
Geographically Weighted Regression (GWR) Analysis
Geographically Weighted Regression (GWR) is a local regression model that accounts for spatial nonstationarity by allowing relationships to vary across the study area. The same explanatory variables and Continuous (Gaussian) model type were used. A neighborhood type of 'number of neighbors' with a golden search method was applied to optimize bandwidth selection. Model performance was evaluated using the Akaike Information Criterion (AICc), with lower values indicating a better model fit compared to the GLR results.
Nonstationarity was assessed using the Koenker (BP) statistic from the GLR model. A statistically significant result indicated that relationships between variables vary across space, justifying the use of GWR.
Assumptions
Several assumptions were made in this analysis. First, all data were assumed to be accurate and representative of underlying population characteristics. Second, explanatory variables were assumed to have meaningful relationships with the dependent variable. Thresholds of R² ≥ 0.5 (for dependent relationships) and R² < 0.5 (to limit multicollinearity) were used based on common practices in regression analysis.
For the GWR model, it was assumed that spatial relationships vary locally and can be effectively captured using neighboring features. The number of neighbors and golden search method were selected to balance model complexity and ensure stable, interpretable results.
Exploratory Data Analysis
Of the initial 12 explanatory variables, three met the selection criteria for regression analysis. Cognitive disability showed the strongest relationship with frequent mental distress (R² = 0.88), followed by depression (R² = 0.55) and short sleep duration (R² = 0.50). Although current smoking had a moderate correlation (R² = 0.57), it was excluded due to strong correlations with other explanatory variables, which could introduce multicollinearity. The Exploratory Regression results also indicated high statistical significance for all three variables, each demonstrating high statistical significance across model iterations. These three variables were therefore selected for use in the GLR and GWR models.
Generalized Linear Regression (GLR) Analysis
The GLR results indicate that the selected explanatory variables strongly predict frequent mental distress, with a multiple R² and adjusted R² of 0.94. All variables were statistically significant, with low VIF values (< 2.5), indicating minimal multicollinearity. The Joint F-statistic and Joint Wald statistic confirmed overall model significance. A significant Koenker (BP) statistic indicated nonstationarity, suggesting that relationships vary across space. The Jarque-Bera statistic indicated non-normal residuals, suggesting potential model bias. Spatial autocorrelation of residuals further indicated clustering, suggesting that additional spatial processes may not be fully captured by the model. The model’s AICc value was 4,115.35.
Geographically Weighted Regression (GWR) Analysis
The GWR model improved model performance, with an R² and adjusted R² of 0.99 and a lower AICc value (1,238.57) compared to the GLR model, indicating a better fit. The cognitive disability map (Figure 2) shows the strongest and most consistent relationships in western Washington, particularly around the Seattle metropolitan area. However, localized clusters of strong relationships are also observed in eastern Washington, especially near Spokane, as well as in parts of southern Washington, indicating regional variability. In contrast, the depression map (Figure 3) shows stronger associations in eastern and southern Washington and weaker relationships in western urban areas. Together, these results demonstrate clear spatial variation, indicating that the relationships between explanatory variables and frequent mental distress are nonstationary.
Figure 2. Spatial distribution of GWR coefficients for cognitive disability, illustrating how its relationship with frequent mental distress varies across census tracts in Washington. Areas with higher coefficients demonstrate stronger localized associations, highlighting spatial nonstationarity in the model.
Figure 3. Spatial distribution of GWR coefficients for depression, illustrating how its relationship with frequent mental distress varies across census tracts in Washington. Areas with higher coefficients demonstrate stronger localized associations, highlighting spatial nonstationarity in the model.
The Exploratory Data Analysis, Generalized Linear Regression (GLR), and Geographically Weighted Regression (GWR) analyses collectively indicate that cognitive disability, depression, and short sleep duration are significant predictors of frequent mental distress across census tracts in Washington. The GLR model demonstrated strong overall model fit, suggesting that these variables explain a substantial portion of variation in mental distress at the global scale. However, a significant Koenker (BP) statistic indicate that relationships between variables vary across space.
The GWR analysis provided additional insight by revealing clear spatial variation in these relationships. Cognitive disability showed the strongest and most consistent association in western Washington, particularly in the Seattle metropolitan area, with additional localized clusters in eastern and southern regions. In contrast, depression exhibited stronger relationships in eastern and southern Washington and weaker associations in western urban areas. These findings indicate that the influence of explanatory variables on mental distress is spatially nonstationary, and that different factors may be more strongly associated with mental distress in different regions of the state. These findings can help inform public health planning by identifying regions where factors associated with mental distress are more strongly expressed, supporting more targeted and spatially informed interventions.
Limitations
Several limitations should be considered when interpreting these results. The Jarque-Bera statistic indicated non-normal residuals, suggesting potential model bias. Additionally, spatial autocorrelation of residuals indicated clustering, which may reflect model misspecification or the omission of relevant explanatory variables. As a result, the models may not fully capture all factors influencing mental distress.
Furthermore, the analysis relies on modeled, aggregate census tract data, which may obscure individual-level variation and introduce ecological bias. Future research could improve model accuracy by incorporating additional variables, such as socioeconomic conditions or access to healthcare, and by testing alternative model specifications.
Centers for Disease Control and Prevention. (2025, April 3). Frequent mental distress. U.S. Department of Health and Human Services. https://www.cdc.gov/disability-inclusion/resources/easy-read-frequent-mental-distress.html
Centers for Disease Control and Prevention. (2022). PLACES Local Data for Better Health 2022 [Feature service]. ArcGIS. https://www.arcgis.com/home/item.html?id=cecf94bfe5cc4650b0e70967d8a8ccda
National Alliance on Mental Illness. (2023, July). Mental health in Washington [Fact sheet]. https://www.nami.org/wp-content/uploads/2023/07/WashingtonStateFactSheet.pdf