When researchers collect data using surveys or assessments, they often gather information about dozens or even hundreds of variables. Analyzing such complex datasets can be overwhelming. Factor analysis offers a solution by identifying underlying patterns in the data and reducing a large set of observed variables to a smaller number of latent factors. This statistical technique reveals hidden structures that explain why certain variables tend to move together, making it easier to understand and interpret research findings.
Table of Contents
- What is factor analysis?
- Understanding latent variables and dimensionality reduction
- Common variance and unique variance
- Factor loadings: measuring relationships
- Eigenvalues: measuring explained variance
- The scree plot method
- Types of factor analysis
- Exploratory factor analysis
- Confirmatory factor analysis
- Applications across research fields
- Data reduction and scale development
- Key considerations for conducting factor analysis
- Interpreting and using factor analysis results
What is factor analysis?
Factor analysis is a family of statistical techniques used to identify the structure and dimensionality of observed data. The fundamental premise is that when multiple observed variables are correlated with each other, this correlation exists because they share common underlying factors. These factors are latent variables-constructs that cannot be directly measured but can be inferred from patterns in the data.
Consider a practical example: if you survey employees about job satisfaction, you might ask questions about pay, work-life balance, relationships with colleagues, career growth opportunities, and job security. Factor analysis can group these variables into broader categories such as “compensation satisfaction” and “workplace environment,” revealing that certain questions cluster together because they measure similar underlying concepts.
Understanding latent variables and dimensionality reduction
The power of factor analysis lies in its ability to simplify complex data. Rather than working with numerous individual variables, researchers can focus on a smaller number of factors that capture the essence of the data. This process is called dimensionality reduction.
The observed variables are modeled as linear combinations of the latent factors plus error terms. Each factor represents an unobserved construct that influences multiple observed variables. For instance, socioeconomic status is a latent variable that cannot be directly measured, but it influences observable indicators like income, education level, and occupation. By analyzing how these indicators correlate, factor analysis can estimate the underlying socioeconomic status factor.
Common variance and unique variance
Every observed variable contains different types of variance. Common variance, also called communality, is the portion of variance that a variable shares with other variables through common factors. Unique variance consists of specific variance (variance unique to that particular variable) and error variance (random measurement error). Factor analysis aims to account for the covariation between observed variables by identifying the common factors that create these relationships.
Factor loadings: measuring relationships
Factor loadings are correlation coefficients that indicate the strength and direction of the relationship between each observed variable and each factor. A factor loading of 0.7 or higher typically indicates that the factor sufficiently captures the variance of that variable, though researchers may accept lower thresholds depending on their specific context.
Think of factor loadings as weights that show how much each variable contributes to a factor. If “salary satisfaction” has a high loading on a “compensation” factor, this means salary satisfaction strongly relates to the underlying compensation construct. Variables with high loadings on a particular factor are considered good indicators of that factor. Researchers examine these loadings to interpret what each factor represents and to name the factors based on the variables that load most heavily on them.
Eigenvalues: measuring explained variance
Eigenvalues represent the total amount of variance that can be explained by a given factor. They help researchers determine how many factors to retain in their analysis. Each factor has an associated eigenvalue, and these values indicate how much of the total variance in the observed variables each factor explains.
A common rule of thumb is the Kaiser criterion, which suggests retaining factors with eigenvalues greater than 1.0. The logic is straightforward: since each standardized variable contributes a variance of 1, a factor with an eigenvalue less than 1 explains less variance than a single variable would. However, this criterion tends to over-extract factors, so researchers often use it in combination with other methods.
The scree plot method
Another popular approach involves creating a scree plot-a line graph that displays eigenvalues in descending order. Researchers look for the “elbow” or bend in the curve where the eigenvalues level off. The number of factors before this bend typically represents the optimal number to retain. This visual method helps balance the goals of explaining sufficient variance while keeping the model as simple as possible.
Types of factor analysis
Researchers use factor analysis for different purposes, which has led to the development of two main approaches.
Exploratory factor analysis
Exploratory factor analysis (EFA) is used when researchers do not have predetermined hypotheses about the factor structure. It identifies underlying structure without prior assumptions and groups correlated variables into factors organically. This approach is ideal for initial investigations where the goal is to discover patterns in the data. For example, when developing a new psychological assessment tool, researchers might use EFA to identify how many distinct constructs the assessment measures.
Confirmatory factor analysis
Confirmatory factor analysis (CFA) tests specific hypotheses about the factor structure. It uses structural equation modeling to check how well the data fit a proposed model. Researchers use CFA when they have theoretical reasons to expect a particular factor structure and want to validate whether their data support this structure. This approach is common when validating existing measurement instruments or testing theories about construct relationships.
Applications across research fields
Factor analysis has found widespread application across numerous disciplines. In psychology, researchers have used it to identify factors underlying intelligence, with studies finding that people who score high on verbal ability tests also tend to perform well on other verbal tasks. This led to the identification of factors like verbal intelligence, spatial reasoning, and numerical ability.
In market research, companies use factor analysis to understand customer preferences and segment their markets. A business might survey customers about numerous product attributes and then use factor analysis to identify the key dimensions that drive purchasing decisions. This information guides product development and marketing strategies.
Healthcare researchers employ factor analysis to develop and validate assessment instruments. For instance, studies examining service quality in hospitals might measure multiple aspects of patient experience and use factor analysis to identify broader dimensions like tangibles, reliability, responsiveness, and empathy.
Data reduction and scale development
One practical application involves reducing questionnaire data to manageable dimensions. When developing survey instruments, researchers need evidence that items truly measure the same underlying construct. Factor analysis provides this evidence by showing which items cluster together and how many distinct constructs the instrument measures.
Key considerations for conducting factor analysis
Successful factor analysis requires adequate sample size. A common guideline suggests at least 20 observations per variable, though the exact requirement depends on the strength of the relationships being studied. The technique also assumes that relationships between variables and factors are linear and that the data meet certain distributional requirements.
Factor rotation is an important step that makes results easier to interpret. Rotation methods like varimax reorganize the factor structure to create clearer patterns, where each variable loads strongly on one factor rather than spreading its loadings across multiple factors. This simplification helps researchers identify and name the underlying constructs more easily.
Interpreting and using factor analysis results
After extracting and rotating factors, researchers must interpret what each factor represents. This requires examining which variables load highly on each factor and finding a common theme among them. The interpretation process involves both statistical analysis and subject-matter expertise.
Once factors are identified and interpreted, they can be used in subsequent analyses. Researchers can calculate factor scores-estimated values of the latent factors for each observation in the dataset. These scores can then serve as variables in other statistical models, such as regression analyses or group comparisons, allowing researchers to work with the simplified factor structure rather than the original complex variable set.
What do you think? How might factor analysis help simplify complex datasets in your field of research? When would exploratory versus confirmatory approaches be most appropriate for your research questions?
References
- https://online.stat.psu.edu/stat505/lesson/12
- https://www.publichealth.columbia.edu/research/population-health-methods/exploratory-factor-analysis
- https://www.driveresearch.com/market-research-company-blog/factor-analysis-definition-types-and-examples/
- https://en.wikipedia.org/wiki/Factor_analysis
- https://www.statisticssolutions.com/free-resources/directory-of-statistical-analyses/factor-analysis/
- https://stats.oarc.ucla.edu/spss/seminars/introduction-to-factor-analysis/a-practical-introduction-to-factor-analysis/
- https://support.minitab.com/en-us/minitab/help-and-how-to/statistical-modeling/multivariate/how-to/factor-analysis/interpret-the-results/all-statistics-and-graphs/
- https://www.geeksforgeeks.org/machine-learning/introduction-to-factor-analytics/
- https://www.philomathresearch.com/blog/2024/09/10/exploring-factor-analysis-in-research-key-types-and-examples/
- https://www.lifescied.org/doi/10.1187/cbe.18-04-0064
- https://support.minitab.com/en-us/minitab/help-and-how-to/statistical-modeling/multivariate/how-to/factor-analysis/methods-and-formulas/methods-and-formulas/
- https://statisticsbyjim.com/basics/factor-analysis/
Leave a Reply