Options
台灣族群身高之遺傳基因特徵:基於機器學習的非優勢群體全基因組關聯研究
Other Title
Genetic Architecture of Body Height in the Taiwan Biobank: A Machine Learning Approach to GWAS in Underrepresented Populations
Type
thesis
Date Issued
2025-12-22
Author(s)
許鈺敏
Advisor
許明暉
Subjects
系所名稱:大數據科技及管理研究所碩士班
Publisher
大數據科技及管理研究所碩士班
Description
學位別:碩士
語文別:英文
口試委員:許明暉; 張詠淳; 張資昊
網際網路,開放日期為2031-01-02
語文別:英文
口試委員:許明暉; 張詠淳; 張資昊
網際網路,開放日期為2031-01-02
Abstract
Background and Objectives
Human height is a highly heritable complex trait extensively studied through genome-wide association studies (GWAS). However, most genetic discoveries have focused on European populations, creating knowledge gaps for non-European ancestries and limiting polygenic risk score applicability across diverse groups. Traditional GWAS approaches face limitations including reliance on additive linear models that fail to capture gene-gene interactions and insufficient statistical power in smaller populations. This study aims to develop a novel machine learning-based GWAS framework to investigate height genetics in the Taiwanese population, with objectives to: (1) develop predictive models capturing non-linear genetic effects, (2) identify novel population- and sex-specific height-associated loci, and (3) establish an optimized analytical workflow for underrepresented populations.
Methods
We employed a four-step machine learning (ML) workflow using 75,198 participants (42,266 females and 32,932 males) from the Taiwan Biobank genotyped with the TWBv2 array (570,000 markers). Step 1: Baseline model construction using Lasso regression with 21 height-related covariates underwent iterative feature selection and 5-fold cross-validation, identifying 9 key factors for males and 12 for females. Step 2: Add-One Test methodology individually integrated each of 510,276 quality-controlled SNPs with the baseline model to assess Mean Squared Error (MSE) improvement. Step 3: Statistical analysis employed geometric knee point detection for ML-enhanced Adaptive Threshold (ML-EAT) determination, and validated the results through comparisons with established false discovery rate control methods. Step 4: Results were externally validated against the Taiwan Precision Medicine Initiative (TPMI) summary statistics. Candidate SNPs are processed to identify novel genetic associations, functionally annotated and filtered using FUMA, CADD, RegulomeDB and GWAS catalog. Sex-specific analyses identified hormonal and pathway-specific effects. Analyses used Python 3.12 with appropriate multiple testing corrections and cross-validation strategies.
Results
The final baseline model explained approximately 19.2% (males) and 15.5% (females) of the phenotypic variance, achieving MSE values of 32.3582 for males and 26.0198 for females. In top 3000 ranked region, 21.6% of loci in males and 22.6% of loci in females had been reported. In gene level analysis, the male 2296 SNPs and female 2398 SNPs in top 3000 list were successfully mapped to corresponding genes, and 1640 (71.4%) SNPs in male and 1528(63.7%) SNPs in female were linked to genes previously reported to be associated with height traits in the GWAS catalog.
The knee point detector identified adaptive cutoffs at the 381st and 1,580th ranked SNP for males and females, respectively, translating into sex-specific optimal p-value thresholds of 3.41 × 10⁻¹⁰ for males and 4.19 × 10⁻⁶ for females. Validation analysis demonstrated exceptional precision: the empirical FDR was maintained at 0.14% for females and negligible (< 1 × 10⁻⁶) for males, far outperforming standard FDR control methods. External validation revealed that while the model replicated established loci, it uniquely reprioritized complex regulatory drivers such as ZBTB38—a gene associated with Idiopathic Short Stature (ISS). A total of two key height-contributing genes were identified in male while 25 key genes were identified in female. Our findings from add-one test are consistent with prior studies, as most of the identified key genes (23/26) were previously associated with height in the GWAS catalog, except for three key genes in females that had not been reported before, namely ERBB3, SUOX and DERA.
Conclusion
Our findings demonstrate that the ML-enhanced GWAS framework—specifically the ML-EAT method—offers a powerful and flexible tool to dissect the complex genetic architecture of human height in the Taiwanese population. Unlike conventional GWAS approaches relying on fixed significance thresholds and additive linear assumptions, our approach accommodates non-linear gene interactions and establishes context-aware adaptive cutoffs, demonstrating high sample efficiency in a cohort of ~75,000. This enables enhanced detection of genetic variants, as evidenced by replication of known height-related loci, successful reprioritization of the ISS-associated gene ZBTB38, and the discovery of novel, sex-specific genes such as ERBB3, SUOX, and DERA. Overall, this study underscores the promise of integrating machine learning adaptive thresholding frameworks with classical genetic epidemiology to advance understanding of polygenic traits. The findings inform the development of more sensitive, interpretable, and population-specific genomic analyses that may ultimately improve risk prediction and personalized interventions.
Human height is a highly heritable complex trait extensively studied through genome-wide association studies (GWAS). However, most genetic discoveries have focused on European populations, creating knowledge gaps for non-European ancestries and limiting polygenic risk score applicability across diverse groups. Traditional GWAS approaches face limitations including reliance on additive linear models that fail to capture gene-gene interactions and insufficient statistical power in smaller populations. This study aims to develop a novel machine learning-based GWAS framework to investigate height genetics in the Taiwanese population, with objectives to: (1) develop predictive models capturing non-linear genetic effects, (2) identify novel population- and sex-specific height-associated loci, and (3) establish an optimized analytical workflow for underrepresented populations.
Methods
We employed a four-step machine learning (ML) workflow using 75,198 participants (42,266 females and 32,932 males) from the Taiwan Biobank genotyped with the TWBv2 array (570,000 markers). Step 1: Baseline model construction using Lasso regression with 21 height-related covariates underwent iterative feature selection and 5-fold cross-validation, identifying 9 key factors for males and 12 for females. Step 2: Add-One Test methodology individually integrated each of 510,276 quality-controlled SNPs with the baseline model to assess Mean Squared Error (MSE) improvement. Step 3: Statistical analysis employed geometric knee point detection for ML-enhanced Adaptive Threshold (ML-EAT) determination, and validated the results through comparisons with established false discovery rate control methods. Step 4: Results were externally validated against the Taiwan Precision Medicine Initiative (TPMI) summary statistics. Candidate SNPs are processed to identify novel genetic associations, functionally annotated and filtered using FUMA, CADD, RegulomeDB and GWAS catalog. Sex-specific analyses identified hormonal and pathway-specific effects. Analyses used Python 3.12 with appropriate multiple testing corrections and cross-validation strategies.
Results
The final baseline model explained approximately 19.2% (males) and 15.5% (females) of the phenotypic variance, achieving MSE values of 32.3582 for males and 26.0198 for females. In top 3000 ranked region, 21.6% of loci in males and 22.6% of loci in females had been reported. In gene level analysis, the male 2296 SNPs and female 2398 SNPs in top 3000 list were successfully mapped to corresponding genes, and 1640 (71.4%) SNPs in male and 1528(63.7%) SNPs in female were linked to genes previously reported to be associated with height traits in the GWAS catalog.
The knee point detector identified adaptive cutoffs at the 381st and 1,580th ranked SNP for males and females, respectively, translating into sex-specific optimal p-value thresholds of 3.41 × 10⁻¹⁰ for males and 4.19 × 10⁻⁶ for females. Validation analysis demonstrated exceptional precision: the empirical FDR was maintained at 0.14% for females and negligible (< 1 × 10⁻⁶) for males, far outperforming standard FDR control methods. External validation revealed that while the model replicated established loci, it uniquely reprioritized complex regulatory drivers such as ZBTB38—a gene associated with Idiopathic Short Stature (ISS). A total of two key height-contributing genes were identified in male while 25 key genes were identified in female. Our findings from add-one test are consistent with prior studies, as most of the identified key genes (23/26) were previously associated with height in the GWAS catalog, except for three key genes in females that had not been reported before, namely ERBB3, SUOX and DERA.
Conclusion
Our findings demonstrate that the ML-enhanced GWAS framework—specifically the ML-EAT method—offers a powerful and flexible tool to dissect the complex genetic architecture of human height in the Taiwanese population. Unlike conventional GWAS approaches relying on fixed significance thresholds and additive linear assumptions, our approach accommodates non-linear gene interactions and establishes context-aware adaptive cutoffs, demonstrating high sample efficiency in a cohort of ~75,000. This enables enhanced detection of genetic variants, as evidenced by replication of known height-related loci, successful reprioritization of the ISS-associated gene ZBTB38, and the discovery of novel, sex-specific genes such as ERBB3, SUOX, and DERA. Overall, this study underscores the promise of integrating machine learning adaptive thresholding frameworks with classical genetic epidemiology to advance understanding of polygenic traits. The findings inform the development of more sensitive, interpretable, and population-specific genomic analyses that may ultimately improve risk prediction and personalized interventions.