-
Machine learning-guided discovery of thermophilic carbonic anhydrases from environmental metagenomes
- Back
Metadata
Document Title
Machine learning-guided discovery of thermophilic carbonic anhydrases from environmental metagenomes
Name from Authors Collection
Affiliations
Enzyme Technology Research Team, Biorefinery Technology and Bioproduct Research Group, National Center for Genetic Engineering and Biotechnology, Khlong Luang, Pathum Thani, 12120, Thailand; Biorefinery Technology and Bioproduct Research Group, National Center for Genetic Engineering and Biotechnology, Khlong Luang, Pathum Thani, 12120, Thailand; Department of Biochemistry and Center for Excellence in Protein and Enzyme Technology, Faculty of Science, Mahidol University, Ratchathewi, Bangkok, 10400, Thailand; Department of Biotechnology, Faculty of Science and Technology, Thammasat University, Rangsit Campus, Phahonyothin Road, Khlong Luang, Pathum Thani, 12120, Thailand
Type
Article
Source Title
Scientific Reports
ISSN
20452322
Year
2025
Volume
15
Issue
1
Open Access
All Open Access; Gold Open Access; Green Open Access
Publisher
Nature Research
DOI
10.1038/s41598-025-24713-1
Abstract
Thermophilic carbonic anhydrases (CAs) are promising biocatalysts for carbon capture utilization and storage (CCUS) due to their stability and efficiency at elevated temperatures. This study presents a machine learning (ML)-guided approach to discover thermostable γ-class CA (γ-CA) from metagenomic datasets derived from Fang Hot Spring, Northern Thailand. To develop classification models, two sets of protein descriptors—dipeptide composition (DPC) and physicochemical/biochemical properties (AAindex)—were used to train classification models. Fourteen ML algorithms were systematically evaluated for each feature set. AdaBoost achieved the best performance for the DPC-based model, while LightGBM performed best with AAindex-based features. External validation with known CA sequences confirmed the ability of the models to discriminate thermophilic from non-thermophilic proteins. Applying the optimized models, we screened 1,534 predicted CAs and identified three high-confidence candidates (TtCA, CrCA, and ToCA). These were heterologously expressed in E. coli, purified, and biochemically validated. All candidates exhibited carbonic anhydrase activity, trimeric oligomeric structures, and high melting temperatures (Tm ranging from 97.0 °C to 109.1 °C). Although their hydration activity was modest compared to α-class CAs, their thermal robustness highlights their potential for industrial CO₂ capture. This study demonstrates an approach in which ML integrated with metagenomics enables efficient discovery and validation of robust enzymes from extreme environments, providing a scalable strategy for CCUS applications. © The Author(s) 2025.
License
CC BY-NC-ND
Rights
Authors
Publication Source
Scopus
Publication Source
Scopus