机器学习方法在小麦表型和基因型研究领域的应用进展

张志贤1 , 孙萌浩1 , 陈博1 , 李志豪1 , 苏培森2 , 孟宪勇1 , 黄思罗3 , 柳平增1 , 颜君1,*
1山东农业大学信息科学与工程学院,农业农村部黄淮海智慧农业技术重点实验室,泰安271018 2聊城大学农学与农业工程学院,聊城 252000 3华中科技大学生命科学与技术学院,武汉 430074

摘 要:

随着高通量测序技术、遥感技术、机器学习技术的迅速发展,小麦(Triticum aestivum L.)研究正从传统的表型评估向数据驱动的精准分析逐渐转变。机器学习作为处理海量数据、发现潜在模式的核心技术,已在小麦研究的多个领域得到广泛应用。本文系统总结了机器学习在小麦四大核心领域的研究进展:生长状态与生理特性监测、生物胁迫与非生物胁迫识别、产量与品质预测、抗病抗逆基因挖掘。通过多源数据融合与算法创新,机器学习算法实现了对小麦生理参数的无损监测、胁迫的早期诊断、产量的精准预测以及育种候选基因的挖掘。然而,当前研究仍面临数据异质性、模型泛化能力不足、可解释性缺失等挑战。未来基于机器学习的小麦相关研究需重点突破多模态数据协同学习、机理-数据双驱动建模、轻量化部署三大方向,构建“从基因到田间”的智能化育种与管理体系,为全球粮食安全与农业可持续发展提供核心驱动力。

通讯作者:颜君 , Email:xinsinian2006@163.com

Advances in the application of machine learning methods in wheat phenotypic and genotypic research
ZHANG Zhi-Xian1 , SUN Meng-Hao1 , CHEN Bo1 , LI Zhi-Hao1 , SU Pei-Sen2 , MENG Xian-Yong1 , HUANG Si-Luo3 , LIU Ping-Zeng1 , YAN Jun1,*
1Key Laboratory of Huang-Huai-Hai Smart Agricultural Technology of the Ministry of Agriculture and Rural Affairs, College of Information Science and Engineering, Shandong Agricultural University, Taian 271018, China 2College of Agronomy and Agricultural Engineering, Liaocheng University, Liaocheng 252000, 3College of Life Science and Technology, Huazhong University of Science and Technology, Wuhan 430074, China

Abstract:

With the rapid advancement of high-throughput sequencing, remote sensing, and machine learning technologies, wheat (Triticum aestivum L.) research is gradually transitioning from traditional phenotypic evaluation to data-driven precision analysis. As a core technology for processing massive high-dimensional data and uncovering underlying complex patterns, machine learning has effectively addressed key research challenges such as wheat’s large and intricate genome (≈16 Gb) and the nonlinear relationships between genotype and phenotype. It has been deeply applied in multiple fields of wheat research and demonstrated revolutionary value. This study systematically summarizes the research progress of machine learning in four core areas of wheat research over the past decade (2015–2025): monitoring of growth status and physiological traits, identification of biotic and abiotic stresses, prediction of yield and quality, and mining of disease-resistance and stress tolerance genes. Through multi-source data fusion (including UAV remote sensing, hyperspectral imaging, 3D point clouds, multi-omics data, etc.) and algorithmic innovation (from traditional machine learning to deep learning model optimization), machine learning has achieved non-destructive and high-precision monitoring of wheat physiological parameters, early and accurate diagnosis of stresses, quantitative prediction of yield and quality, as well as efficient screening of candidate breeding genes. However, current research still faces prominent challenges: strong data heterogeneity and lack of unified standards, insufficient cross-environment generalization ability and poor interpretability of models, as well as high computational power requirements and high costs for field deployment. Future wheat-related research based on machine learning should focus on breaking through three key directions: advancing the collaborative learning of ″space-air-ground″ integrated multi-modal data, constructing a mechanism-data dual-driven modeling framework, and developing lightweight models and edge computing deployment solutions. Ultimately, by integrating technological achievements, an integrated intelligent breeding and management system ″from gene to field″ will be built, providing a core driving force for global food security and sustainable agricultural development.

Communication Author:YAN Jun , Email:xinsinian2006@163.com

Back to top