An intelligent breeding method for economic crops based on big data analysis

Through multimodal data fusion and plant electrical signal analysis, a seed vigor assessment model was constructed, which solved the problem of inaccurate seed vigor assessment in traditional breeding methods, realized dynamic monitoring and evaluation of economic crops from seeds to early growth process, and improved breeding efficiency and variety quality.

CN120356515BActive Publication Date: 2025-09-09SHANXI ZHONGNONG NEW ERA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510816238.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-09
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Traditional economic crop breeding methods rely on single data to evaluate seed vitality, which makes it difficult to comprehensively and accurately measure seed vitality. In addition, there is a lack of means to evaluate the physiological state in the early growth stage, resulting in low breeding efficiency and inability to screen out high-quality varieties.

Method used

By collecting multimodal data of seeds, including physical properties, physiological indicators and genomic information, a deep learning model is constructed to evaluate seed vitality. In combination with high-precision plant electrical signal acquisition equipment, changes in electrical signals during crop growth stages are monitored, and a model for the association between electrical signals and traits is established. Data fusion analysis is performed to achieve continuous monitoring and evaluation of the seed to early growth process.

Benefits of technology

It improves breeding efficiency, can accurately evaluate seed vitality and early growth status, ensure the scientific nature of breeding decisions, and cultivate high-quality economic crop varieties to meet the needs of modern agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356515B_ABST
    Figure CN120356515B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent breeding method for economic crops based on big data analysis, which relates to the field of agricultural breeding technology. The method comprises the following specific steps: seed multimodal data collection and processing: collecting physical property data of seeds, obtaining physiological indicators and genome information data through biochemical experiments and gene sequencing, and preprocessing them; the present invention integrates multimodal data and, after preprocessing, uses deep learning and data fusion technology to construct a precise seed vitality assessment model, which can more accurately measure the seed vitality level. At the same time, high-precision plant electrical signal acquisition equipment is used to monitor the changes in electrical signals in the early growth stage of crops in real time, and establishes a correlation model between electrical signal characteristics and multiple crop traits, and fuses the seed vitality assessment data with the plant electrical signal data for analysis, which greatly improves the breeding efficiency and provides strong support for the cultivation of high-quality varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural breeding technology, and specifically to an intelligent breeding method for economic crops based on big data analysis. Background Art

[0002] In the continuous development of agricultural breeding technology, cash crop breeding has always been a key and extremely challenging field. With the improvement of scientific and technological levels, a variety of emerging technologies have made it possible to apply multimodal data fusion and plant electrical signal analysis technology in cash crop breeding. Among them, multimodal data fusion technology can integrate data from different sources and in different forms, thereby providing more comprehensive and rich information. Plant electrical signal analysis technology can deeply understand the physiological state and internal mechanism of plants by monitoring and analyzing plant electrical signals. The development of these technologies has brought new ideas and methods to cash crop breeding, and promoted breeding work to develop in a more efficient and accurate direction.

[0003] Traditional economic crop breeding methods have many limitations. In terms of seed vigor assessment, traditional methods often rely only on a single physical property or physiological indicator of the seed. For example, seed vigor is judged only by physical properties such as seed size and weight, or is evaluated based on only a few physiological indicators such as respiration rate and enzyme activity. However, seed vigor is a complex comprehensive characteristic affected by many factors. A single type of data is difficult to comprehensively and accurately measure the actual vitality level of the seed. In addition, traditional methods lack effective means to evaluate the physiological state and trait potential of economic crops in the early growth stage. During the early growth of crops, it is difficult to timely discover individuals that have high seed vigor but cannot show good trait potential in actual growth. This leads to low breeding efficiency and difficulty in screening out truly high-quality seeds and varieties, which restricts the development of economic crop breeding and cannot meet the needs of modern agriculture for efficient and precise breeding. Summary of the Invention

[0004] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide an intelligent breeding method for economic crops based on big data analysis. It can collect physical property data, physiological indicator data and genomic information data of seeds and perform preprocessing to provide an accurate data basis for subsequent evaluation, and use deep learning and data fusion technology to construct a precise seed vitality evaluation model, which can more accurately evaluate the vitality level of seeds. At the same time, high-precision plant electrical signal acquisition equipment is used to monitor the changes in electrical signals of crops from seedling emergence to early growth stage in real time, and machine learning technology is used for analysis and decoding to extract effective electrical signal features. By establishing a correlation model between electrical signal features and multiple crop traits, the intrinsic connection between electrical signals and crop traits is explored. Finally, the multimodal data of seed vitality evaluation and plant electrical signal data in the early growth stage of crops are fused and analyzed to achieve continuous and dynamic monitoring and evaluation of economic crops from seed to early growth process.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: an intelligent breeding method for economic crops based on big data analysis, the method comprising the following specific steps:

[0006] Seed multimodal data collection and processing: Collect physical property data of seeds, obtain physiological indicators and genomic information data through biochemical experiments and gene sequencing, and pre-process them;

[0007] Construction of a precise seed vitality assessment model: Build a model based on a deep learning framework, divide multimodal data into training sets, validation sets, and test sets to train the model, and adjust parameters based on the validation set;

[0008] Plant electrical signal acquisition and analysis: Use high-precision equipment to collect electrical signals from different parts of the plant. After denoising preprocessing, extract electrical signal features at different scales.

[0009] Establishment of a model for the association between electrical signals and traits: Collect and analyze electrical signals from crop samples with different traits under the same environment, compare characteristic differences, identify relevant characteristic indicators through statistical analysis, and establish a model to describe the quantitative relationship between electrical signals and traits;

[0010] Data fusion analysis and comprehensive evaluation: Integrate and fuse seed multimodal data with plant electrical signal data, deeply explore potential connections, and build a comprehensive evaluation model to dynamically monitor crop growth status;

[0011] Breeding screening and decision-making: Breeding screening decisions are made based on the comparison of comprehensive evaluation values ​​with set thresholds.

[0012] Furthermore, in the seed multimodal data collection and processing step, the physical property data include the size, weight and morphology of the seeds, and the physiological indicator data of the seeds include the respiration rate and enzyme activity.

[0013] The model formula is: ,in, Represents the seed vitality evaluation value, which is the quantitative result of seed vitality based on comprehensive multimodal data. A data set representing the physical properties of a seed, Represents a set of physiological indicator data of seeds, A data set representing the genomic information of seeds, They are transformation functions for physical property data, physiological index data, and genomic information data, respectively. are the weight coefficients of the corresponding transformation functions, reflecting the importance of different data features to seed vitality assessment, Respectively represent the number of transformation functions involved in physical property data, physiological indicator data, and genomic information data.

[0014] Furthermore, in the plant electrical signal collection and analysis step, high-precision equipment is used to collect electrical signals from different parts of the plant, including roots, stems and leaves. The collected original plant electrical signals are preprocessed and feature extraction is performed. The characteristics of the electrical signals in the time dimension and the frequency dimension are analyzed to extract effective electrical signal features that can reflect the physiological state and growth characteristics of the crop.

[0015] Furthermore, in the plant electrical signal collection and analysis step, the collected original plant electrical signals are preprocessed and feature extraction is performed, and the feature extraction formula is: ,in, Indicates at time The extracted electrical signal eigenvalues, It is at the moment The collected plant electrical signal values, is the number of data points in the sliding window, is the moment when the feature is currently calculated, is the size of the sliding window.

[0016] Furthermore, in the step of establishing the electrical signal and trait association model, economic crop samples with different traits, including disease resistance, drought resistance and nutrient absorption capacity, are selected. Under the same growth environment and conditions, the plant electrical signals of the samples are collected and analyzed, and the differences in the electrical signal characteristics of crops with different traits are compared and analyzed. Through statistical analysis, the electrical signal characteristic indicators related to each trait are found. Based on the relationship between the electrical signal characteristics and the traits obtained by analysis, an association model that can describe the quantitative relationship between the electrical signal characteristics and the crop traits is established.

[0017] Furthermore, in the step of establishing the electrical signal and trait association model, an association model is established that can describe the quantitative relationship between electrical signal characteristics and crop traits, and the model formula is: ,in, Represents a certain trait value of the crop, including disease resistance, drought tolerance and nutrient absorption capacity, are different eigenvalues ​​extracted from plant electrical signals, They are the corresponding electrical signal characteristics and The weight coefficient of Represent the number of linear features and quadratic features, is the bias term, which is used to adjust the output of the model. It is an activation function used to introduce nonlinearity so that the model can learn the complex relationship between electrical signal features and crop traits.

[0018] Furthermore, in the data fusion analysis and comprehensive evaluation step, the multimodal data of seed vitality evaluation are fused with the plant electrical signal data of the early growth stage of the crop, and the fused data are deeply analyzed to explore the potential connection between seed vitality-related indicators and subsequent plant electrical signal changes, establish a comprehensive evaluation model, and continuously and dynamically monitor and evaluate the status of economic crops from seeds to early growth processes, giving comprehensive evaluation results of crops at different growth stages.

[0019] Furthermore, in the data fusion analysis and comprehensive evaluation step, the potential connection between seed vitality-related indicators and subsequent plant electrical signal changes is explored to establish a comprehensive evaluation model, the model formula of which is: ,in, Represents the comprehensive assessment value of economic crops from seeds to early growth stages, is the seed vigor assessment value, Indicates the Crop trait values, represents the total number of crop traits considered, is the seed vigor assessment value In the comprehensive evaluation value The weight coefficient in , , is the weight coefficient of the crop trait evaluation value in the comprehensive evaluation value, , It is Crop trait values Weight coefficients in comprehensive evaluation of crop traits , , is a constant term, representing the comprehensive evaluation value of other factors not considered except the association between seed vigor and electrical signal and traits impact.

[0020] Compared with the existing technology, this intelligent breeding method for economic crops based on big data analysis has the following beneficial effects:

[0021] 1. The present invention integrates multimodal data and, after preprocessing, uses deep learning and data fusion technology to construct an accurate seed vigor assessment model, which can more accurately measure the seed vigor level. At the same time, it uses high-precision plant electrical signal acquisition equipment to monitor the changes in electrical signals in the early growth stage of crops in real time, and establishes a correlation model between electrical signal characteristics and multiple crop traits. The seed vigor assessment data and plant electrical signal data are integrated and analyzed, which greatly improves breeding efficiency and provides strong support for the cultivation of high-quality varieties.

[0022] 2. The present invention deeply explores the intrinsic connection between seed vitality and the early growth status of crops through multimodal data fusion and plant electrical signal analysis. In terms of seed vitality assessment, the precise model can accurately predict the actual vitality level of seeds. In the early growth stage of crops, changes in plant electrical signals can reflect their physiological state and trait potential. Through data fusion analysis and comprehensive evaluation, breeders can fully understand the vitality of seeds and the growth of crops, so as to make more scientific breeding decisions, which will help to cultivate higher-quality economic crop varieties and meet the demand of modern agriculture for high-quality varieties.

[0023] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0025] Figure 1 This is a process operation diagram of an intelligent breeding method for economic crops based on big data analysis;

[0026] Figure 2 This is a flow chart of an intelligent breeding method for economic crops based on big data analysis. DETAILED DESCRIPTION

[0027] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0028] Example 1

[0029] like Figure 1-2 As shown, at a tomato seed breeding base, for a batch of newly cultivated tomato seed samples, the size of each seed is first accurately measured, and the weight of each tomato seed is weighed one by one using an electronic balance, and the data is recorded in detail. The seeds are photographed from multiple angles using a high-resolution microscope to obtain clear seed morphological images. In order to obtain physiological index data, a certain number of tomato seeds are placed in a strictly sealed transparent container with a known volume. The container is equipped with high-precision oxygen and carbon dioxide sensors. Under a specific temperature (such as 25°C) and humidity (such as 60%) environment, the changes in the oxygen and carbon dioxide concentrations in the container are continuously monitored within a certain time interval (such as every hour). The reduction in oxygen or carbon dioxide per unit time is used to determine the change in the concentration of oxygen and carbon dioxide. Increase the amount and accurately calculate the respiration rate of the seeds. At the same time, catalase and amylase, which are closely related to the vitality of tomato seeds, are selected, and these enzymes are extracted from the seeds using professional biochemical experimental methods. The corresponding substrates and reaction buffers are prepared, and the enzymes are allowed to react with the substrates under appropriate temperature (such as 37°C) and pH value (such as 7.0). By measuring the amount of product generated or the amount of substrate consumed during the reaction, the activity of the enzyme is accurately determined. For the collection of genomic information data, a DNA extraction kit is used, and strict operating procedures are followed to efficiently extract genomic DNA from tomato seeds. The extracted genomic DNA is sequenced using the next-generation gene sequencing technology to obtain the complete genomic sequence information of the tomato seeds.

[0030] After obtaining all the data, these multimodal data are cleaned. By setting a reasonable data range and statistical methods, outliers (such as measurements that deviate significantly from the normal range) and missing values ​​(data that were not successfully obtained) are identified and removed. For missing values, they are filled according to the distribution characteristics and correlation of the data. Finally, the normalization processing method is used to uniformly map all data to the interval of [0, 1] to eliminate the dimensional differences between different data features and provide a standardized data basis for subsequent seed vigor assessment.

[0031] Construct a seed vitality evaluation formula:

[0032] ,in represents the seed vigor evaluation value, A data set representing the physical properties of tomato seeds, Represents a set of physiological indicator data of tomato seeds, A data set representing the genomic information of tomato seeds, 、 、 They are transformation functions for physical property data, physiological index data, and genomic information data, respectively. 、 、 are the weight coefficients of the corresponding transformation functions, 、 、 Represents the number of transformation functions involved in physical property data, physiological indicator data, and genomic information data, respectively. The processed multimodal data are divided into training set, validation set, and test set in a ratio of 7:2:1. The training set is used to train the model, and the gradient descent method is used as the optimization algorithm to calculate the seed vigor value predicted by the model. The mean square error between the actual seed vigor value measured by traditional methods and the weight parameter is calculated based on the mean square error. 、 、 The gradient of the gradient is then updated in the opposite direction of the gradient. After each training iteration, the performance of the model is evaluated using the validation set. The weight parameters are fine-tuned based on the mean square error and other performance indicators on the validation set. For example, if the mean square error on the validation set does not decrease significantly or even increases in several consecutive iterations, it indicates that overfitting may have occurred. In this case, the value of the weight parameter should be appropriately reduced. If the mean square error decreases slowly, the step size of the weight parameter adjustment may be too small. In this case, the step size should be appropriately increased. When the performance of the model on the validation set reaches a stable and satisfactory state, the test set is used to perform a final verification of the model to ensure that the model can accurately evaluate the vitality of tomato seeds.

[0033] During the emergence stage after tomato sowing, when the tomato seedlings grow to a certain height (such as about 5 cm), a specially customized high-precision plant electrical signal acquisition device is used. The device consists of highly sensitive electrodes, signal amplifiers and data acquisition cards. The electrodes are carefully placed at appropriate positions on the roots of the tomato plants (about 1 cm from the root tip), stems (2 cm from the ground) and leaves (select mature functional leaves and avoid leaf veins) to ensure that the electrodes are in close contact with the plant tissues and do not cause damage to the plants. The acquisition equipment is placed in an environment with good shielding effect to reduce the impact of external electromagnetic interference on electrical signal acquisition. During the entire stage from emergence to early growth of tomato plants, plant electrical signals are collected in real time at a fixed sampling frequency (such as 100 Hz), and the collected electrical signal data is stored in a computer. The collected raw electrical signals are preprocessed, and a digital filtering algorithm (such as a low-pass filter) is used to remove high-frequency noise interference in the signal to improve the signal quality. Then, a feature calculation formula based on a sliding window is used. The preprocessed electrical signal is subjected to feature extraction, where Indicates at time The extracted electrical signal eigenvalues, It is at the moment The collected electrical signal values ​​of tomato plants, is the number of data points in the sliding window, is the moment of current feature calculation, by adjusting the window size (Tried from 10 sampling points to 100 sampling points), calculated the mean, variance, frequency and other characteristic values ​​of the electrical signal in each sliding window to obtain the electrical signal characteristics at different time scales, and provided rich data information for subsequent analysis of the relationship between electrical signals and tomato traits.

[0034] Constructing a formula to associate electrical signals with traits ,in Represents a certain trait value of tomato, such as disease resistance, drought tolerance, nutrient absorption capacity, etc. 、 are the different eigenvalues ​​extracted from the electrical signals of tomato plants, 、 They are the corresponding electrical signal characteristics and The weight coefficient of 、 Represent the number of linear features and quadratic features, is a bias term, It is an activation function used to introduce nonlinearity so that the model can learn the complex relationship between electrical signal features and tomato traits. When determining the weights corresponding to each electrical signal feature in the model, the extracted electrical signal features are first 、 and the trait values ​​of tomato Perform correlation analysis and calculate the Pearson correlation coefficient between them. According to the size of the correlation coefficient, the weight corresponding to the electrical signal feature with higher correlation is 、 The initial setting was a relatively high value. Meanwhile, the weights were further adjusted and optimized with reference to relevant research results in the fields of plant physiology and electrophysiology to ensure that the model could accurately reflect the potential relationship between electrical signal characteristics and various tomato traits.

[0035] The collected tomato sample data was used to train the association model, and the stochastic gradient descent method was used as the optimization algorithm to calculate the trait values ​​predicted by the model. The mean square error between the measured trait value and the actual measured trait value is used to calculate the error with respect to the weight parameter 、 、 The gradient of the model is obtained, and then the weight parameters are updated in the opposite direction of the gradient. During the training process, the cross-validation method is used to divide the sample data into multiple subsets, which are used as training sets and validation sets in turn to improve the generalization ability of the model. When the performance of the model on the validation set reaches stability and meets certain accuracy requirements, the test set is used to perform a final verification of the model to ensure that the model can accurately predict the various traits of tomatoes based on the electrical signal characteristics.

[0036] Constructing a comprehensive evaluation formula ,in Represents the comprehensive evaluation value of tomatoes from seeds to early growth stages, is the seed vigor assessment value, which is calculated by the seed vigor precision assessment model. Indicates the The trait values ​​of various tomatoes are calculated by the association model between electrical signals and traits. represents the total number of tomato traits considered, is the seed vigor assessment value In the comprehensive evaluation value The weight coefficient in , is the tomato trait evaluation value ( In the comprehensive evaluation value The weight coefficient in , and , It is Tomato trait values Weight coefficient in comprehensive evaluation of tomato traits, , is a constant term that represents the comprehensive evaluation value of other factors other than seed vigor and the association between electrical signals and traits. The weight of seed vigor evaluation value was preliminarily determined based on the actual needs of tomato breeding, such as paying more attention to fruit quality and disease resistance, and referring to the contribution ratio of seed vigor and various traits to the final breeding effect in historical breeding data. , the weight of disease resistance evaluation value , the weight of drought tolerance assessment value , the weight of the nutrient absorption capacity assessment value , In practical applications, genetic algorithms are used to optimize weight parameters. Genetic algorithms simulate the selection, crossover, and mutation operations in the biological evolution process to search for the optimal weight combination within a certain range. The fitness function is set as the correlation coefficient between the comprehensive evaluation value S and the actual tomato growth performance (such as yield, quality indicators, etc.). Through continuous iterative evolution, the weight combination that maximizes the fitness function is found. At the same time, according to the average deviation between the model prediction value and the actual value, the constant term is Fine-tuning was performed so that the comprehensive assessment results could more accurately reflect the true status of tomatoes from seeds to early growth stages.

[0037] With reference to the distribution of comprehensive evaluation values ​​of high-quality tomato varieties cultivated in the past, combined with current breeding goals and market demand, and through multiple experiments, a screening threshold is set. In the tomato seed production and sales process, all seeds are evaluated one by one according to the comprehensive evaluation model, and seeds with comprehensive evaluation values ​​greater than or equal to the screening threshold are screened out. These seeds are marked as high-quality seeds with high vigor and good trait potential and are given priority for promotion and sales. In the early growth stage of tomatoes, the planted tomato plants are continuously monitored and evaluated, and the electrical signal data of the plants are regularly collected. Combined with other growth indicators (such as plant height, leaf color, flowering time, etc.), the comprehensive evaluation value of each tomato is recalculated. For those tomato plants with high seed vigor but a comprehensive evaluation value below the screening threshold in actual growth, which cannot show good trait potential, they are marked and eliminated in time to ensure that the planted tomato population has high quality and yield potential. Through this precise breeding screening method, high-quality tomato varieties with high seed vigor, excellent physiological characteristics and multiple resistances are eventually bred to meet the market demand for high-quality tomatoes.

[0038] Example 2

[0039] In the experimental field for wheat seed cultivation, the physical dimensions of each seed, such as length, width and thickness, were measured for wheat seeds of different varieties. A large number of randomly selected wheat seeds were weighed one by one using an electronic balance, and the weight data of each seed was recorded in detail. With the help of a high-resolution industrial-grade scanner, clear appearance images of the wheat seeds were obtained in high-pixel mode. Then, professional image analysis software was used to carefully extract and quantitatively analyze the morphological characteristics of the seeds, such as shape, color, and texture. In order to obtain physiological index data, a certain number of wheat seeds were selected and placed in a special sealed glass container. The container was equipped with a high-precision gas concentration sensor that can monitor the changes in oxygen and carbon dioxide concentrations in real time. The container was placed in a constant temperature and humidity incubator with a set temperature of 20°C and a relative humidity of 50%. The oxygen and carbon dioxide concentration data were recorded every 30 minutes within 4 hours, and the respiration rate of the seeds per unit time was calculated based on the difference in concentration changes. At the same time, a professional biochemical experimental process was used to extract superoxide dismutase (SOD) and peroxidase (POD), which are closely related to the vitality of wheat seeds, from the seeds. The rate of the enzyme-catalyzed reaction was measured using a specific enzyme activity detection kit under specified reaction conditions (such as temperature 30°C and pH 7.5) to evaluate the level of enzyme activity. For the collection of genomic information data, the modified CTAB method was used to extract genomic DNA from wheat seeds to ensure the purity and integrity of the DNA. The extracted genomic DNA was sequenced using second-generation sequencing technology to obtain complete and high-precision genome sequence information of the wheat seeds. After completing data collection, the multimodal data is comprehensively cleaned. By setting reasonable statistical thresholds, outliers are identified and eliminated. For missing values, multiple imputation methods are used to fill them according to data type and relevance to ensure the integrity and accuracy of the data. Subsequently, a normalization algorithm is used to normalize all data to the interval of [0, 1] to eliminate the influence of different data dimensions and dimensions, providing a standardized data basis for subsequent seed vigor assessment.

[0040] The model is built based on the deep learning framework, and its model formula is: ,in, Represents the seed vitality evaluation value, which is the quantitative result of seed vitality based on comprehensive multimodal data. A data set representing the physical properties of a seed, Represents a set of physiological indicator data of seeds, A data set representing the genomic information of seeds, They are transformation functions for physical property data, physiological index data, and genomic information data, respectively. are the weight coefficients of the corresponding transformation functions, reflecting the importance of different data features to seed vitality assessment, Respectively represent the number of transformation functions involved in physical property data, physiological indicator data, and genomic information data. When determining the weights corresponding to each data feature in the model, the equal weight method is used for preliminary setting, that is, it is assumed that physical properties, physiological indicators, and genomic information data have the same importance in evaluating seed vigor, and they are given equal initial weights. The preprocessed multimodal data are divided into training set, validation set, and test set in a ratio of 8:1:1. The training set is used to train the model, and the stochastic gradient descent algorithm is used as the optimization strategy. The seed vigor value predicted by the model is compared with the actual seed vigor that passes the standard. The mean square error between the true values ​​obtained by the test methods (such as germination rate test, vitality index calculation, etc.) is calculated, the gradient of the error with respect to the weight parameter is calculated, and the weight parameter is updated in the opposite direction of the gradient. During the training process, the performance of the model is evaluated using the validation set every certain number of iterations (such as 100 times). The weight parameters are fine-tuned based on the MSE and other performance indicators on the validation set to prevent the model from overfitting. When the performance of the model on the validation set tends to be stable and meets the expected accuracy requirements, the model is finally verified using the test set to ensure that the model can accurately and reliably evaluate the vitality of wheat seeds.

[0041] During the three-leaf stage after wheat seedlings are sown, a self-developed high-precision plant electrical signal acquisition system is used. The system consists of high-impedance, low-noise electrodes, high-gain signal amplifiers, and high-speed data acquisition cards. The electrodes are placed in a specific layout at the roots of the wheat plants (0.5 cm from the tip of the main root), the base of the stem (1 cm from the ground), and the leaves (the middle of the second fully expanded leaf) respectively. Conductive glue is used to ensure good contact between the electrodes and the plant tissues. The acquisition system is placed in an electromagnetic shielding test box to reduce the impact of external electromagnetic interference.

[0042] During the critical growth stage of wheat plants from seedling emergence to jointing, plant electrical signals were collected in real time at a sampling frequency of 200 Hz. The collected data was transmitted to a data storage server via optical fiber. The collected raw electrical signals were preprocessed to remove noise components in the signals and improve the signal-to-noise ratio. Then, by changing the length of the sliding window (experiments were conducted from 20 sampling points to 200 sampling points), features of the preprocessed electrical signals were extracted, and the time domain features (such as mean, variance, peak-to-peak value) and frequency domain features (such as power spectral density and frequency center of gravity) of the electrical signals within the sliding window were calculated. The electrical signal characteristics reflecting the physiological state of wheat plants at different scales were obtained, providing rich feature information for the subsequent analysis of the relationship between electrical signals and wheat traits.

[0043] Wheat samples with different disease resistance (such as wheat plants that are immune, resistant, or susceptible to stripe rust and powdery mildew), drought tolerance (wheat plants grown under drought stress and normal water supply conditions), and nutrient absorption capacity (determined by measuring the nitrogen, phosphorus, and potassium content in the aboveground parts of the plants and root activity) were selected, and these samples were planted in an artificial climate chamber with strictly controlled environmental conditions to ensure consistency of growth factors such as light, temperature, humidity, and soil nutrients.

[0044] A model for associating electrical signals with traits was constructed. The input layer received the extracted electrical signal feature data, the hidden layer performed nonlinear transformation on the features through activation functions, and the output layer outputted the values ​​of various wheat traits (such as disease resistance grade, drought tolerance score, nutrient absorption efficiency index, etc.). When determining the weights corresponding to each electrical signal feature in the model, a typical correlation analysis was first performed on the extracted electrical signal features and the values ​​of various wheat traits to find the typical correlation variables between the electrical signal features and the traits. According to the size of the typical correlation coefficient, the weights corresponding to the electrical signal features with higher correlation with the traits were initially set to larger values. At the same time, combined with the theoretical knowledge of wheat physiology and electrophysiology, the weights were further adjusted and optimized so that the model could more accurately reflect the intrinsic connection between the electrical signal features and the various wheat traits.

[0045] The association model was trained using the collected wheat sample data. The weight parameters were adjusted according to the root mean square error between the trait values ​​predicted by the model and the actual measured trait values. During the training process, the sample data was divided into 5 subsets, which were used as training sets and validation sets in turn to improve the generalization ability of the model. When the root mean square error of the model on the validation set reached the minimum and tended to be stable, the test set was used to perform final verification of the model to ensure that the model could accurately predict the various traits of wheat based on the electrical signal characteristics.

[0046] A comprehensive evaluation formula was constructed to integrate the seed vigor evaluation value and the wheat trait evaluation values ​​obtained through the electrical signal and trait association model. According to the goals of wheat breeding, such as the pursuit of high-yield, high-quality, and multi-resistant varieties, combined with the contribution ratio of each factor to breeding success in historical breeding data, the weight of the seed vigor evaluation value was preliminarily determined to be 0.3, the weight of the disease resistance evaluation value was 0.3, the weight of the drought tolerance evaluation value was 0.2, and the weight of the nutrient absorption capacity evaluation value was 0.2. In practical applications, the particle swarm optimization algorithm was used to optimize the weight parameters. At the same time, the constant term in the comprehensive evaluation formula was dynamically adjusted according to the deviation between the model predicted value and the actual value, so that the comprehensive evaluation result can more accurately reflect the comprehensive state of wheat from seed to early growth stage.

[0047] With reference to the distribution range of comprehensive evaluation values ​​of high-yield and high-quality wheat varieties cultivated in the past, combined with current market demand and actual agricultural production conditions, a reasonable screening threshold is set through a large number of field experiments and data analysis. In the wheat seed production and sales links, all seeds are evaluated and screened using a comprehensive evaluation model, and seeds with comprehensive evaluation values ​​greater than or equal to the screening threshold are selected, marked as high-quality seeds, and supplied to farmers for planting on a priority basis. In the early growth stage of wheat, a real-time monitoring system is established to regularly collect electrical signal data and other growth indicators of wheat plants, and calculate the comprehensive evaluation value of each wheat plant. Wheat plants with high seed vitality but a comprehensive evaluation value lower than the screening threshold in actual growth, which show the potential for adverse traits, are marked and removed in a timely manner to ensure the overall quality and yield potential of the wheat population.

[0048] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. An intelligent breeding method for economic crops based on big data analysis, characterized in that: The method comprises the following specific steps: Seed multimodal data collection and processing: Collect physical property data of seeds, obtain physiological indicators and genomic information data through biochemical experiments and gene sequencing, and pre-process them; Construction of a precise seed vitality assessment model: The model is built based on a deep learning framework, and its model formula is: ,in, Represents the seed vitality evaluation value, which is the quantitative result of seed vitality based on comprehensive multimodal data. A data set representing the physical properties of a seed, Represents a set of physiological indicator data of seeds, A data set representing the genomic information of a seed, They are transformation functions for physical property data, physiological index data, and genomic information data, respectively. are the weight coefficients of the corresponding transformation functions, reflecting the importance of different data features to seed vitality assessment, Represents the number of transformation functions involved in physical property data, physiological indicator data, and genomic information data respectively. The multimodal data is divided into training set, validation set, and test set to train the model, and the parameters are adjusted according to the validation set; Plant electrical signal collection and analysis: High-precision equipment is used to collect electrical signals from different parts of the plant. After denoising preprocessing, the electrical signal features of different scales are extracted. The feature extraction formula is: ,in, Indicates at time The extracted electrical signal eigenvalues, It is at the moment The collected plant electrical signal values, is the number of data points in the sliding window, is the moment when the feature is currently calculated, is the size of the sliding window; Establishment of the association model between electrical signals and traits: Collect and analyze electrical signals of crop samples with different traits under the same environment, compare the characteristic differences, find relevant characteristic indicators through statistical analysis, and establish a model to describe the quantitative relationship between electrical signals and traits. The model formula is: ,in, Represents a certain trait value of the crop, including disease resistance, drought tolerance and nutrient absorption capacity, are different eigenvalues ​​extracted from plant electrical signals, They are the corresponding electrical signal characteristics and The weight coefficient of Represent the number of linear features and quadratic features, is the bias term, which is used to adjust the output of the model. It is an activation function used to introduce nonlinearity, enabling the model to learn the complex relationship between electrical signal characteristics and crop traits; Data fusion analysis and comprehensive evaluation: Integrate and fuse seed multimodal data and plant electrical signal data, deeply explore potential connections, and build a comprehensive evaluation model to dynamically monitor crop growth status. The model formula is: ,in, Represents the comprehensive assessment value of economic crops from seeds to early growth stages, is the seed vigor assessment value, Indicates the Crop trait values, represents the total number of crop traits considered, is the seed vigor assessment value In the comprehensive evaluation value The weight coefficient in , , The crop trait evaluation value is the comprehensive evaluation value The weight coefficient in , , It is Crop trait values The weight coefficient in the comprehensive evaluation of crop traits, , is a constant term, representing the comprehensive evaluation value of other factors not considered except the association between seed vigor and electrical signal and traits the impact of; Breeding screening and decision-making: Breeding screening decisions are made based on the comparison of comprehensive evaluation values ​​with set thresholds.

2. The method for intelligent breeding of economic crops based on big data analysis according to claim 1, characterized in that: In the seed multimodal data collection and processing step, the physical property data include the size, weight and morphology of the seeds, and the physiological indicator data of the seeds include the respiratory rate and enzyme activity.

3. The method for intelligent breeding of economic crops based on big data analysis according to claim 1, characterized in that: In the plant electrical signal collection and analysis step, high-precision equipment is used to collect electrical signals from different parts of the plant, including roots, stems and leaves. The collected original plant electrical signals are preprocessed and feature extraction is performed. The characteristics of the electrical signals in the time dimension and the frequency dimension are analyzed to extract effective electrical signal features that can reflect the physiological state and growth characteristics of the crop.

4. The method for intelligent breeding of economic crops based on big data analysis according to claim 1, characterized in that: In the step of establishing the electrical signal and trait association model, economic crop samples with different traits, including disease resistance, drought tolerance and nutrient absorption capacity, are selected. Plant electrical signals of the samples are collected and analyzed under the same growth environment and conditions. The differences in electrical signal characteristics of crops with different traits are compared and analyzed. Through statistical analysis, electrical signal characteristic indicators related to each trait are found. Based on the relationship between the electrical signal characteristics and the traits obtained by analysis, an association model that can describe the quantitative relationship between the electrical signal characteristics and the crop traits is established.

5. The method for intelligent breeding of economic crops based on big data analysis according to claim 1, characterized in that: In the data fusion analysis and comprehensive evaluation step, the multimodal data of seed vitality evaluation are fused with the plant electrical signal data of the early growth stage of the crop, and the fused data are deeply analyzed to explore the potential connection between seed vitality-related indicators and subsequent plant electrical signal changes. A comprehensive evaluation model is established to continuously and dynamically monitor and evaluate the status of economic crops from seeds to early growth processes, and provide comprehensive evaluation results of crops at different growth stages.

Citation Information

Patent Citations

  • Forest tree planting article selection prediction method and system based on big data

    CN119358821A

  • Comprehensive evaluation method for quality traits of plant breeding

    CN119782885A