Intelligent breeding method for commercial crops based on big data analysis

Through multimodal data fusion and plant electrical signal analysis, a seed vitality assessment model was constructed, which solved the problem of inaccurate seed vitality assessment in traditional breeding methods, and achieved dynamic monitoring and evaluation of the growth process of economic crops from seed to early stages, improving breeding efficiency and variety quality.

CN120356515AActive Publication Date: 2025-07-22SHANXI ZHONGNONG NEW ERA TECH CO LTD

Patent Information

Application Number
CN202510816238.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Traditional economic crop breeding methods lack precision in seed vitality assessment, making it difficult to comprehensively measure seed vitality levels, and insufficient assessment of the physiological status and trait potential of crop early growth stages, resulting in low breeding efficiency and inability to meet the efficient and precise breeding needs of modern agriculture.

Method used

By collecting multimodal data of seeds (physical characteristics, physiological indicators and genomic information), a deep learning model is constructed to evaluate seed vitality, and combined with high-precision plant electrical signal acquisition equipment to monitor the electrical signal changes in the early growth stage of crops, establish an electrical signal and trait correlation model, and perform data fusion analysis to achieve continuous and dynamic monitoring and evaluation of the economic crops from seed to early growth.

Benefits of technology

It improves breeding efficiency, can more accurately evaluate seed vitality level and predict early growth status of crops, ensure the scientific nature of breeding decisions, cultivate high-quality varieties, and meet the needs of modern agriculture for high-quality varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356515A_ABST
    Figure CN120356515A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent breeding method for commercial crops based on big data analysis, and relates to the technical field of agricultural breeding, the method comprises the following specific steps: seed multi-modal data acquisition and processing: acquiring physical characteristic data of seeds, acquiring physiological indexes and genome information data through biochemical experiments and gene sequencing, and determining the seed multi-modal data according to the physiological indexes and the genome information data; carrying out pretreatment on the raw materials; according to the method, after multi-modal data are integrated and preprocessed, a seed vigor accurate evaluation model is constructed by using deep learning and data fusion technologies, the seed vigor level can be measured more accurately, and meanwhile, the electric signal change of the early growth stage of crops is monitored in real time by using high-precision plant electric signal acquisition equipment, so that the accuracy of the seed vigor evaluation is improved. And a correlation model of the electric signal characteristics and multiple traits of crops is established, and fusion analysis is performed on seed vigor evaluation data and plant electric signal data, so that the breeding efficiency is greatly improved, and powerful support is provided for cultivation of high-quality varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural breeding, and particularly to an intelligent breeding method for cash crops based on big data analysis. Background Art

[0002] In the process of the continuous development of agricultural breeding technology, cash crop breeding has always been a key and extremely challenging field. With the improvement of the scientific and technological level, a variety of emerging technologies have made it possible to apply multi-modal data fusion and plant electrical signal analysis technology in cash crop breeding. Among them, multi-modal data fusion technology can integrate data from different sources and different forms, so as to provide more comprehensive and rich information. Plant electrical signal analysis technology, through the monitoring and analysis of plant electrical signals, deeply understands the physiological state and internal mechanism of plants. The development of these technologies has brought new ideas and methods to cash crop breeding, and promoted the breeding work to develop in a more efficient and accurate direction.

[0003] Traditional cash crop breeding methods have many limitations. In terms of seed vigor assessment, traditional methods often rely only on a single physical characteristic or physiological index of the seeds. For example, the seed vigor is judged only by physical characteristics such as the size and weight of the seeds, or only evaluated based on a few physiological indexes such as respiration rate and enzyme activity. However, seed vigor is a complex comprehensive characteristic affected by multiple factors, and single-type data is difficult to comprehensively and accurately measure the actual vigor level of the seeds. In addition, for the assessment of the physiological state and trait potential of cash crops in the early growth stage, traditional methods lack effective means. During the early growth process of the crops, it is difficult to timely discover those individuals with high seed vigor but unable to show good trait potential in actual growth, which results in low breeding efficiency, difficult to screen out truly high-quality seeds and varieties, limits the development of cash crop breeding, and cannot meet the requirements of modern agriculture for efficient and accurate breeding. Summary of the Invention

[0004] The objective of the present invention is to make up for the deficiencies of the prior art and provide an intelligent breeding method for cash crops based on big data analysis. It can collect physical characteristic data, physiological index data, and genomic information data of seeds, and perform preprocessing to provide an accurate data basis for subsequent evaluation. By using deep learning and data fusion technologies, a precise seed vigor evaluation model is constructed to more accurately evaluate the vigor level of seeds. At the same time, a high-precision plant electrical signal acquisition device is used to continuously monitor the changes in electrical signals of crops from emergence to the early growth stage, and machine learning technologies are used for analysis and decoding to extract effective electrical signal features. By establishing an association model between electrical signal features and various crop traits, the internal relationship between electrical signals and crop traits is explored. Finally, the multi-modal data for seed vigor evaluation and the plant electrical signal data during the early growth stage of crops are fused and analyzed to achieve continuous and dynamic monitoring and evaluation of cash crops from seeds to the early growth process.

[0005] To solve the above technical problems, the present invention provides the following technical solution: An intelligent breeding method for cash crops based on big data analysis, the method comprising the following specific steps: Collection and processing of seed multi-modal data: Collect physical characteristic data of seeds, obtain physiological index and genomic information data through biochemical experiments and gene sequencing, and perform preprocessing on them; Construction of a precise seed vigor evaluation model: Build a model based on a deep learning framework, divide the multi-modal data into a training set, a validation set, and a test set to train the model, and adjust the parameters according to the validation set; Collection and analysis of plant electrical signals: Use a high-precision device to collect electrical signals from different parts of the plant, and after denoising preprocessing, extract electrical signal features at different scales; Establishment of an association model between electrical signals and traits: Collect and analyze electrical signals of crop samples with different traits under the same environment, compare the characteristic differences, and through statistical analysis, find relevant characteristic indicators to establish a model describing the quantitative relationship between electrical signals and traits; Data fusion analysis and comprehensive evaluation: Integrate the seed multi-modal data and the plant electrical signal data and perform fusion, deeply explore potential connections, and build a comprehensive evaluation model to dynamically monitor the growth status of crops; Breeding screening and decision-making: Make breeding screening decisions based on the comparison between the comprehensive evaluation value and the set threshold.

[0006] Further, in the step of collection and processing of seed multi-modal data, the physical characteristic data includes the size, weight, and morphology of the seeds, and the physiological index data of the seeds includes the respiration rate and enzyme activity.

[0007] Its model formula is: , where, represents the seed vigor evaluation value, which is the quantitative result of seed vigor by integrating multi-modal data. represents the set of physical characteristic data of seeds. represents the set of physiological index data of seeds. represents the set of genomic information data of seeds. are transformation functions for physical characteristic data, physiological index data, and genomic information data respectively. are the weight coefficients corresponding to the respective transformation functions, reflecting the importance of different data characteristics for seed vigor evaluation. respectively represent the number of transformation functions involved in physical characteristic data, physiological index data, and genomic information data.

[0008] Furthermore, in the step of plant electrical signal acquisition and analysis, high-precision equipment is used to collect electrical signals from different parts of the plant, including the roots, stems, and leaves. The collected original plant electrical signals are preprocessed and feature extraction is performed. The features of the electrical signals in the time dimension and frequency dimension are analyzed, and effective electrical signal features that can reflect the physiological state and growth characteristics of the crop are extracted.

[0009] Furthermore, in the step of plant electrical signal acquisition and analysis, the collected original plant electrical signals are preprocessed and feature extraction is performed. The feature extraction formula is: , where represents the electrical signal feature value extracted at time is the plant electrical signal value collected at time is the number of data points within the sliding window. is the current time for calculating the feature. is the size of the sliding window.

[0010] Furthermore, in the step of establishing the electrical signal and trait association model, economic crop samples with different traits are selected, including disease resistance, drought tolerance, and nutrient absorption ability. Under the same growth environment and conditions, plant electrical signals of the samples are collected and analyzed. The differences in electrical signal features of crops with different traits are compared and analyzed. Through statistical analysis, electrical signal feature indicators related to each trait are found. According to the relationship between the electrical signal features and traits obtained from the analysis, an association model that can describe the quantitative relationship between electrical signal features and crop traits is established.

[0011] Furthermore, in the step of establishing the electrical signal and trait association model, an association model that can describe the quantitative relationship between electrical signal features and crop traits is established. The model formula is: , where ​​Represents a certain trait value of the crop, including disease resistance, drought tolerance and nutrient absorption capacity. are different feature values extracted from plant electrical signals, They are the corresponding electrical signal characteristics and The weight coefficient of Represent the number of linear and quadratic features, respectively. is the bias term, which is used to adjust the output of the model. It is an activation function used to introduce nonlinearity so that the model can learn the complex relationship between electrical signal features and crop traits.

[0012] Furthermore, in the data fusion analysis and comprehensive evaluation step, the multimodal data of seed vitality evaluation is fused with the plant electrical signal data of the early growth stage of the crop, and the fused data is deeply analyzed to explore the potential connection between seed vitality-related indicators and subsequent plant electrical signal changes, establish a comprehensive evaluation model, and continuously and dynamically monitor and evaluate the status of economic crops from seeds to early growth processes, giving comprehensive evaluation results of crops at different growth stages.

[0013] Furthermore, in the data fusion analysis and comprehensive evaluation step, the potential connection between seed vitality-related indicators and subsequent plant electrical signal changes is explored to establish a comprehensive evaluation model, and the model formula is: ,in, Represents the comprehensive evaluation value of cash crops from seeds to early growth stages, is the seed vigor assessment value, Indicates Crop trait values, represents the total number of crop traits considered, is the seed vigor assessment value In the comprehensive evaluation The weight coefficient in , is the weight coefficient of the crop trait evaluation value in the comprehensive evaluation value, , It is Crop trait values Weight coefficients in comprehensive evaluation of crop traits , , is a constant term, representing the comprehensive evaluation value of factors other than seed vigor and the association between electrical signals and traits. impact.

[0014] Compared with the existing technology, this economic crop intelligent breeding method based on big data analysis has the following beneficial effects: 1. After integrating and preprocessing multi-modal data, the present invention constructs an accurate evaluation model for seed vigor by using deep learning and data fusion technologies, which can more accurately measure the seed vigor level. At the same time, a high-precision plant electrical signal acquisition device is used to monitor the changes in electrical signals during the early growth stage of crops in real time, and an association model between electrical signal characteristics and various traits of crops is established. By fusing and analyzing the seed vigor evaluation data and plant electrical signal data, the breeding efficiency is greatly improved, providing strong support for cultivating high-quality varieties.

[0015] 2. Through multi-modal data fusion and plant electrical signal analysis, the present invention deeply explores the internal relationship between seed vigor and the early growth state of crops. In terms of seed vigor evaluation, the accurate model can accurately predict the actual vigor level of seeds. During the early growth stage of crops, the changes in plant electrical signals can reflect their physiological state and trait potential. Through data fusion analysis and comprehensive evaluation, breeders can comprehensively understand the vigor of seeds and the growth of crops, thereby making more scientific breeding decisions, which helps to cultivate higher-quality economic crop varieties and meet the needs of modern agriculture for high-quality varieties.

[0016] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a flow operation diagram of an intelligent breeding method for economic crops based on big data analysis; Figure 2 It is a flow chart of an intelligent breeding method for economic crops based on big data analysis. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention objective, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and their effects of the present invention as follows.

[0020] Example 1

[0021] As Figure 1-2As shown, in the cultivation base of tomato seeds, for a batch of newly cultivated tomato seed samples, first, accurately measure the size of each seed, use an electronic balance to weigh each tomato seed one by one, and record the data in detail. Through a high-resolution microscope, take multi-angle photos of the seeds to obtain clear images of the seed morphology. To obtain physiological index data, place a certain number of tomato seeds in a transparent container with a known volume that has been strictly sealed, and the container is equipped with high-precision oxygen and carbon dioxide sensors. Under specific temperature (such as 25°C) and humidity (such as 60%) environments, continuously monitor the changes in the oxygen and carbon dioxide concentrations in the container at certain time intervals (such as every hour). According to the decrease in oxygen or the increase in carbon dioxide per unit time, accurately calculate the respiration rate of the seeds. At the same time, select catalase and amylase that are closely related to the vitality of tomato seeds, and use professional biochemical experimental methods to extract these enzymes from the seeds. Prepare the corresponding substrates and reaction buffers, and let the enzymes react with the substrates under suitable temperature (such as 37°C) and pH value (such as 7.0) conditions. By measuring the generation amount of products or the consumption amount of substrates during the reaction process, accurately determine the activity level of the enzymes. For the collection of genomic information data, use a DNA extraction kit, and according to a strict operation process, efficiently extract genomic DNA from tomato seeds. Use a new generation of gene sequencing technology to perform whole-genome sequencing on the extracted genomic DNA to obtain the complete genomic sequence information of tomato seeds.

[0022] After obtaining all the data, clean these multi-modal data. By setting reasonable data ranges and statistical methods, identify and remove outliers (such as measurement values that deviate significantly from the normal range) and missing values (data that were not successfully obtained). For missing values, fill them according to the distribution characteristics and correlations of the data. Finally, use a normalization processing method to map all the data uniformly to the interval [0, 1], eliminating the dimensional differences between different data characteristics, and providing a standardized data basis for subsequent seed vitality assessment.

[0023] Construct a seed vitality assessment formula: , where represents the seed vitality assessment value, represents the set of physical characteristic data of tomato seeds, represents the set of physiological index data of tomato seeds, represents the set of genomic information data of tomato seeds, 、 、 are transformation functions for physical characteristic data, physiological index data, and genomic information data respectively, 、 、 They are the weight coefficients of the corresponding transformation functions, , , respectively represent the number of transformation functions involved in physical characteristic data, physiological index data, and genomic information data. The processed multimodal data is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The training set is used to train the model, and the gradient descent method is used as the optimization algorithm to calculate the seed vigor value predicted by the model and the mean square error between the actual seed vigor value measured by the traditional method. The gradient of the error with respect to the weight parameters , , is calculated, and then the weight parameters are updated along the opposite direction of the gradient. After each training iteration, the performance of the model is evaluated using the validation set. According to the mean square error and other performance indicators on the validation set, the weight parameters are fine-tuned. For example, if the mean square error on the validation set does not decrease significantly or even increases in several consecutive iterations, it indicates that overfitting may have occurred. At this time, the value of the weight parameters is appropriately reduced. If the mean square error decreases slowly, it indicates that the step size of the weight parameter adjustment may be too small, and the step size is appropriately increased. When the performance of the model on the validation set reaches a stable and satisfactory state, the test set is used to finally verify the model to ensure that the model can accurately evaluate the vigor of tomato seeds.

[0024] During the seedling emergence stage after tomato sowing, when the tomato seedlings grow to a certain height (such as about 5 cm), a specially customized high-precision plant electrical signal acquisition device is used. This device consists of highly sensitive electrodes, a signal amplifier, and a data acquisition card. The electrodes are carefully placed at appropriate positions on the roots (about 1 cm from the root tip), stems (2 cm from the ground), and leaves (select mature functional leaves, avoiding leaf veins) of the tomato plants respectively, ensuring that the electrodes are in close contact with the plant tissues and will not cause damage to the plants. The acquisition device is placed in an environment with good shielding effect to reduce the influence of external electromagnetic interference on the electrical signal acquisition. During the entire stage from seedling emergence to early growth of the tomato plants, the plant electrical signals are collected in real time at a fixed sampling frequency (such as 100 Hz), and the collected electrical signal data is stored in a computer. The collected original electrical signals are preprocessed, and a digital filtering algorithm (such as a low-pass filter) is used to remove the high-frequency noise interference in the signals to improve the signal quality. Then, a sliding window-based feature calculation formula is used to extract features from the preprocessed electrical signals, where represents the electrical signal feature value extracted at time , is the tomato plant electrical signal value collected at time , is the number of data points within the sliding window. is the moment for calculating the current feature. By adjusting the window size (trying from 10 sampling points to 100 sampling points), calculate the eigenvalue such as the mean, variance, frequency, etc. of the electrical signal within each sliding window, so as to obtain the electrical signal features at different time scales, and provide rich data information for subsequent analysis of the relationship between the electrical signal and tomato traits.

[0025] Construct the association formula between the electrical signal and the trait , where represents a certain trait value of the tomato, such as disease resistance, drought tolerance, nutrient absorption ability, etc., 、 are different eigenvalue extracted from the electrical signal of the tomato plant, 、 are the weight coefficients corresponding to the electrical signal features and respectively, 、 represent the number of linear term features and quadratic term features respectively, is a bias term, is an activation function, which is used to introduce non-linearity so that the model can learn the complex relationship between the electrical signal features and tomato traits. When determining the weights corresponding to each electrical signal feature in the model, first perform a correlation analysis on the extracted electrical signal features 、 and each trait value of the tomato, calculate their Pearson correlation coefficients. According to the magnitude of the correlation coefficients, initially set the weights 、 corresponding to the electrical signal features with higher correlation to relatively high values. At the same time, refer to the relevant research results in the fields of plant physiology and electrophysiology to further adjust and optimize the weights to ensure that the model can accurately reflect the potential relationship between the electrical signal features and each trait of the tomato.

[0026] Use the collected tomato sample data to train the association model, adopt the stochastic gradient descent method as the optimization algorithm, calculate the mean square error between the predicted trait value of the model and the actually measured trait value, and calculate the error with respect to the weight parameters 、 、 The gradient is then used to update the weight parameters in the opposite direction of the gradient. During the training process, the cross-validation method is employed to divide the sample data into multiple subsets and alternately use them as the training set and the validation set to improve the generalization ability of the model. When the performance of the model on the validation set reaches stability and meets certain accuracy requirements, the test set is used to finally validate the model to ensure that the model can accurately predict the various traits of tomatoes based on the electrical signal characteristics.

[0027] Construct a comprehensive evaluation formula , where represents the comprehensive evaluation value of tomatoes from the seed to the early growth stage, is the seed vigor evaluation value, which is calculated by the seed vigor precise evaluation model, represents the th tomato trait value, which is calculated by the electrical signal and trait association model, represents the total number of types of tomato traits considered, is the seed vigor evaluation value in the comprehensive evaluation value weight coefficient, is the tomato trait evaluation value ( in the comprehensive evaluation value weight coefficient, and , represents the th tomato trait value weight coefficient in the comprehensive evaluation of tomato traits, , is a constant term, representing the influence of other unconsidered factors on the comprehensive evaluation value except for seed vigor and the association between electrical signals and traits. According to the actual requirements of tomato breeding, such as paying more attention to fruit quality and disease resistance, and referring to the contribution ratios of seed vigor and various traits to the final breeding effect in historical breeding data, the weight of the seed vigor evaluation value is initially determined , the weight of the disease resistance evaluation value , the weight of the drought tolerance evaluation value , the weight of the nutrient absorption ability evaluation value , , in practical applications, a genetic algorithm is used to optimize the weight parameters. The genetic algorithm searches for the optimal weight combination within a certain range by simulating operations such as selection, crossover, and mutation in the biological evolution process. The fitness function is set as the correlation coefficient between the comprehensive evaluation value S and the actual tomato growth performance (such as yield, quality indicators, etc.). Through continuous iterative evolution, the weight combination that maximizes the fitness function is found. At the same time, according to the average deviation between the model prediction value and the actual value, the constant term Fine-tuning is carried out to make the comprehensive evaluation results more accurately reflect the true state of tomatoes from seeds to the early growth stage.

[0028] Referring to the distribution of comprehensive evaluation values of high-quality tomato varieties cultivated in the past, combining the current breeding goals and market demands, through multiple experiments, a screening threshold is set. In the production and sales links of tomato seeds, all seeds are evaluated one by one according to the comprehensive evaluation model, and the seeds with comprehensive evaluation values greater than or equal to the screening threshold are selected and marked as high-quality seeds with high vigor and good trait potential, and are preferentially used for promotion and sales. In the early growth stage of tomatoes, continuous monitoring and evaluation are carried out on the planted tomato plants, the electrical signal data of the plants are collected regularly, and combined with other growth indicators (such as plant height, leaf color, flowering time, etc.), the comprehensive evaluation value of each tomato plant is recalculated. For those tomato plants with relatively high seed vigor but with comprehensive evaluation values lower than the screening threshold in actual growth and unable to show good trait potential, timely marking and elimination are carried out to ensure that the planted tomato population has high quality and yield potential. Through this precise breeding screening method, high-quality tomato varieties with high seed vigor, excellent physiological characteristics and multiple resistances are finally cultivated to meet the market demand for high-quality tomatoes.

[0029] Example 2

[0030] In the experimental field for wheat seed cultivation, the physical dimensions of each seed, such as length, width and thickness, were measured for wheat seeds of different strains. A large number of randomly selected wheat seeds were weighed one by one using an electronic balance, and the weight data of each seed was recorded in detail. With the help of a high-resolution industrial-grade scanner, clear appearance images of wheat seeds were obtained in high-pixel mode. Then, professional image analysis software was used to carefully extract and quantitatively analyze the morphological characteristics of the seeds, such as shape, color, and texture. In order to obtain physiological indicator data, a certain number of wheat seeds were selected and placed in a special sealed glass container. The container was equipped with a high-precision gas concentration sensor, which can monitor the changes in oxygen and carbon dioxide concentrations in real time. The container was placed in a constant temperature and humidity incubator with a set temperature of 20°C and a relative humidity of 50%. The concentration data of oxygen and carbon dioxide were recorded every 30 minutes within 4 hours, and the respiration rate of seeds per unit time was calculated based on the difference in concentration changes. At the same time, professional biochemical experimental procedures were used to extract superoxide dismutase (SOD) and peroxidase (POD), which are closely related to the vitality of wheat seeds, from the seeds. The rate of enzyme-catalyzed reaction was measured under specified reaction conditions (such as temperature 30°C and pH 7.5) using a specific enzyme activity detection kit to evaluate the enzyme activity. For the collection of genomic information data, the modified CTAB method was used to extract genomic DNA from wheat seeds to ensure the purity and integrity of the DNA. The second-generation sequencing technology was used to perform whole-genome sequencing on the extracted genomic DNA to obtain complete and high-precision genome sequence information of wheat seeds. After completing data collection, the multimodal data is comprehensively cleaned. By setting reasonable statistical thresholds, outliers are identified and eliminated. For missing values, multiple imputation methods are used to fill them according to data type and relevance to ensure data integrity and accuracy. Subsequently, a normalization algorithm is used to normalize all data to the interval of [0, 1] to eliminate the influence of different data dimensions and dimensions, providing a standardized data basis for subsequent seed vigor evaluation.

[0031] The model is built based on the deep learning framework, and its model formula is: ,in, Represents the seed vitality assessment value, which is the quantitative result of seed vitality based on comprehensive multimodal data. A collection of data representing the physical properties of a seed, Represents a set of physiological index data of seeds, A data set representing the genomic information of seeds, They are transformation functions for physical property data, physiological index data, and genomic information data, respectively. are the weight coefficients of the corresponding transformation functions, reflecting the importance of different data features to seed vitality assessment, respectively represent the number of transformation functions involved in physical characteristic data, physiological index data, and genomic information data. When determining the weights corresponding to each data feature in the model, the equal-weight method is used for preliminary setting, that is, it is assumed that physical characteristics, physiological indexes, and genomic information data have the same importance in evaluating seed vigor, and equal initial weights are assigned to them. The preprocessed multimodal data is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set is used to train the model, and the stochastic gradient descent algorithm is used as the optimization strategy. According to the mean square error between the predicted seed vigor value of the model and the true value obtained by the actual standard seed vigor test method (such as germination rate test, vigor index calculation, etc.), the gradient of the error with respect to the weight parameter is calculated, and the weight parameter is updated along the opposite direction of the gradient. During the training process, at regular intervals of a certain number of iterations (such as 100 times), the validation set is used to evaluate the performance of the model. According to the MSE and other performance indicators on the validation set, the weight parameters are fine-tuned to prevent the model from overfitting. When the performance of the model on the validation set tends to be stable and meets the expected accuracy requirements, the test set is used to conduct the final validation of the model to ensure that the model can accurately and reliably evaluate the vigor of wheat seeds.

[0032] At the three-leaf stage after wheat sowing and emergence, an independently developed high-precision plant electrical signal acquisition system is used. This system consists of high-impedance, low-noise electrodes, high-gain signal amplifiers, and high-speed data acquisition cards. The electrodes are placed at the root (0.5 cm from the tip of the main root), the base of the stem (1 cm from the ground), and the leaf (the middle part of the second fully expanded leaf) of the wheat plant according to a specific layout method. Conductive glue is used to ensure good contact between the electrodes and the plant tissue. The acquisition system is placed in a test chamber with electromagnetic shielding treatment to reduce the influence of external electromagnetic interference.

[0033] During the critical growth stage of wheat plants from emergence to jointing stage, plant electrical signals are collected in real time at a sampling frequency of 200 Hz. The collected data is transmitted to the data storage server through optical fibers. The collected original electrical signals are preprocessed to remove the noise components in the signals and improve the signal-to-noise ratio. Then, by changing the length of the sliding window (testing from 20 sampling points to 200 sampling points), feature extraction is performed on the preprocessed electrical signals, and the time-domain features (such as mean, variance, peak-to-peak value) and frequency-domain features (such as power spectral density, frequency centroid) of the electrical signals within the sliding window are calculated to obtain electrical signal features reflecting the physiological state of wheat plants at different scales, providing rich feature information for subsequent analysis of the relationship between electrical signals and wheat traits.

[0034] Select wheat samples with different disease resistances (such as wheat plants showing immunity, disease resistance, and susceptibility to stripe rust and powdery mildew), drought tolerances (wheat plants grown under drought stress and normal water supply conditions), and nutrient absorption capacities (judged by measuring the nitrogen, phosphorus, and potassium contents in the above-ground parts of the plants and root activity). Plant these samples in an artificial climate chamber with strictly controlled environmental conditions to ensure that growth factors such as light, temperature, humidity, and soil nutrients are consistent.

[0035] Construct a correlation model between electrical signals and traits. The input layer receives the extracted electrical signal feature data. The hidden layer performs non-linear transformation on the features through an activation function. The output layer outputs the trait values of wheat (such as disease resistance levels, drought tolerance scores, nutrient absorption efficiency indicators, etc.). When determining the weights corresponding to each electrical signal feature in the model, first perform a canonical correlation analysis on the extracted electrical signal features and the trait values of wheat to find the canonical correlation variables between the electrical signal features and the traits. According to the magnitude of the canonical correlation coefficients, initially set the weights corresponding to the electrical signal features with higher correlations with the traits to larger values. At the same time, combined with the theoretical knowledge of wheat physiology and electrophysiology, further adjust and optimize the weights so that the model can more accurately reflect the internal relationship between the electrical signal features and the traits of wheat.

[0036] Use the collected wheat sample data to train the correlation model. Adjust the weight parameters according to the root mean square error between the trait values predicted by the model and the actually measured trait values. During the training process, divide the sample data into 5 subsets and take turns as the training set and the validation set to improve the generalization ability of the model. When the root mean square error of the model on the validation set reaches the minimum and tends to be stable, use the test set to finally validate the model to ensure that the model can accurately predict the traits of wheat based on the electrical signal features.

[0037] Construct a comprehensive evaluation formula to fuse the seed vigor evaluation value and the trait evaluation values of wheat obtained from the correlation model between electrical signals and traits. According to the goals of wheat breeding, such as pursuing high-yield, high-quality, and multi-resistant varieties, combined with the contribution ratios of various factors to breeding success in historical breeding data, initially determine the weight of the seed vigor evaluation value to be 0.3, the weight of the disease resistance evaluation value to be 0.3, the weight of the drought tolerance evaluation value to be 0.2, and the weight of the nutrient absorption capacity evaluation value to be 0.2. In practical applications, use the particle swarm optimization algorithm to optimize the weight parameters. At the same time, dynamically adjust the constant term in the comprehensive evaluation formula according to the deviation between the model prediction value and the actual value, so that the comprehensive evaluation result can more accurately reflect the comprehensive state of wheat from the seed to the early growth stage.

[0038] Referring to the distribution range of the comprehensive evaluation values of previously cultivated high-yield and high-quality wheat varieties, combining with the current market demand and the actual situation of agricultural production, through a large number of field trials and data analysis, a reasonable screening threshold is set. In the wheat seed production and sales links, all seeds are evaluated and screened using the comprehensive evaluation model, and the seeds with comprehensive evaluation values greater than or equal to the screening threshold are selected and marked as high-quality seeds, which are preferentially supplied to farmers for planting. In the early growth stage of wheat, a real-time monitoring system is established to regularly collect the electrical signal data and other growth indicators of wheat plants, and calculate the comprehensive evaluation value of each wheat plant. For wheat plants with high seed vigor but with comprehensive evaluation values lower than the screening threshold in actual growth, showing the potential for poor traits, they are promptly marked and removed to ensure the overall quality and yield potential of the wheat population.

[0039] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes, but as long as the technical content of the present invention is not departed from, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An intelligent breeding method for cash crops based on big data analysis, characterized in that, The method includes the following specific steps: Seed multi-modal data collection and processing: Collect physical characteristic data of seeds, obtain physiological index and genomic information data through biochemical experiments and gene sequencing, and preprocess them; Precise evaluation model construction of seed vigor: Build a model based on a deep learning framework, divide the multi-modal data into a training set, a validation set, and a test set to train the model, and adjust the parameters according to the validation set; Plant electrical signal collection and analysis: Use high-precision equipment to collect electrical signals from different parts of the plant, and after denoising preprocessing, extract electrical signal features at different scales; Establishment of the correlation model between electrical signals and traits: Collect and analyze electrical signals of crop samples with different traits in the same environment, compare the characteristic differences, find relevant characteristic indicators through statistical analysis, and establish a model describing the quantitative relationship between electrical signals and traits; Data fusion analysis and comprehensive evaluation: Integrate seed multi-modal data and plant electrical signal data and perform fusion, deeply explore potential connections, and construct a comprehensive evaluation model to dynamically monitor the growth state of crops; Breeding screening and decision-making: Make breeding screening decisions based on the comparison between the comprehensive evaluation value and the set threshold.

2. The intelligent breeding method of cash crops based on big data analysis according to claim 1, characterized in that In the step of seed multi-modal data collection and processing, the physical characteristic data includes the size, weight, and morphology of the seeds, and the physiological index data of the seeds includes the respiration rate and enzyme activity.

3. The intelligent breeding method for cash crops based on big data analysis according to claim 1, wherein, In the steps of constructing the precise seed vigor evaluation model, a model is built based on a deep learning framework, and its model formula is: , where represents the seed vigor evaluation value, which is the quantification result of seed vigor based on multi-modal data. represents the set of physical characteristic data of the seeds. represents the set of physiological index data of the seeds. represents the set of genomic information data of the seeds. are the transformation functions for physical characteristic data, physiological index data, and genomic information data respectively. are the weight coefficients of the corresponding transformation functions, reflecting the importance of different data characteristics for seed vigor evaluation. represent the number of transformation functions involved in physical characteristic data, physiological index data, and genomic information data respectively.

4. The intelligent breeding method for cash crops based on big data analysis according to claim 1, wherein, In the step of plant electrical signal collection and analysis, use high-precision equipment to collect electrical signals from different parts of the plant, including the roots, stems, and leaves, preprocess the collected original plant electrical signals and perform feature extraction, analyze the features of the electrical signals in the time dimension and the frequency dimension, and extract effective electrical signal features that can reflect the physiological state and growth characteristics of the crops.

5. The intelligent breeding method for cash crops based on big data analysis according to claim 4, wherein, In the step of plant electrical signal collection and analysis, preprocess the collected original plant electrical signals and perform feature extraction. The feature extraction formula is as follows: , where represents the eigenvalue of the electrical signal extracted at time , is the value of the plant electrical signal collected at time , is the number of data points within the sliding window, is the time when the current feature is calculated, is the size of the sliding window.

6. The intelligent breeding method for cash crops based on big data analysis according to claim 1, wherein, In the step of establishing the correlation model between electrical signals and traits, select economic crop samples with different traits, including disease resistance, drought tolerance, and nutrient absorption ability. Under the same growth environment and conditions, collect and analyze the plant electrical signals of the samples, compare and analyze the differences in the electrical signal characteristics of crops with different traits, and through statistical analysis, find the electrical signal characteristic indicators related to each trait. According to the relationship between the electrical signal characteristics and traits obtained from the analysis, establish a correlation model that can describe the quantitative relationship between the electrical signal characteristics and crop traits.

7. The intelligent breeding method of cash crops based on big data analysis according to claim 6, characterized in that, In the step of establishing the electro-signal and trait association model, an association model that can describe the quantitative relationship between electro-signal characteristics and crop traits is established, and its model formula is: , where represents a certain trait value of the crop, including disease resistance, drought tolerance, and nutrient absorption capacity, are different characteristic values extracted from plant electro-signals, are the weight coefficients corresponding to the electro-signal characteristics and respectively, respectively represent the numbers of linear term characteristics and quadratic term characteristics, is the bias term, used to adjust the output of the model, is the activation function, used to introduce non-linearity so that the model can learn the complex relationship between electro-signal characteristics and crop traits.

8. The intelligent breeding method for cash crops based on big data analysis according to claim 1, characterized in that, In the step of data fusion analysis and comprehensive evaluation, fuse the multi-modal data for seed vigor evaluation and the plant electrical signal data in the early growth stage of the crop, and deeply analyze the fused data, explore the potential connection between the indicators related to seed vigor and the subsequent changes in plant electrical signals, establish a comprehensive evaluation model, continuously and dynamically monitor and evaluate the state of the economic crop from the seed to the early growth process, and give the comprehensive evaluation results of the crop at different growth stages.

9. An intelligent breeding method for cash crops based on big data analysis according to claim 8, characterized in that, In the data fusion analysis and comprehensive evaluation steps, potential connections between indicators related to seed vigor and subsequent plant electrical signal changes are mined, and a comprehensive evaluation model is established. The model formula is as follows: , where represents the comprehensive evaluation value of cash crops from the seed stage to the early growth stage, is the seed vigor evaluation value, represents the th crop trait value, represents the total number of types of crop traits considered, is the seed vigor evaluation value in the comprehensive evaluation value weight coefficient, , is the weight coefficient of the crop trait evaluation value in the comprehensive evaluation value , , is the th crop trait value weight coefficient in the comprehensive evaluation of crop traits, , is the constant term, representing the influence of other unconsidered factors on the comprehensive evaluation value except for the associations between seed vigor, electrical signals, and traits .

Citation Information

Patent Citations

  • Forest tree planting article selection prediction method and system based on big data

    CN119358821A

  • Method for rapidly and nondestructively detecting seed quality based on near infrared spectrum technology

    CN119595590A

  • Seed quality evaluation system and device based on multi-modal data fusion

    CN119715414A

  • Comprehensive evaluation method for quality traits of plant breeding

    CN119782885A

  • Intelligent crop seed vigor detection method and system

    CN119969003A

Cited By

  • High-resistance plant screening method and system based on ecological environment simulation

    CN120562718A

  • A method and system for screening highly resistant plants based on ecological environment simulation

    CN120562718B

  • Operation and maintenance management method and device, electronic equipment and storage medium

    CN120744791A

  • Cucumber root-knot nematode resistance grade identification method and system based on characteristic electric signals

    CN121499600A

  • Method and system for identifying resistance grade of cucumber to meloidogyne hapla based on characteristic electrical signal

    CN121499600B