Rape seed quality evaluation model based on big data
Through the big data-based rape seed quality evaluation model, combined with the graph neural network and multi-head attention mechanism, deep learning of the gene-environment interaction relationship, the problem of data integration and complex interaction relationship in rape seed quality evaluation is solved, efficient and accurate quality evaluation is achieved, and the new variety characteristics can be quickly adapted.
Patent Information
- Application Number
- CN202510678901.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
There are data integration problems in the evaluation of rapeseed seed quality, complex gene-environment interaction relationships, and it is difficult for traditional methods to fully integrate multi-dimensional feature information, resulting in inaccurate evaluation results and it is difficult for traditional models to quickly adapt to the characteristics of new varieties.
The rape seed quality evaluation model based on big data is adopted, including multi-source heterogeneous data acquisition and preprocessing, cross-domain feature interaction map construction, multi-modal feature fusion and optimization, feature evaluation analysis and transfer learning adaptation deployment, and other modules are used to deeply learn gene-environment interaction relationships through graph neural network and multi-head attention mechanism to generate and optimize comprehensive feature vectors, build quality evaluation models, and adapt to new varieties through transfer learning.
A comprehensive and accurate evaluation of rapeseed seed quality has been achieved, data utilization efficiency has been improved, the model's processing ability of complex data has been enhanced, and the new variety characteristics can be quickly adapted to the characteristics of the new variety, which has improved the evaluation efficiency.
Smart Images

Figure CN120196911A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of seed treatment analysis, and particularly relates to a rapeseed quality evaluation model based on big data. Background Art
[0002] There are still some challenges in the current rapeseed quality evaluation, including the following aspects: The rapeseed quality evaluation involves multi-source heterogeneous data such as genes, environment, and phenotypes. These data are widely sourced and in various formats, lacking effective integration means, resulting in low data utilization efficiency and difficulty in fully realizing the data value; The rapeseed quality is jointly affected by genes and the environment, and the gene-environment interaction relationship is complex. Traditional methods are difficult to deeply explore this complex interaction, leading to an insufficient understanding of the rapeseed quality formation mechanism; When evaluating rapeseed quality, multi-dimensional characteristics such as genes, environment, and phenotypes need to be considered, but traditional methods have deficiencies in feature fusion and are difficult to fully integrate the information of different-dimensional characteristics, resulting in inaccurate evaluation results; With the continuous update of rapeseed varieties, traditional quality evaluation models are difficult to quickly adapt to the feature space of new varieties, resulting in inaccurate quality evaluation of new varieties. Therefore, the present invention proposes a rapeseed quality evaluation model based on big data. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems in the background art and propose a rapeseed quality evaluation model based on big data.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A rapeseed quality evaluation model based on big data, comprising: Multi-source heterogeneous data acquisition and preprocessing module: Obtain multi-source heterogeneous rapeseed data, process and analyze the obtained multi-source heterogeneous rapeseed data, and generate a multi-source heterogeneous rapeseed data pool; Cross-domain feature interaction map construction module: Based on the multi-source heterogeneous rapeseed data pool, construct a cross-domain feature interaction map, use a graph neural network to perform deep feature learning on the constructed cross-domain feature interaction map, analyze the gene-environment interaction relationship, and output an updated cross-domain feature interaction map; Multi-modal feature fusion and optimization module: Analyze the cross-domain feature interaction map through a preset four-branch multi-head attention architecture, generate double weights and fuse features in combination with a cross-modal attention mechanism, and perform dual optimization through multi-head attention screening to generate an optimized comprehensive feature vector that integrates multi-dimensional information; Feature evaluation and analysis module: Decode and analyze the optimized comprehensive feature vector, thereby constructing a rapeseed quality evaluation model and extracting high-yield rapeseed variety data; Transfer learning adaptation and deployment module: Deploy a generative adversarial network using the extracted data of high-yield rapeseed varieties, and adapt the quality evaluation model to the new variety feature space through transfer learning to complete the intelligent evaluation of the quality of new rapeseed seeds.
[0005] Furthermore, the multi-source heterogeneous data collection and preprocessing module obtains multi-source heterogeneous rapeseed data and processes and analyzes the obtained multi-source heterogeneous rapeseed data. The process of generating a multi-source heterogeneous rapeseed data pool includes: Extract rapeseed whole-genome SNP marker data from the gene database; collect field phenotypic data through the Internet of Things sensor network; access the real-time monitoring data stream of the weather station to obtain field environmental parameters; use a drone equipped with a multispectral camera to obtain dynamic image data of the rapeseed growth cycle; Unify and store all the collected data sets in the HBase distributed database and perform preprocessing. Specifically: Remove duplicate sites and low-quality sites in the rapeseed whole-genome SNP marker data through data cleaning, retain the set of valid SNP sites to obtain a genotype data set; Parse the time series records of field environmental parameters and establish an environmental parameter time series index; Call the drone aerial image data, and use computer vision technology to extract features from the dynamic image data, segment and label the image timestamps according to the growth cycle, and extract the vegetation index as the initial phenotypic feature; At the same time, align the collected field phenotypic data and the extracted initial phenotypic features in space and time to eliminate the sampling time deviation, so as to obtain collaborative phenotypic features; Construct an SNP-chromosome association matrix based on the genotype data set as the gene dimension basic framework of the rapeseed original data pool; Map the environmental parameter time series index and collaborative phenotypic features to the gene dimension basic framework according to the timestamp to establish a cross-domain data association relationship; Adopt the column family storage model of the HBase distributed database, divide the data column families according to genotype, environment, and phenotype, and set the composite coding rule of the row key; Perform data integrity verification, verify the consistency of cross-domain data through the hash verification code, and complete the construction of the rapeseed original data pool; Perform standardization processing on the rapeseed original data pool to obtain a rapeseed standardized data pool.
[0006] Furthermore, the cross-domain feature interaction graph construction module constructs a cross-domain feature interaction graph based on the multi-source heterogeneous rapeseed data pool, and uses a graph neural network to perform deep feature learning on the constructed cross-domain feature interaction graph to analyze the gene-environment interaction relationship. The process of outputting the updated cross-domain feature interaction graph includes: Separate genotype nodes, environment nodes, and phenotype nodes from the rapeseed standardized data pool, and perform one-hot encoding on each node to generate an initial feature vector as the initial feature representation of the graph; After completing the initial feature representation of nodes, perform an association analysis on gene nodes and environmental nodes: calculate the mutual information value between gene nodes and environmental nodes, where the mutual information value is used to quantify the degree of association between these two node types; based on the calculated mutual information value, set a mutual information threshold for screening out significantly associated node pairs; construct a gene-environment binary edge set according to the screening process; Generate a static graph based on the constructed gene-environment binary edge set and the initial feature representation of nodes: aggregate adjacent node features through graph convolutional layers to enable it to capture the information of adjacent nodes around the node, and fuse the captured information into the feature representation of the target node, thereby generating a node embedding representation; finally, the output static graph includes a node embedding matrix and an adjacency matrix, where the node embedding matrix is used to record the feature representation of each node after feature aggregation, and the adjacency matrix is used to represent the connection relationship between nodes in the static graph; Input the generated static graph into a three-layer graph attention network and use a gated recurrent unit to capture the temporal changes of gene-environment edge weights; Define a gene-environment interaction intensity index, and the calculation method of this index is the edge weight multiplied by the cosine similarity of node embeddings. Among them, the edge weight represents the association strength between gene nodes and environmental nodes, and the cosine similarity of node embeddings is used to measure the similarity degree of two nodes in the feature space; based on the gene-environment interaction intensity index, screen the top N strongly associated edges with index values, where N is a preset positive integer used to limit the number of strongly associated edges selected from many gene-environment edges after sorting by the interaction intensity index value; Introduce a graph contrast learning framework to further analyze the top N strongly associated edges with index values: perform perturbation operations on the static graph, and the perturbation operations include randomly modifying node features and edge weights; calculate the change rate of interaction intensity by comparing the differences in node representations of the graph before and after perturbation; set a change rate threshold M to screen out the key feature edges whose interaction intensity change rate is greater than or equal to the change rate threshold M; integrate all the screened key feature edges into the static graph to form a final dynamic graph, that is, an updated cross-domain feature interaction graph.
[0007] Furthermore, the process of inputting the generated static graph into a three-layer graph attention network and using a gated recurrent unit to capture the temporal changes of gene-environment edge weights includes: Each layer of the graph attention network adopts a multi-head attention mechanism, and the multi-head attention mechanism learns the attention relationship between nodes from multiple different subspaces; in each layer, calculate the attention coefficient between nodes, where the attention coefficient is used to reflect the importance and association strength between nodes; according to the calculated attention coefficient, update the representation of the node to enable the node to fuse the information of its adjacent nodes; Taking the environmental parameter change rate as the gating signal, the gated recurrent unit dynamically corrects the weight value of the gene-environment edge according to the change of environmental parameters.
[0008] Further, the process of the multi-modal feature fusion and optimization module parsing the cross-domain feature interaction map through a preset four-branch multi-head attention architecture, generating double weights and fusing features in combination with the cross-modal attention mechanism, and generating an optimized comprehensive feature vector that integrates multi-dimensional information after being screened by multi-head attention includes: Inputting the constructed cross-domain feature interaction map into a preset four-branch multi-head attention sub-module: dividing the units of the four-branch multi-head attention sub-module, specifically including a gene focus head, an environment focus head, a phenotype focus head, and an interaction focus head, to process gene, environment, phenotype, and gene-environment interaction features respectively; After completing the unit division, internal calculations start independently within each focus head; meanwhile, a cross-modal attention mechanism is introduced between the gene and environment branches: calculating the cross-correlation strength between the gene feature and the environment feature through dot product operation; generating weighted cross features based on the calculated cross-correlation strength value; Concatenating the output results of the four branches, and the concatenated feature vector contains comprehensive information of genes, environment, phenotype, and gene-environment interaction; inputting the concatenated feature vector into a multi-layer perceptron, and the multi-layer perceptron further analyzes the input feature vector to finally generate a four-dimensional weight matrix; normalizing the four-dimensional weight matrix through the Softmax function to obtain the gene-environment interaction weight and the phenotype response weight; Performing a weighted summation operation on gene, environment, phenotype, and gene-environment interaction features based on the gene-environment interaction weight and the phenotype response weight to form an initial fusion vector; Performing a multi-head self-attention operation on the initial fusion vector: the multi-head self-attention mechanism captures the correlation relationship between elements in the initial fusion vector from multiple different subspaces, thereby calculating the global correlation estimation value; setting a global correlation threshold K, and screening out the core feature subset whose global correlation estimation value is greater than or equal to K by comparing the global correlation estimation value with the threshold K; Setting a residual connection sub-module, inputting the features of the cross-domain feature interaction map and the screened core feature subset into a convolutional neural network together to generate a residual correction term; performing non-negative matrix factorization operation on the residual correction term to extract latent factors as the constituent units of the optimized comprehensive feature vector; concatenating the extracted latent factors with the previously screened core features, and generating the final optimized comprehensive feature vector through attention scoring weighting.
[0009] Furthermore, the feature evaluation and analysis module decodes and analyzes the optimized comprehensive feature vector to construct a rapeseed quality evaluation model, and the process of extracting high-yield rapeseed variety data includes: Use a sparse autoencoder to achieve feature domain decoupling. Through L1 regularization constraints, force the activation values of the hidden layer to be sparse, and separate: gene feature sub-vectors, environmental feature sub-vectors, phenotypic feature sub-vectors, and interaction feature sub-vectors; among them, each sub-vector corresponds to an independent feature domain; Construct corresponding capsule processing units for each decoupled feature domain to form a parallelized feature domain processing pipeline, that is, determine the information transfer weights between sub-vectors through an iterative routing algorithm; at the same time, construct a corresponding processing pipeline for each feature domain; Based on the routing weights, perform weighted aggregation on the decoupled feature sub-vectors and transmit them layer by layer in the capsule network: First, in the local feature fusion stage, pair the capsules of adjacent feature domains, and refine the local interaction pattern through dynamic routing; Subsequently, in the global feature fusion stage, integrate all feature domain capsules into a unified high-level representation to form a progressive feature refinement path from local to global; According to the progressive feature refinement processing results, set a multi-scale feature fusion mechanism, and finally output fusion features including two levels, namely local feature fusion and global feature fusion; among them, local feature fusion: retain the detailed information of the interaction through the weighted combination of adjacent feature domain capsules; global feature fusion: aggregate all feature domain capsules into a unified global representation through top-level routing to form an overall description of rapeseed; Adopt a feature importance evaluation method to quantitatively analyze the local fusion features and global fusion features obtained under the multi-scale feature fusion mechanism, and determine the contribution degree of each feature domain and its combination to the rapeseed variety annotation; According to the quantitative analysis results, screen out the feature domains and their combinations that play a key role in the annotation of high-yield rapeseed varieties as the criteria for subsequent annotation and screening; finally generate the annotation results of each rapeseed variety.
[0010] Furthermore, the transfer learning adaptation and deployment module deploys a generative adversarial network using the extracted high-yield rapeseed variety data, and adapts the quality evaluation model to the new variety feature space through transfer learning. The process of completing the intelligent evaluation of the quality of new rapeseed seeds includes: Take the extracted high-yield rapeseed variety data as the source domain data, input it into the preset adversarial training sub-module for pre-training of the quality evaluation model, so that it learns the general feature representation related to rapeseed seed quality; adopt a transfer learning strategy to transfer the knowledge of the pre-trained model from the high-yield variety data set to the new variety rapeseed seed quality evaluation task; at the same time, adjust and optimize the parameters of the quality evaluation model according to the characteristics of the new variety data.
[0011] Compared with the prior art, the present invention has the following beneficial effects: by acquiring multi-source heterogeneous rapeseed data and constructing a data pool, data from different channels and structures are integrated, data silos are broken, and a comprehensive and rich data foundation is provided for subsequent analysis, which is helpful to mine more potential information; by constructing a cross-domain feature interaction map and using graph neural network learning, the interaction relationship between genes and environmental characteristics can be intuitively presented, the gene-environment interaction mechanism can be deeply mined, and a more accurate basis can be provided for subsequent analysis, and the output updated map also provides strong support for subsequent analysis; by parsing the map through a four-branch multi-head attention architecture and combining a cross-modal attention mechanism, the importance of features of different dimensions can be fully paid attention to, multi-dimensional information can be integrated and dual-optimized to generate optimized comprehensive feature vectors, the quality and representativeness of the features can be effectively improved, and the model's processing ability for complex data can be enhanced; by decoding the optimized feature vector to construct an evaluation model and extract high-yield variety data, it can be directly applied to rapeseed seed quality evaluation, and at the same time provide an important reference for subsequent new variety research, which is helpful to quickly screen out rapeseed varieties with high-yield potential; by deploying a generative adversarial network using high-yield data and adapting the model through transfer learning, the model can be quickly applied to new varieties, intelligent evaluation can be achieved, and evaluation efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a module diagram of a rapeseed seed quality assessment model based on big data proposed in the present invention. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the implementation regulations described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0014] Reference Figure 1 , a rapeseed seed quality evaluation model based on big data, including: Multi-source heterogeneous data acquisition and preprocessing module: obtain multi-source heterogeneous rapeseed data, process and analyze the obtained multi-source heterogeneous rapeseed data, and generate a multi-source heterogeneous rapeseed data pool; Cross-domain feature interaction map construction module: Based on the multi-source heterogeneous rapeseed data pool, a cross-domain feature interaction map is constructed. The graph neural network is used to perform deep feature learning on the constructed cross-domain feature interaction map, analyze the gene-environment interaction relationship, and output the updated cross-domain feature interaction map; Multi-modal Feature Fusion and Optimization Module: Analyze the cross-domain feature interaction map through a preset four-branch multi-head attention architecture, generate double weights and fuse features in combination with the cross-modal attention mechanism, and generate an optimized comprehensive feature vector that integrates multi-dimensional information after being screened by multi-head attention; Feature Evaluation and Analysis Module: Decode and analyze the optimized comprehensive feature vector to construct a rapeseed quality evaluation model and extract high-yield rapeseed variety data; Transfer Learning Adaptation and Deployment Module: Use the extracted high-yield rapeseed variety data to deploy a generative adversarial network, and adapt the quality evaluation model to the new variety feature space through transfer learning to complete the intelligent evaluation of the quality of new rapeseed seeds.
[0015] It should be further noted that in the specific implementation process, the process of the multi-source heterogeneous data acquisition and preprocessing module obtaining multi-source heterogeneous rapeseed data, processing and analyzing the obtained multi-source heterogeneous rapeseed data, and generating a multi-source heterogeneous rapeseed data pool is as follows: Extract rapeseed whole-genome SNP marker data from the gene database; collect field phenotypic data through the Internet of Things sensor network; access the real-time monitoring data stream of the meteorological station to obtain field environmental parameters; use a drone equipped with a multi-spectral camera to obtain dynamic image data of the rapeseed growth cycle; it can be understood that the phenotypic data collected by the field sensors includes the number of branches, yield, etc., and the field environmental parameters include environmental monitoring data such as soil humidity, temperature, and light intensity; Unify and store all the collected data sets in the HBase distributed database and perform preprocessing. Specifically: Remove duplicate sites and low-quality sites in the rapeseed whole-genome SNP marker data through data cleaning, retain the set of valid SNP sites to obtain a genotype data set; analyze the time series records of the field environmental parameters and establish a time series index of the environmental parameters; call the drone aerial image data, and use computer vision technology to extract features from the dynamic image data, segment and label the image timestamps according to the growth cycle, and extract the vegetation index as the preliminary phenotypic feature; at the same time, align the collected field phenotypic data and the extracted preliminary phenotypic features in space and time to eliminate the sampling time deviation, so as to obtain the collaborative phenotypic feature; Construct an SNP-chromosome association matrix based on the genotype data set as the gene dimension basic framework of the rapeseed original data pool; map the environmental parameter time series index and the collaborative phenotypic feature to the gene dimension basic framework according to the timestamp to establish a cross-domain data association relationship; adopt the column family storage model of the HBase distributed database, divide the data column families according to genotype, environment, and phenotype, and set the composite coding rule of the row key; perform data integrity verification, verify the consistency of the cross-domain data through the hash verification code, and complete the construction of the rapeseed original data pool; Standardize the original rapeseed data pool to obtain a standardized rapeseed data pool; specifically, perform clustering analysis on the original rapeseed data pool, use the DBSCAN algorithm to identify data outliers caused by extreme climates, and eliminate abnormal rapeseed data samples based on the local density threshold; for missing measurement values, use the time series interpolation algorithm to fill in the missing measurement values; perform Z-score standardization on the filled data to eliminate the dimensional difference and generate a standardized data pool that conforms to the normal distribution.
[0016] It should be further noted that in the specific implementation process, the process of the cross-domain feature interaction map construction module constructing a cross-domain feature interaction map based on the multi-source heterogeneous rapeseed data pool, using the graph neural network to perform deep feature learning on the constructed cross-domain feature interaction map, analyzing the gene-environment interaction relationship, and outputting the updated cross-domain feature interaction map is as follows: Separate the genotype nodes (SNP locus sequences), environmental nodes (time series data such as soil moisture and temperature), and phenotypic nodes (time series features such as NDVI and plant height) from the rapeseed standardized data pool, and perform one-hot encoding on each node to generate an initial feature vector as the initial feature representation of the map; for example, obtain the genotype, environment, and phenotype data of 1000 samples from the rapeseed standardized data pool. The genotype data contains 50 gene locus information, the environment data includes 10 environmental parameters such as temperature, humidity, and light, and the phenotype data has 8 phenotypic indicators such as plant height and yield; the genotype node has 50 possible states, so it is encoded as a 50-dimensional vector; the environmental node is 10-dimensional, and the phenotypic node is 8-dimensional to generate the initial feature vector; After completing the initial feature representation of the nodes, perform correlation analysis on the gene nodes and environmental nodes: calculate the mutual information value between the gene nodes and environmental nodes , where the mutual information value is used to quantify the degree of association between these two node types, and the formula is: ; In the formula, is the gene node, is the set of gene node states, is the environmental node, is the set of environmental node states, is the joint probability, is the marginal probability; based on the calculated mutual information value (assumed to be 0.3), set the mutual information threshold (such as 0.5) to screen out significantly associated node pairs; construct a gene-environment binary edge set according to the screening process (a total of 200 edges are obtained); Based on the constructed gene-environment binary edge set and the initial feature representation of nodes, a static graph is generated: The adjacent node features are aggregated through graph convolutional layers to enable it to capture the information of adjacent nodes around the node, and the captured information is fused into the feature representation of the target node, thereby generating a node embedding representation; It can be understood that when constructing the static graph, graph convolutional layers are used to aggregate adjacent node features, and this process involves information interaction among multiple nodes in the graph structure; In each step of graph convolution, a node is selected as the object to be processed currently, and this node is the target node; For example, assume that the gene-environment interaction graph contains gene nodes j1, j2, j3 and environment nodes e1, e2; When performing graph convolution operations on the gene node j1, the gene node j1 is the target node. If the gene node j1 is connected to the environment nodes e1 and e2, the graph convolutional layer first captures the feature information of the environment nodes e1 and e2, and then fuses this feature information into the feature representation of the gene node j1. The updated feature of the gene node j1 contains its own initial feature and the feature information of the environment nodes e1 and e2, thereby more accurately describing the interaction between the gene node j1 and the environment; Finally, the output static graph contains a node embedding matrix and an adjacency matrix, where the node embedding matrix is used to record the feature representation of each node after feature aggregation, and the adjacency matrix is used to represent the connection relationship between nodes in the static graph; The generated static graph is input into a three-layer graph attention network, and a gated recurrent unit is used to capture the temporal changes of gene-environment edge weights: Each layer of the graph attention network adopts a multi-head attention mechanism (such as an 8-head attention mechanism). The multi-head attention mechanism learns the attention relationship between nodes from multiple different subspaces, thereby more comprehensively capturing the interaction information between nodes; In each layer, the attention coefficient between nodes is calculated, and the attention coefficient is used to reflect the importance and correlation strength between nodes; According to the calculated attention coefficient, the representation of the node is updated to enable the node to fuse the information of its adjacent nodes; Using the environmental parameter change rate as a gating signal, the gated recurrent unit dynamically corrects the weight value of the gene-environment edge according to the change of the environmental parameter, so that the edge weight in the graph is updated in real time as the environment changes, thereby more accurately describing the dynamic interaction relationship between genes and the environment; Define the gene-environment interaction strength index , and the calculation method of this index is the edge weight multiplied by the cosine similarity of node embeddings , that is , where the edge weights represent the association strength between gene nodes and environmental nodes, and the cosine similarity of node embeddings is used to measure the similarity between two nodes in the feature space. The product of the two comprehensively reflects the strength of gene-environment interaction. Based on the gene-environment interaction strength index, the top N (such as 50) strongly associated edges with index values are selected. Here, N is a preset positive integer used to limit the number of strongly associated edges selected from numerous gene-environment edges sorted by the interaction strength index value, so as to focus on the edges that play a key role in gene-environment interaction and lay a quantitative range foundation for the subsequent accurate screening of key feature edges. Introduce a graph contrast learning framework to further analyze the top N strongly associated edges with index values: perform perturbation operations on the static graph, and the perturbation operations include randomly modifying node features and edge weights. Calculate the change rate of interaction strength by comparing the node representations of the graph before and after perturbation. : ; Set a change rate threshold M, and screen for key feature edges with an interaction strength change rate greater than or equal to the change rate threshold M. Integrate all the screened key feature edges into the static graph to form a final dynamic graph, that is, the updated cross-domain feature interaction graph. For example, the temperature changes by 10% within a month, and the humidity changes by 5%. Dynamically correct the weight values of gene-environment edges according to these change rates. Define the gene-environment interaction strength index, and screen the top 50 strongly associated edges with index values. Perform perturbation operations on the static graph, such as randomly modifying 5% of the node features and 3% of the edge weights. Calculate the change rate of interaction strength by comparing the node representations of the graph before and after perturbation. Set the change rate threshold M to 0.1, and screen out 30 key feature edges with a change rate greater than or equal to 0.1, and integrate them into the static graph to form a final dynamic graph.
[0017] It should be further noted that in the specific implementation process, the process of the multi-modal feature fusion and optimization module parsing the cross-domain feature interaction graph through a preset four-branch multi-head attention architecture, generating double weights and fusing features through a cross-modal attention mechanism, and double-optimizing to generate an optimized comprehensive feature vector integrating multi-dimensional information is as follows: Input the constructed cross-domain feature interaction graph into a preset four-branch multi-head attention sub-module: divide the units of the four-branch multi-head attention sub-module, specifically including a gene focus head, an environment focus head, a phenotype focus head, and an interaction focus head, to process gene, environment, phenotype, and gene-environment interaction features respectively. After the unit division is completed, internal calculations start independently within each focus head, including calculation queries, keys, and value matrices. At the same time, a cross-modal attention mechanism is introduced between the gene and environment branches: the dot product operation is performed on the gene features and environment features to calculate the cross-correlation strength between the two. Based on the calculated cross-correlation strength value, weighted cross features are generated. The output results of the four branches (gene focus head, environment focus head, phenotype focus head, interaction focus head) are concatenated. The concatenated feature vector contains comprehensive information on genes, environment, phenotype, and gene-environment interaction. The concatenated feature vector is input into a multi-layer perceptron, which further analyzes the input feature vector (non-linear transformation and feature extraction), and finally generates a four-dimensional weight matrix. The four-dimensional weight matrix is normalized through the Softmax function to obtain the gene-environment interaction weight and the phenotype response weight. Among them, the gene-environment interaction weight is used to reflect the importance of the interaction between genes and the environment, and the phenotype response weight reflects the sensitivity of the phenotype to changes in genes and the environment. Based on the gene-environment interaction weight and the phenotype response weight, weighted summation operations are performed on the gene, environment, phenotype, and gene-environment interaction features to form an initial fusion vector. Perform multi-head self-attention operations on the initial fusion vector: the multi-head self-attention mechanism captures the correlation relationships between elements in the initial fusion vector from multiple different subspaces, thereby calculating the global correlation estimation value. The global correlation estimation value reflects the importance degree of each element in the entire vector. Set the global correlation threshold K (such as 0.8). By comparing the global correlation estimation value with the threshold K, a core feature subset with a global correlation estimation value greater than or equal to K is selected. The core feature subset contains the most critical and important information in the initial fusion vector. A residual connection sub-module is set up. The features of the cross-domain feature interaction map and the selected core feature subset are input into a convolutional neural network together to generate a residual correction term. Among them, the features of the cross-domain feature interaction map include the feature information of genes, environment, phenotypes, and gene-environment interactions. Perform non-negative matrix factorization on the residual correction term: decompose the high-dimensional residual correction term into a low-dimensional latent factor matrix and a coefficient matrix, and use the latent factor as the latent feature pattern in the residual correction term. It can be understood that in the gene feature dimension, the latent factor corresponds to the core components of the gene expression pattern in the gene regulatory network; for environmental features, the latent factor reflects the key environmental factor combinations affecting rapeseed traits or the potential driving factors of environmental changes; in terms of phenotypic features, the latent factor represents the potential sources of phenotypic variation or phenotypic classification; for gene-environment interaction features, the latent factor reflects the potential mechanism of gene-environment interaction. Extract the latent factor as a component unit for optimizing the comprehensive feature vector. Concatenate the extracted latent factor with the previously selected core features, and generate the final optimized comprehensive feature vector through attention score weighting.
[0018] It should be further noted that in the specific implementation process, the process of the feature evaluation and analysis module decoding and analyzing the optimized comprehensive feature vector to construct a rapeseed seed quality evaluation model and extract high-yield rapeseed variety data is as follows: The construction uses a sparse autoencoder to achieve feature domain decoupling. Through L1 regularization constraints, the activation values of the hidden layer are forced to be sparse, and the following are separated: gene feature sub-vectors and environmental feature sub-vectors and phenotypic feature sub-vectors as well as interaction feature sub-vectors . Among them, \(R\) represents the real number space (i.e., the eigenvalues are composed of specific numerical values), and \(d\) represents the dimension of the corresponding feature vector (i.e., the number of parameters describing gene, environment, phenotype, or interaction features). Each sub-vector corresponds to an independent feature domain, representing a set of features with clear semantics in a certain category in the original feature space. Construct corresponding capsule processing units for each decoupled feature domain (gene / environment / phenotype / interaction) to form a parallel feature domain processing pipeline, that is, determine the information transfer weights between sub-vectors through an iterative routing algorithm. At the same time, construct a corresponding processing pipeline for each feature domain: set the initial routing coefficient, and adjust the routing coefficient according to the prediction consistency between the lower-layer capsules and the higher-layer capsules. At the same time, construct a corresponding processing pipeline for each feature domain. Specifically: obtain the gene pipeline by initializing the 1D-CNN convolution kernel; obtain the environment pipeline by loading the pre-trained ResNet-18 weights to process the spatio-temporal sequence; obtain the phenotype pipeline by constructing a knowledge distillation sub-module and loading the prior knowledge of plant phenomics; by initializing the bilinear fusion matrix Gene-environment interaction pipeline acquisition; Weighted aggregation of the decoupled feature sub-vectors is performed based on routing weights and passed layer by layer in the capsule network: First, in the local feature fusion stage, capsules in adjacent feature domains (such as gene-environment, environment-phenotype) are combined pairwise, and the local interaction pattern is refined through dynamic routing; Subsequently, in the global feature fusion stage, all feature domain capsules are integrated into a unified high-level representation, forming a progressive feature refinement path from local to global, specifically: For gene feature refinement, the input gene feature sub-vector G is fed into a 1D-CNN, and the primary gene pattern is output , sequence annotation is performed through the CRF layer to identify disease-resistant related SNP combination patterns and generate a gene quality score , where is the activation function, is the weight matrix of gene features, is the bias term of gene features; For environmental feature refinement, E is input into the spatio-temporal Transformer, decomposed into static attributes Es (such as soil type) and dynamic time series Ed (precipitation / temperature sequence) through the separable attention mechanism, and an environmental adaptability score is generated , where ME is a multi-layer decision processor (mining the laws of static environmental features through a multi-layer neural network), LE is a time series memory sub-module (used to process environmental dynamic data such as temperature / precipitation that changes over time), and ⊕ is a feature fusion operation (merging and analyzing static attributes and dynamic time series information); For phenotype feature refinement, a cross-layer graph structure is constructed, with phenotype features as nodes and genetic correlations between phenotypes as edges. Information is propagated through the GAT network to update the phenotype representation P' and generate a phenotype integrity score , where is the attention weight, represents different phenotype feature dimensions (such as specific trait indicators such as plant height, number of siliques, 1000-grain weight, etc.); For interaction feature refinement, bilinear pooling is performed: , and an interaction intensity score is generated , where tanh is a bidirectional activation function, represents matrix transpose, is the interaction feature weight matrix, is the reference correction parameter for interaction feature calculation (i.e., the interaction feature bias term), is the interaction feature matrix after bilinear pooling (containing multi-dimensional interaction information between genes and the environment), and F is the norm in the interaction intensity score; Based on the refined results of progressive features, a multi-scale feature fusion mechanism is set, and finally, fused features at two levels are output, namely local feature fusion and global feature fusion. Among them, local feature fusion: By weighted combination of adjacent feature domain capsules (such as gene-environment fusion capsules), detailed information of interactions is retained. Global feature fusion: All feature domain capsules are aggregated into a unified global representation through top-level routing to form an overall description of rapeseed. Using a feature importance evaluation method (such as calculating feature importance based on a tree model), quantitative analysis is carried out on the local fusion features and global fusion features obtained under the multi-scale feature fusion mechanism to determine the contribution degree of each feature domain and its combination to the rapeseed variety annotation. According to the results of quantitative analysis, feature domains and their combinations that play a key role in the annotation of high-yield rapeseed varieties are selected (such as certain gene locus combinations in the gene feature sub-vector and combinations of temperature and humidity in the environmental feature sub-vector, etc.) as the criteria for subsequent annotation and screening, and at the same time, features with less contribution or irrelevance are removed. For example, 200 high-yield rapeseed variety data that meet these key features are extracted from the data, and finally, the annotation results of each rapeseed variety are generated.
[0019] It should be further noted that in the specific implementation process, the process of the transfer learning adaptation and deployment module using the extracted high-yield rapeseed variety data to deploy a generative adversarial network and adapting the quality evaluation model to the new variety feature space through transfer learning to complete the intelligent evaluation of the quality of new rapeseed seeds is as follows: Taking the extracted high-yield rapeseed variety data as the source domain data, input it into the preset adversarial training sub-module for pre-training of the quality evaluation model to enable it to learn the general feature representation related to the quality of rapeseed seeds. Adopt a transfer learning strategy to transfer the knowledge of the pre-trained model from the high-yield variety dataset to the quality evaluation task of new rapeseed seeds. At the same time, adjust and optimize the parameters of the quality evaluation model according to the characteristics of the new variety data to ensure that the quality evaluation model can accurately evaluate the quality of rapeseed seeds on the new variety data.
[0020] It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0021] It should be understood that determining B according to A does not mean determining B only according to A, but also B can be determined according to A and / or other information.
[0022] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.
[0023] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A rapeseed quality evaluation model based on big data, characterized in that: Multi-source heterogeneous data acquisition and preprocessing module: Obtain multi-source heterogeneous rapeseed data, process and analyze the obtained multi-source heterogeneous rapeseed data, and generate a multi-source heterogeneous rapeseed data pool; Cross-domain feature interaction map construction module: Based on the multi-source heterogeneous rapeseed data pool, construct a cross-domain feature interaction map, use a graph neural network to perform deep feature learning on the constructed cross-domain feature interaction map, analyze the gene-environment interaction relationship, and output an updated cross-domain feature interaction map; Multi-modal feature fusion and optimization module: Parse the cross-domain feature interaction map through a preset four-branch multi-head attention architecture, generate double weights and fuse features in combination with a cross-modal attention mechanism, and generate an optimized comprehensive feature vector that integrates multi-dimensional information after being screened by multi-head attention; Feature evaluation and analysis module: Decode and analyze the optimized comprehensive feature vector, thereby constructing a rapeseed quality evaluation model and extracting high-yield rapeseed variety data; Transfer learning adaptation and deployment module: Use the extracted high-yield rapeseed variety data to deploy a generative adversarial network, and adapt the quality evaluation model to the new variety feature space through transfer learning to complete the intelligent evaluation of the quality of new rapeseeds.
2. The rapeseed seed quality evaluation model based on big data according to claim 1, wherein: The process of the multi-source heterogeneous data acquisition and preprocessing module obtaining multi-source heterogeneous rapeseed data, processing and analyzing the obtained multi-source heterogeneous rapeseed data, and generating a multi-source heterogeneous rapeseed data pool includes: Extract rapeseed whole-genome SNP marker data from the gene database; Collect field phenotype data through the Internet of Things sensor network; Access the real-time monitoring data stream of the weather station to obtain field environmental parameters; Use a drone equipped with a multi-spectral camera to obtain dynamic image data of the rapeseed growth cycle; Uniformly store all the collected data sets in the HBase distributed database and perform preprocessing. Specifically: Remove duplicate sites and low-quality sites in the rapeseed whole-genome SNP marker data through data cleaning, retain the set of valid SNP sites to obtain a genotype data set; Parse the time series record of the field environmental parameters and establish an environmental parameter time series index; Call the drone aerial image data, and use computer vision technology to extract features from the dynamic image data, segment and label the image timestamps according to the growth cycle, and extract the vegetation index as the preliminary phenotype feature; At the same time, align the collected field phenotype data and the extracted preliminary phenotype features in space and time to eliminate the sampling time deviation, thereby obtaining a collaborative phenotype feature; Construct an SNP-chromosome association matrix based on the genotype data set as the gene dimension basic framework of the rapeseed original data pool; Map the environmental parameter time series index and the collaborative phenotype feature to the gene dimension basic framework according to the timestamp to establish a cross-domain data association relationship; Adopt the column family storage model of the HBase distributed database, divide the data column families according to genotype, environment, and phenotype, and set the composite coding rule of the row key; Execute data integrity verification, verify the consistency of the cross-domain data through the hash verification code, and complete the construction of the rapeseed original data pool; Perform standardization processing on the rapeseed original data pool to obtain a rapeseed standardized data pool.
3. The rapeseed seed quality evaluation model based on big data according to claim 2, characterized in that: The process of the cross - domain feature interaction graph construction module constructing a cross - domain feature interaction graph based on a multi - source heterogeneous rapeseed data pool, performing deep feature learning on the constructed cross - domain feature interaction graph using a graph neural network, analyzing gene - environment interaction relationships, and outputting the updated cross - domain feature interaction graph includes: Separating genotype nodes, environment nodes, and phenotype nodes from the rapeseed standardized data pool, and performing one - hot encoding on each node to generate an initial feature vector as the initial feature representation of the graph; After completing the initial feature representation of the nodes, performing an association analysis between gene nodes and environment nodes: calculating the mutual information value between gene nodes and environment nodes, where the mutual information value is used to quantify the degree of association between these two node types; based on the calculated mutual information value, setting a mutual information threshold for screening out significantly associated node pairs; constructing a gene - environment binary edge set according to the screening process; Generating a static graph based on the constructed gene - environment binary edge set and the initial feature representation of the nodes: aggregating adjacent node features through a graph convolutional layer to enable it to capture the information of adjacent nodes around the node, and fusing the captured information into the feature representation of the target node, thereby generating a node embedding representation; finally, the output static graph includes a node embedding matrix and an adjacency matrix, where the node embedding matrix is used to record the feature representation of each node after feature aggregation, and the adjacency matrix is used to represent the connection relationship between nodes in the static graph; Inputting the generated static graph into a three - layer graph attention network and using a gated recurrent unit to capture the temporal changes of gene - environment edge weights; Defining a gene - environment interaction intensity index, the calculation method of which is the edge weight multiplied by the cosine similarity of node embeddings, where the edge weight represents the association strength between gene nodes and environment nodes, and the cosine similarity of node embeddings is used to measure the similarity degree of two nodes in the feature space; based on the gene - environment interaction intensity index, screening the top N strongly associated edges by index value, where N is a preset positive integer used to limit the number of strongly associated edges selected according to the interaction intensity index value from among many gene - environment edges; Introducing a graph contrast learning framework to further analyze the top N strongly associated edges by index value: performing perturbation operations on the static graph, and the perturbation operations include randomly modifying node features and edge weights; calculating the change rate of interaction intensity by comparing the node representation differences of the graph before and after perturbation; setting a change rate threshold M, and screening out key feature edges with an interaction intensity change rate greater than or equal to the change rate threshold M; integrating all the screened key feature edges into the static graph to form a final dynamic graph, that is, the updated cross - domain feature interaction graph.
4. The rapeseed seed quality evaluation model based on big data according to claim 3, characterized in that: The process of inputting the generated static graph into a three - layer graph attention network and using a gated recurrent unit to capture the temporal changes of gene - environment edge weights includes: Each layer of the graph attention network adopts the multi-head attention mechanism, which learns the attention relationships between nodes from multiple different subspaces; in each layer, the attention coefficients between nodes are calculated, where the attention coefficients are used to reflect the importance and association strength between nodes; according to the calculated attention coefficients, the representation of the nodes is updated to enable the nodes to fuse the information of their adjacent nodes. Taking the environmental parameter change rate as the gating signal, the gated recurrent unit dynamically corrects the weight value of the gene-environment edge according to the change of the environmental parameters.
5. The rapeseed quality evaluation model based on big data according to claim 1, characterized in that: The process of the multi-modal feature fusion and optimization module parsing the cross-domain feature interaction graph through a preset four-branch multi-head attention architecture, generating double weights and fusing features in combination with the cross-modal attention mechanism, and finally generating an optimized comprehensive feature vector that integrates multi-dimensional information after being screened by the multi-head attention includes: Input the constructed cross-domain feature interaction graph into a preset four-branch multi-head attention sub-module: perform unit division on the four-branch multi-head attention sub-module, specifically including the gene focus head, the environment focus head, the phenotype focus head, and the interaction focus head, to process gene, environment, phenotype, and gene-environment interaction features respectively. After completing the unit division, internal calculations start independently within each focus head; at the same time, a cross-modal attention mechanism is introduced between the gene and environment branches: through the dot product operation of the gene feature and the environment feature, calculate the cross-correlation strength between the two; based on the calculated cross-correlation strength value, generate weighted cross features. Concatenate the output results of the four branches. The concatenated feature vector contains comprehensive information of genes, environment, phenotype, and gene-environment interaction; input the concatenated feature vector into a multi-layer perceptron. The multi-layer perceptron further analyzes the input feature vector and finally generates a four-dimensional weight matrix; perform normalization processing on the four-dimensional weight matrix through the Softmax function to obtain the gene-environment interaction weight and the phenotype response weight. Perform a weighted summation operation on the gene, environment, phenotype, and gene-environment interaction features based on the gene-environment interaction weight and the phenotype response weight to form an initial fusion vector. Perform a multi-head self-attention operation on the initial fusion vector: the multi-head self-attention mechanism captures the association relationships between elements in the initial fusion vector from multiple different subspaces, thereby calculating the global association estimation value; set a global association threshold K, and by comparing the global association estimation value with the threshold K, screen out the core feature subset whose global association estimation value is greater than or equal to K. Set up a residual connection sub-module, input the features of the cross-domain feature interaction graph and the screened core feature subset into a convolutional neural network together to generate a residual correction term; perform non-negative matrix factorization operation on the residual correction term to extract latent factors as the constituent units of the optimized comprehensive feature vector; concatenate the extracted latent factors with the previously screened core features, and generate the final optimized comprehensive feature vector through attention scoring weighting.
6. The rapeseed seed quality evaluation model based on big data according to claim 1, characterized in that: Feature evaluation The process of the analysis module decoding and analyzing the optimized comprehensive feature vector to construct a rapeseed quality evaluation model and extract high-yield rapeseed variety data includes: Feature domain decoupling is achieved by using a sparse autoencoder. Through L1 regularization constraints, the activation values of the hidden layer are forced to be sparse, separating out: gene feature sub-vectors, environmental feature sub-vectors, phenotypic feature sub-vectors, and interaction feature sub-vectors; where each sub-vector corresponds to an independent feature domain. A corresponding capsule processing unit is constructed for each decoupled feature domain to form a parallel feature domain processing pipeline, that is, the information transfer weights between sub-vectors are determined by an iterative routing algorithm; at the same time, a corresponding processing pipeline is constructed for each feature domain. Based on the routing weights, the decoupled feature sub-vectors are weighted and aggregated and passed layer by layer in the capsule network: First, in the local feature fusion stage, capsules in adjacent feature domains are combined pairwise, and the local interaction pattern is refined through dynamic routing; subsequently, in the global feature fusion stage, all feature domain capsules are integrated into a unified high-level representation, forming a progressive feature refinement path from local to global. According to the progressive feature refinement processing results, a multi-scale feature fusion mechanism is set, and finally, fusion features at two levels are output, namely local feature fusion and global feature fusion; among them, local feature fusion: through the weighted combination of capsules in adjacent feature domains, the detailed information of the interaction is retained; global feature fusion: all feature domain capsules are aggregated into a unified global representation through top-level routing, forming an overall description of rapeseed. A feature importance evaluation method is used to quantitatively analyze the local fusion features and global fusion features obtained under the multi-scale feature fusion mechanism, and determine the contribution degree of each feature domain and its combination to the rapeseed variety annotation. According to the quantitative analysis results, the feature domains and their combinations that play a key role in the high-yield rapeseed variety annotation are selected as the criteria for subsequent annotation and screening; finally, the annotation results of each rapeseed variety are generated.
7. The rapeseed seed quality evaluation model based on big data according to claim 1, wherein: The migration learning adaptation and deployment module uses the extracted high-yield rapeseed variety data to deploy a generative adversarial network, and adapts the quality evaluation model to the new variety feature space through migration learning. The process of completing the intelligent evaluation of the quality of new rapeseed seeds includes: Taking the extracted high-yield rapeseed variety data as the source domain data, inputting it into a preset adversarial training sub-module for pre-training of the quality evaluation model, so that it learns the general feature representation related to rapeseed seed quality; adopting a migration learning strategy to transfer the knowledge of the pre-trained model from the high-yield variety dataset to the new variety rapeseed seed quality evaluation task; at the same time, adjusting and optimizing the parameters of the quality evaluation model according to the characteristics of the new variety data.
Citation Information
Patent Citations
Aquatic organism germplasm resource evaluation method and system based on deep learning
CN118155721A
Deep learning network-based corn seed quality evaluation method
CN119048812A
Ramie seed vigor assessment method and system based on multispectral image analysis
CN119131604A
Method, device and equipment for detecting seed quality and storage medium
CN119540231A
System and method for leveraging neural network based hybrid feature extraction model for grain quality analysis
WO2023084543A1
Cited By
Feed processing prediction and real-time regulation and control method based on deep learning
CN120494211A
Rape variety environmental adaptability evaluation system based on machine learning
CN120509606A
Intelligent comprehensive management method and system for germplasm resources of Chinese orchid
CN120744104A
Intelligent breeding optimization method and system based on artificial intelligence
CN120808903A
An intelligent planting optimization method and system based on artificial intelligence
CN120808903B