Flaxseed variety screening method based on comprehensive data analysis
By constructing a multidimensional dataset and performing feature recombination and adaptive weighted scoring, flaxseed varieties that produce stably under varying environments were identified. This solved the problem of insufficient cross-environment adaptability in existing screening methods, and enabled efficient and accurate variety screening and promotion.
Patent Information
- Application Number
- CN202510933371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-31
AI Technical Summary
Existing flaxseed variety screening methods cannot accurately identify varieties with stable yield and quality indicators under variable environments or extreme climatic conditions. They lack the ability to integrate multidimensional characteristics and predict cross-environmental adaptability, making it difficult to meet the evaluation needs of multi-trait coupling and multi-environmental synergy.
By collecting multidimensional datasets, performing multi-scale feature recombination and variable enhancement, constructing high-dimensional feature vectors, establishing an adaptive weighted scoring model, and combining a graph structure attention enhancement module and a robust dynamic parameter tuning network, we can conduct multiquantile distribution analysis and multi-objective decision-making to identify the flaxseed variety with the best overall performance.
It enables highly accurate trait performance and adaptability analysis of flaxseed varieties under different growth environments, improving breeding selection efficiency and promotion success rate, and possessing cross-environmental robustness and multi-scenario adaptability.
Smart Images

Figure CN120873643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of crop variety screening technology, specifically to a method for screening flaxseed varieties based on comprehensive data analysis. Background Technology
[0002] Flax, a crop with both economic and functional nutritional value, exhibits significant differences in seed quality among different varieties and is highly sensitive to environmental conditions. Currently, flaxseed variety selection primarily relies on plot trials and field observations, obtaining results through manual scoring or single-index statistical analysis. While these methods can reflect varietal characteristics to some extent, they cannot accurately identify varieties that maintain stable yields and high-quality indicators under variable environments or extreme climatic conditions. Especially given the increasingly diversified goals of target breeding, traditional methods are insufficient to meet the evaluation needs of flaxseed varieties that involve multi-trait coupling, multi-environmental synergy, and multi-scenario applications.
[0003] Existing screening techniques typically employ horizontal index comparisons to evaluate a small number of samples under a single environment, neglecting the coupling relationships and response patterns among varieties in a multidimensional feature space. Furthermore, they lack in-depth modeling of the nonlinear, multi-scale, and inter-variable dependencies within the data, making it difficult to accurately identify varieties with potential advantages but exhibiting "atypical" behavior. In addition, most methods fail to integrate molecular, phenotypic, and ecological data for modeling, lacking cross-regional adaptive prediction capabilities.
[0004] Therefore, there is an urgent need for a systematic method that combines multi-environmental test data, phenotypic performance data, and environmental variable data, and introduces a multi-layer feature fusion and weighted discrimination mechanism to conduct a global, objective, and multi-dimensional dynamic evaluation of flaxseed varieties. This method would enable rapid and intelligent screening of superior flaxseed varieties, thereby significantly improving screening efficiency and expanding their adaptability and promotion capabilities. Summary of the Invention
[0005] The purpose of this invention is to provide a method for screening flaxseed varieties based on comprehensive data analysis, in order to address the shortcomings of the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a flaxseed variety screening method based on comprehensive data analysis, comprising:
[0007] S100. Collect phenotypic characteristic data, grain physicochemical index data and environmental meteorological data of multiple flaxseed varieties under multiple growth environments, and construct a multidimensional dataset D;
[0008] S200. Based on the multidimensional dataset D, perform multi-scale feature recombination and variable enhancement to construct a high-dimensional feature vector T of flaxseed varieties under various environments. The feature vector is generated by a feature compression mechanism with a principal component retention rate of not less than 90%.
[0009] S300. Based on the feature vector T of each variety and the corresponding growth environment feature distribution, an adaptive weighted scoring model W is established, in which the scoring weight coefficient is adjusted according to the performance volatility and stability of the variety under extreme environment, and includes a phenotypic robustness discrimination submodule.
[0010] S400. Using model W, dynamic scoring is performed on each flaxseed variety to obtain a variety score set S. Multiquantile distribution analysis is performed on the score data in S to identify the candidate variety set C with the best fit in different target application scenarios.
[0011] S500 outputs the flaxseed variety with the best overall performance in candidate variety set C as the target screening result.
[0012] Preferably, S100 specifically includes: the phenotypic characteristic data includes the dynamic change curve of the growth cycle, leaf area index and biomass change rate; the physicochemical index data includes α-linolenic acid content, lignan accumulation rate and grain water metabolism parameters; and the environmental meteorological data includes temperature, sunshine, precipitation frequency and climate anomaly factors.
[0013] Preferably, S200 specifically includes:
[0014] The structure of different types of data in the multidimensional dataset D is defined, and the scale levels are divided according to the consistency of time series and the correlation of variables. Phenotypic-physicochemical dual-channel feature paths and meteorological influencing factor coupling channels are constructed respectively.
[0015] Nonlinear mapping and coupling transformation are performed on variables at different scales to output candidate feature groups, and the first round of candidate feature matrix T′ is generated based on the importance score of the variables;
[0016] The feature matrix T′ is subjected to dimension-wise principal component analysis to extract the principal component vectors with a cumulative contribution rate of not less than 90%, forming a high-dimensional compressed feature vector T, while simultaneously retaining the cross-scale collaborative change trend factors;
[0017] The high-dimensional feature vector T is mapped into the variety feature space and used as the input feature set for the subsequent scoring model.
[0018] Preferably, S300 specifically includes:
[0019] Based on the high-dimensional feature vector T corresponding to each flaxseed variety and the feature distribution density of its growth environment, an environment-specific response factor matrix E is constructed to characterize the adaptability surface of the variety to different environmental variables.
[0020] The stability of each principal component in T under typical and extreme environments is simulated by interval perturbation, the performance volatility index V is calculated, and this index is used as a dynamic adjustment factor for the scoring weight.
[0021] A multi-core fusion scoring model W with a phenotypic robust discrimination submodule is constructed, in which the submodule adopts a bidirectional feature stability filter to identify and suppress phenotypic outliers and short-period fluctuation signals;
[0022] By combining E, V, and robust output parameters, the weights of each scoring channel in model W are iteratively updated using a Bayesian optimization algorithm to form a weighted scoring structure with adaptive adjustment capabilities, thereby achieving accurate evaluation of flaxseed varieties under different climatic conditions.
[0023] Preferably, the scoring model W further includes a graph-structured attention enhancement module GAM-R and a robust two-layer dynamically tuned network DRN, specifically including:
[0024] The feature vectors T among flaxseed varieties are constructed into a weighted graph structure using Euclidean distance and environmental similarity index. Nodes represent the features of each variety, and edge weights reflect the degree of environmental interaction. An improved graph attention mechanism is introduced to redistribute the weights of different varieties in heterogeneous environments to generate an attention-enhanced feature matrix A_T.
[0025] A_T is input into a robust two-layer dynamic parameter tuning network DRN. The first layer is used to extract the nonlinear mismatch pattern between phenotype and environmental variables. The second layer uses the error propagation path as the adjustment benchmark to dynamically adjust the weight coefficients in W and optimize the robustness of the scoring model.
[0026] Finally, the optimized parameter set output by DRN is applied to the scoring model W.
[0027] Preferably, S400 specifically includes:
[0028] The scoring model W is used to dynamically score each flaxseed variety under different growth environments, generating a scoring time series matrix S;
[0029] A multiquantile distribution fitting analysis was performed on the scoring matrix S to construct a quantile-environmental factor response model Q, which was used to identify the stable performance range of each variety under a set environmental combination.
[0030] Based on the output of the Q model, a variety-scenario matching function F is designed. The function F uses the median score, the upper and lower quartiles, and the target application environment weight factor as parameters to calculate the variety suitability score.
[0031] Based on the fitness scores, the set of varieties C with the best overall performance under the target environment and target requirements is selected from the scoring matrix S, forming a set of highly adaptable candidate varieties.
[0032] Preferably, further improvements to the adaptability of candidate variety set C include:
[0033] An improved density-sensitive clustering algorithm, DS-DBC, is applied to a subset of samples with a highly stable rating structure in the rating matrix S. The algorithm constructs a gradient field of the rating structure based on the rate of change of the rating quantile distribution density, and identifies cluster center varieties with potential multi-objective adaptability.
[0034] An environmental heterogeneity variation simulator is introduced into each cluster to reconstruct the target scenario of the scoring trajectory of candidate varieties. The fitting scoring response of candidate varieties in a non-sampling environment is generated by perturbation inference.
[0035] Based on the reconstruction results, the generalization fitness index G of each candidate variety is calculated, and the varieties in C are screened a second time, retaining only the subset that performs best in multi-scenario simulations, forming the final set of candidate varieties for promotion level C'.
[0036] Preferably, S500 specifically includes:
[0037] Based on the score distribution characteristics of each variety in the candidate variety set C, an improved hierarchical multi-objective decision model IM-HMOD is constructed. The model integrates phenotypic stability index, physicochemical characteristic index and environmental adaptability score as the input vector of the decision layer.
[0038] A confidence interval is constructed for the score value of each variety, and a stability weight function is introduced to correct the range of score fluctuations and exclude samples with high volatility and low confidence.
[0039] The IM-HMOD model was used to assign weighted scores to each variety in C, and the environmental heterogeneity response ability of the top 5% of the varieties was tested. Finally, the target flaxseed variety T* with the best overall performance and strong cross-environment versatility was selected.
[0040] Output the target flaxseed variety T* as the screening endpoint.
[0041] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0042] 1. The technical solution provided by this invention significantly improves the accuracy of phenotypic performance and adaptability analysis of flaxseed varieties under different growth environments by introducing multi-source data fusion and multi-scale feature extraction mechanisms. Compared with traditional variety screening methods that rely on single indicators or single-point experiments, this invention achieves high-dimensional fusion of phenotypic information, physicochemical indicators, and meteorological variables, and combines an improved principal component compression and graph attention mechanism scoring model, thus possessing stronger environmental robustness and performance identification capabilities.
[0043] 2. The quantile response model, heterogeneous environment variation simulator, and improved multi-objective decision-making mechanism constructed in this invention solve the problem of insufficient identification of cross-environment adaptability in existing methods. By dynamically adjusting the score prediction and confidence interval under simulated conditions, the stability and generalizability of the target variety screening results are achieved, significantly improving the efficiency of breeding selection and the success rate of promotion, and has obvious practical application value. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0045] Figure 1 This is a mind map of the method of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Example 1, please refer to Figure 1 As shown in this embodiment, a flaxseed variety screening method based on comprehensive data analysis includes:
[0048] S100. Collect phenotypic characteristic data, grain physicochemical index data and environmental meteorological data of multiple flaxseed varieties under multiple growth environments, and construct a multidimensional dataset D;
[0049] S200. Based on the multidimensional dataset D, perform multi-scale feature recombination and variable enhancement to construct a high-dimensional feature vector T of flaxseed varieties under various environments. The feature vector is generated by a feature compression mechanism with a principal component retention rate of not less than 90%.
[0050] S300. Based on the feature vector T of each variety and the corresponding growth environment feature distribution, an adaptive weighted scoring model W is established, in which the scoring weight coefficient is adjusted according to the performance volatility and stability of the variety under extreme environment, and includes a phenotypic robustness discrimination submodule.
[0051] S400. Using model W, dynamic scoring is performed on each flaxseed variety to obtain a variety score set S. Multiquantile distribution analysis is performed on the score data in S to identify the candidate variety set C with the best fit in different target application scenarios.
[0052] S500 outputs the flaxseed variety with the best overall performance in candidate variety set C as the target screening result.
[0053] In this embodiment of the invention, the key to the flaxseed variety screening method based on comprehensive data analysis lies in establishing a high-quality, spatiotemporally unified, and structurally standardized multidimensional dataset D. This dataset consists of three data sources: phenotypic feature data, seed physicochemical index data, and environmental meteorological data. To ensure that this dataset possesses comprehensive feature extraction capabilities with high accuracy, high dimensionality, and strong adaptability, step S100 specifically includes:
[0054] This invention prioritizes collecting spatiotemporal growth data of flaxseed varieties in typical and marginal ecological zones through a multi-environmental joint experimental platform deployed in multiple locations. Using phenotypic monitoring devices with automatic image recognition and multispectral processing capabilities (such as high-throughput field phenotypic acquisition platforms and UAVs carrying visible light and infrared imaging equipment), continuous image acquisition is performed on each test variety at a set frequency (e.g., once daily or at each key growth stage) throughout the growing season. The system automatically extracts the phenotypic feature change trajectories from high-frequency images, including but not limited to:
[0055] Dynamic curve of reproductive cycle: Extract growth node time points based on the main stem development features in image sequences and calculate the length of reproductive stages;
[0056] Leaf Area Index (LAI): Leaf area is inferred from the image segmentation algorithm and the near-infrared channel vegetation index.
[0057] Biomass change rate: An estimation model is constructed by using the trend of plant height and crown width changes in time series images and combining it with the standard dry matter conversion coefficient, and the estimated biomass value is dynamically output.
[0058] After being encoded with multidimensional features, the above phenotypic data constitutes a subset of phenotypic features, providing high-resolution structural support for subsequent multi-scale feature fusion.
[0059] To reflect the differences in nutritional functionality among flaxseed varieties, this invention employs multi-channel near-infrared spectral decomposition (MC-NIR) technology for rapid, non-destructive detection of field-harvested seed samples. This method acquires spectral information of flaxseed in different component response peak regions through multi-channel wavelength settings, and, combined with an established inversion calibration model, accurately calculates the following key indicators:
[0060] α-Linolenic acid content: The intensity of the characteristic band located in the lipid absorption region is highly correlated with the content;
[0061] Lignin accumulation rate: obtained by multi-period sampling and calculation of the component time series change rate;
[0062] Grain water metabolism parameters: Grain dehydration rate and endpoint water stability were estimated using a water migration model.
[0063] Through the above physicochemical analysis, a standardized subset of physicochemical index data is formed, which has high repeatability and comparability, and can support the longitudinal tracking and horizontal comparison of grain quality.
[0064] To comprehensively evaluate the adaptability and resilience of various varieties to the ecological environment, this invention further integrates a standard meteorological data interface and a remote sensing satellite data module to retrieve environmental meteorological data that strictly corresponds to the test location and time. Specifically, this includes:
[0065] Hourly temperature variation sequence: Collect hourly temperature data at each test point and analyze the temperature peak and valley amplitude and duration of continuous hot / cold conditions;
[0066] Effective illumination duration: The total daily effective illumination is retrieved by inverting the illumination intensity from remote sensing images;
[0067] Precipitation frequency statistics: Statistical analysis of the number of rainfall days and intensity distribution per unit time during the growing season to assess water stress;
[0068] Climate Anomaly Index: Based on multi-year meteorological averages, it calculates the degree of deviation from the target test year and marks the level of environmental stress factors.
[0069] The above data uses a unified format encoding and is strictly aligned with the time point of phenotypic data collection to ensure the feasibility of data time synchronization and closed-loop modeling of environmental response.
[0070] To achieve data consistency and integrity, this invention introduces the following data processing mechanism:
[0071] Time alignment processing: unify the sampling frequency and timestamp of various types of data, supplement non-simultaneous sampling data, and construct a complete time series;
[0072] Spatial synchronous calibration: The remote sensing data is accurately registered to the center coordinates of the test plot according to latitude and longitude to ensure that the data reflects the actual plot conditions;
[0073] Missing data imputation and outlier removal: A time-series interpolation method based on multivariate collaborative prediction is used to imput missing data, and outliers are removed by Mahalanobis distance detection.
[0074] Generate a multidimensional dataset D using a unified coding structure: Standardize and structurally encode the three types of sub-datasets processed above, and merge them to construct a multidimensional dataset D for subsequent modeling and analysis.
[0075] Dataset D contains complete performance characteristics of each variety in multiple environments and at multiple time points, forming a structured, multi-scale, and highly semantic data structure, laying a solid data foundation for subsequent feature reorganization, scoring modeling, and target variety selection.
[0076] In this invention, to achieve a comprehensive evaluation of flaxseed varieties under different growth environments and multiple feature dimensions, it is necessary to perform in-depth data structure reconstruction, feature extraction, and variable compression on the constructed multidimensional dataset D. The core of step S200 lies in constructing a high-dimensional feature generation mechanism that integrates time series consistency, variable coupling, and scale hierarchical structure, and outputs a high-dimensional feature vector T that can accurately represent the characteristics of the variety, serving as the key input to the subsequent scoring model W.
[0077] First, the three types of data (phenotypic feature data, physicochemical index data, and environmental meteorological data) contained in dataset D are structurally labeled and logically hierarchically divided. To overcome the inconsistencies in time frequency, data scale, and variable dimensions among different data sources, this invention adopts the following labeling logic:
[0078] Time series consistency check: Automatically align the sampling frequency and time distribution of each variable to ensure that subsequent analysis is based on a unified time reference;
[0079] Correlation analysis: Spearman rank correlation and mutual information matrix are used to model the strength of the association between variables, thereby identifying variable groups with coupling relationships;
[0080] Scale hierarchy classification: Phenotypic characteristics and physicochemical indicators are classified into the "phenotypic-physicochemical dual-channel path" because they jointly describe individual traits and grain intrinsic quality; while environmental meteorological data form an independent "meteorological influencing factor coupling channel".
[0081] This structural calibration mechanism breaks away from conventional linear fusion methods and effectively enhances the model's ability to identify potential interactions between variables through channel separation modeling.
[0082] To further extract the collaborative variation features among deeper-level variables, this invention employs a nonlinear mapping algorithm to transform the channel-specific variables. Specific technical means include:
[0083] Kernel principal component analysis (PCA) and independent component analysis (ICA) were performed on each channel to extract nonlinear principal features and source signal features;
[0084] A cross-channel coupling layer is introduced to bidirectionally combine important features from the two channels (phenotypic and physicochemical) and work together with meteorological factor channels to construct a coupled feature space.
[0085] Based on the variable importance scoring mechanism (such as Gini coefficient ranking or information gain index), variables that contribute significantly to the differences in variety characteristics are screened to form the first round of candidate feature matrix T′.
[0086] T′ not only contains a compressed representation of the original variables, but also incorporates the interaction factors between cross channels, giving it stronger discriminative power.
[0087] After obtaining T′, in order to control the complexity of subsequent models and improve computational efficiency, this invention uses principal component analysis (PCA) to efficiently compress T′:
[0088] By calculating the cumulative contribution rate of each principal component, the number of principal components is dynamically selected to ensure that the cumulative explained variance is not less than 90%.
[0089] To maintain the scale coherence characteristics between variables, we extract cross-scale trend factors (such as linear fitting slope, periodic fluctuation amplitude, etc.) and embed them into the principal component space.
[0090] The principal component vector of the final output is concatenated with the co-variation factor to construct a compressed high-dimensional feature vector T.
[0091] T not only represents the aggregate performance of each variety under multiple variable dimensions, but also retains the synergistic and perturbation response characteristics among the original variables, providing support for the robustness and generalization ability of the scoring model.
[0092] Finally, the constructed high-dimensional feature vector T is mapped onto the variety feature space, forming a standardized and structurally uniform input dataset. This dataset has the following characteristics:
[0093] Each row represents the characteristic behavior of a variety under a specific environment;
[0094] Each column represents a highly discriminative feature dimension after compression and fusion processing;
[0095] The data matrix structure has good numerical stability and trainability, making it suitable for subsequent multi-model evaluation and scoring.
[0096] The establishment of this standard input structure significantly improves the convergence speed of the scoring model W during training and the accuracy during evaluation, and can effectively identify flaxseed varieties that still have potential advantages in highly complex environments.
[0097] To achieve a comprehensive assessment of the adaptability of flaxseed varieties to varying climatic environments, this invention proposes a multi-core weighted scoring model W that integrates high-dimensional feature modeling and a robust scoring mechanism. Model W aims to fully explore the multidimensional phenotypes, physicochemical responses, and stability differences exhibited by various varieties in different growth environments, in order to achieve an accurate evaluation of the overall performance of the target varieties.
[0098] This invention first constructs an environment-specific response factor matrix E based on the high-dimensional feature vector T of each flaxseed variety and its corresponding growth environment variables. The technical approach includes:
[0099] Calculate the response gradient of each eigenvector T under multiple growth environments;
[0100] Environmental factors (such as temperature, light, and precipitation) are modeled according to their spatial distribution density to form an environmental feature space;
[0101] The kernel density estimation (KDE) method is used to quantify the response intensity of each variety to different environmental variables, and the continuous response surface is output.
[0102] Matrix E can be viewed as an adaptive surface model in the binary interaction space of "variety-environment" to support the subsequent adjustment mechanism of scoring weights.
[0103] To capture the stability performance of varieties under typical and extreme climatic conditions, this invention designs an interval perturbation simulation mechanism, the specific process of which is as follows:
[0104] Based on the original environmental data, boundary condition perturbations are introduced to construct extreme scenarios (such as high temperature stress and drought delay);
[0105] Reassess the variability of each principal component in T under this scenario, and calculate the mean deviation and variance diffusion ratio.
[0106] A variety performance volatility index V is constructed by comprehensively considering both the magnitude of variation and the consistency of response.
[0107] V, as an important reflection of variety stability, is used to dynamically adjust the weights of the corresponding feature channels in model W to improve the model's sensitivity in identifying unstable varieties.
[0108] Based on the aforementioned data, a multi-core fusion scoring model W, including a phenotypic robustness discrimination submodule, is constructed, with the following structure:
[0109] Model W consists of multiple kernel functions in parallel, each responsible for scoring different dimensions (phenotypic, physicochemical, and environmental adaptability) of features.
[0110] The phenotypic robustness submodule embeds a "bidirectional feature stability filter". This module detects abnormal abrupt changes in the scoring sequence based on a sliding window mechanism and suppresses short-period fluctuation noise through an adaptive filtering algorithm.
[0111] The output of each kernel function is normalized and then fed into the fusion layer to form the initial scoring matrix S0;
[0112] This model structure enhances sensitivity control over noise, drift, and instability in the scoring, thereby improving scoring accuracy.
[0113] To further enhance the generalization ability of the scoring model in cross-environment applications, this invention introduces Bayesian optimization algorithm to iteratively learn the channel weights in model W. Technical details include:
[0114] The scoring error is used as the objective function, with stability and accuracy as constraints.
[0115] Predicting the performance of different weight combinations using a Gaussian process model;
[0116] The weights of the kernel function combination are automatically adjusted after each round of scoring to achieve dynamic self-adaptation;
[0117] This mechanism can quickly approximate the optimal weight solution, ensuring that the scoring model has high transfer performance in new environments.
[0118] To further enhance the model's structural recognition capabilities, this invention introduces the Graph Structure Attention Enhancement (GAM-R) module, which explicitly models the structural relationships between variety feature vectors T:
[0119] We constructed a weighted graph structure among varieties using Euclidean distance and environmental similarity as dual indicators. Nodes represent variety characteristics, and edge weights represent the degree of environmental interaction.
[0120] An improved graph attention mechanism is introduced to dynamically weight the information transmission between nodes and identify key varieties that perform stably or prominently in heterogeneous environments.
[0121] The attention-enhanced feature matrix A_T after output redistribution is used as the basis for the new input of model W;
[0122] This module effectively improves the model's ability to capture potential key features and prevents the score from being affected by sparse samples.
[0123] Finally, a robust two-layer dynamically tuned network (DRN) is constructed to further optimize the scoring model parameters:
[0124] The first layer of the network focuses on nonlinear mismatch pattern recognition between phenotype and environmental variables, and uses a convolutional normalized feedback structure to extract conflict signals.
[0125] The second layer network, guided by the error propagation path, adjusts the channel parameters of each kernel function in W in real time, enhancing the model's ability to correct score deviations.
[0126] The final set of parameters output by DRN will be applied to W simultaneously to complete the overall robustness iterative optimization of the scoring model;
[0127] This two-layer structure significantly improves the stability and accuracy of the scoring model under noise interference and high environmental complexity.
[0128] In summary, step S300, through the construction of a triple mechanism of high-order graph structure modeling, dynamic robust modulation and nonlinear stability modeling, not only breaks through the sensitivity of traditional scoring algorithms to data fluctuations and multi-source coupling, but also forms an innovative scoring framework with high transferability, adaptive adjustment capability and structure recognition capability, laying a technical foundation for the accurate determination of the final target variety.
[0129] To accurately assess the comprehensive adaptability and multi-objective performance of flaxseed varieties, this invention, based on the constructed weighted scoring model W, further proposes a set of scoring time series analysis, quantile model response modeling, and multi-objective adaptability screening mechanisms. These mechanisms aim to identify the set of candidate varieties C with optimal performance in specific application environments from the global scoring distribution. To improve the stability and practical generalizability of the results, step S400 also introduces density-sensitive clustering and heterogeneous environmental variation simulation modules to refine the selection of the final generalizable candidate variety set C'.
[0130] After the scoring model W is constructed, the system calls model W to score the comprehensive feature vector T of each flaxseed variety under all set growth environments (including typical and marginal environments), generating a scoring time series matrix S, where:
[0131] The rows represent the various flaxseed varieties;
[0132] The columns represent environmental conditions;
[0133] The cell represents the score output by model W; a higher value indicates better overall performance.
[0134] This scoring matrix not only reflects the overall performance of the varieties, but also reveals the scoring fluctuation structure of the varieties in multiple environmental scenarios.
[0135] To identify varieties that maintain stable output under different environmental combinations, this invention performs quantile distribution fitting analysis on each row vector of the scoring matrix S:
[0136] Construct a score distribution density function for each variety's score sequence;
[0137] Extract key statistical quantiles: such as the 25th percentile (Q1), median (Q2), and 75th percentile (Q3);
[0138] Establish a response relationship between the quantile structure and its corresponding environmental variables, and construct a quantile-environmental factor response model Q;
[0139] Model Q is used to identify whether a variety has a stable output range under specific ecological scenarios, and is a basic tool for judging the stability and resistance to disturbances of a variety.
[0140] To evaluate the scenario matching degree of candidate varieties, this invention designs a matching function F, with input parameters including:
[0141] The three elements of quantile structure (Q1, Q2, Q3);
[0142] Environmental weighting factor matrix (set according to important environmental variables of the target application area);
[0143] Function F calculates the fit score A_score for each variety in the target scenario using a weighted formula. The higher the score, the better its stability and median performance match the target requirements.
[0144] The system sorts all varieties in descending order based on their A_score, and selects the top N% as the high-adaptability candidate variety set C. This set possesses characteristics such as high suitability for specific application scenarios, stable scoring structure, and strong potential for generalization, laying the foundation for further in-depth screening.
[0145] To further enhance the wide-area generalization capability of the candidate variety set C, this invention introduces an improved density-sensitive clustering algorithm (DS-DBC), the main steps of which are as follows:
[0146] Distribution density analysis was performed on the subset of C with high structural stability of the scoring sequence.
[0147] Construct a score quantile density gradient field and identify potential cluster centers based on the local density change rate;
[0148] Several clusters are formed on the variety scoring structure diagram, and each cluster center represents a typical multi-objective adaptability model.
[0149] This clustering method can identify groups of varieties with common scoring patterns, thereby improving structural representativeness and screening breadth.
[0150] For each cluster, this invention further introduces an environmental heterogeneity variation simulator to simulate the potential performance of candidate varieties under unobserved (or newly introduced) environmental conditions:
[0151] Multiple fitting environment scenarios are constructed using the original environmental variable perturbation module;
[0152] The original scoring trajectory of each variety is perturbed and extrapolated to generate its fitted scoring response under the variable environment;
[0153] Calculate the stability index of the fit score and the magnitude of the variation response to form the generalization fitness index G;
[0154] Only those varieties that maintain high adaptability in most simulation environments are retained to form the final set of candidate varieties for promotion, C'.
[0155] This mechanism effectively avoids the model's reliance on "environment overfitting" and improves the adaptability and stability of the target varieties across ecological regions in actual promotion.
[0156] In the screening process of this invention, step S500 is the final decision-making stage of the entire flaxseed variety screening process. Its goal is to further select target flaxseed varieties T* from the already screened candidate variety set C, which exhibit excellent performance in phenotypic stability, physicochemical characteristics, and environmental adaptability, and have the potential for cross-ecological zone promotion. To this end, this step introduces an improved hierarchical multi-objective decision model (IM-HMOD) and a scoring confidence interval adjustment mechanism to achieve high-precision and high-robust target variety identification.
[0157] In the candidate variety set C, each variety already possesses complete scoring records and multidimensional feature descriptions. To systematically integrate information from different feature dimensions, this invention designs the IM-HMOD model, specifically including:
[0158] Model Structure: A three-tiered decision-making structure is adopted: the bottom layer is the grouping of indicators (phenotypic stability, physicochemical quality, and environmental adaptability); the middle layer is the normalized weighted calculation of sub-weights; and the top layer is the fusion of total scores.
[0159] Input feature vector construction: Each variety is treated as a sample, and its feature sub-vector is constructed from its three major indicator groups, where:
[0160] Phenotypic stability metrics include coefficient of variation (CV), continuous score slope, and short-period fluctuation frequency.
[0161] Physicochemical characteristics include α-linolenic acid content, lignan accumulation rate, and grain dry matter concentration;
[0162] The environmental adaptability score is the mean and distribution characteristics of the weighted score sequence output by the scoring model W under various environmental conditions.
[0163] Model features: IM-HMOD introduces a self-learning weight correction module on the basis of the traditional AHP (Analog-Hybrid Hierarchical Analysis) model and allows non-linear mapping of the decision function, so that it can automatically adjust the importance of dimensions according to the data characteristics;
[0164] The IM-HMOD model can efficiently aggregate complex multidimensional attributes to generate a target score that comprehensively reflects the overall performance of a variety.
[0165] To avoid the scoring being affected by extreme values or data noise, this invention performs confidence interval construction and stability adjustment on the score values of each variety before IM-HMOD evaluation. The process is as follows:
[0166] For each variety, construct a 95% confidence interval [L,U] based on the multi-environmental score data, and calculate the interval length Δ.
[0167] A stability weight function W_s = 1 / Δ is introduced, where the smaller Δ is, the higher the stability and the larger the weight. W_s is used as a multiplicative factor in the final scoring result of the IM-HMOD model to penalize samples with large scoring fluctuations.
[0168] This mechanism ensures the stability of the final score and prevents model misjudgment when data fluctuations are severe.
[0169] After the IM-HMOD model outputs a score and adjusts it for confidence intervals, this invention performs a step-by-step scoring and priority ranking on the candidate variety set C, and executes the following screening steps:
[0170] The varieties in C are ranked from highest to lowest according to the weighted scores, and the top 5% of the samples are selected for the final test.
[0171] For these high-scoring varieties, an "environmental diversity coping ability test" was conducted, including the following methods:
[0172] A comparative analysis of the score distribution in typical and marginal environments was conducted.
[0173] If a sample’s score drops significantly (greater than a set threshold) under extreme conditions, it will be excluded.
[0174] If its scoring curve is stable in multiple environments, it is considered to have cross-environment general capabilities.
[0175] Finally, the varieties that passed the above verification were retained, and the target flaxseed variety T* with the best overall performance was selected, which is the endpoint result of the screening process.
[0176] This decision-making mechanism avoids the misselection of varieties that are "averagely high but extremely unstable," ensuring that the target variety T* has high predictability and stable performance under different climates and cultivation regions.
[0177] Finally, the system outputs T* to the database and application interface for subsequent breeding recommendations, regional planting layout models, or digital agriculture platforms. The complete record of T* indicators, scoring trajectory, and feature structure will be stored in the database to support later tracking and feedback optimization.
[0178] Example 2: To verify the screening method (including steps S100–S500) of the present invention in actual screening of high-yield, high-quality flaxseed varieties that are adaptable to multiple environments, 15 candidate flaxseed varieties (labeled P1–P15) were tested in three environmental regions (E1: temperate humid region, E2: subtropical arid region, E3: high-altitude low-temperature region), and the target varieties were screened using the complete process.
[0179] Experimental Materials and Experimental Design:
[0180] Candidate varieties: P1–P15, from different breeding institutions, with diverse fertility and genetic backgrounds.
[0181] Test area:
[0182] E1: Average annual temperature 15℃, rainfall 800mm;
[0183] E2: Average annual temperature 25℃, rainfall 400mm;
[0184] E3: Average annual temperature 10℃, rainfall 600mm.
[0185] Sampling frequency: Image acquisition once a day, grain sampling three times per variety per environment, and meteorological data recorded hourly.
[0186] Experimental procedure:
[0187] The technology of this invention was used to complete image monitoring, near-infrared analysis and meteorological fusion data acquisition, and a multidimensional dataset D with dimensions of 15×3×120 (days)×(12+8+4)=15×3×120×24 was constructed. The 50-dimensional compressed feature vector of each variety was extracted through the S200 process.
[0188] A scoring model W was constructed. The model weights were adjusted using 5-fold cross-validation, ultimately generating 50 scoring values for each variety under different environments.
[0189] Each variety was scored P1–P15 under three environmental conditions:
[0190]
[0191] The Q-model was used to identify the stable interval of the scores. The median and IQR were calculated, and the corresponding fitness score F was obtained. The top 40% of the highly fit varieties were selected to form a set C (selecting five varieties: P1, P4, P7, P9, and P12).
[0192] Applying DS-DBC clustering analysis in C++, the results show clustering into two classes:
[0193] Cluster A (P1, P4, P7): The score density gradients within the clusters are consistent, and the fitting generalization is good;
[0194] Cluster B (P9, P12): The scores fluctuate in a concentrated manner but the response is excellent in a single environment.
[0195] After simulating environmental perturbations, the generalization fitness G was calculated. In cluster A, P1 and P7 both obtained G≥0.8 in all 10 simulated environments, meeting the selection criteria. The final set C'={P1,P7} was formed.
[0196] The two varieties in C' were further optimized using IM-HMOD:
[0197] Phenotypic stability index (CV): P1: 0.12; P7: 0.08;
[0198] Physicochemical indicators (α-linolenic acid content %): P1: 35; P7: 33;
[0199] Mean scores for environmental adaptability: P1: 80; P7: 81;
[0200] The model calculated confidence intervals (P1: [78–82]; P7: [79–83]), with stability weights W_s of 1 / 4 = 0.25 and 1 / 4 = 0.25, respectively. The overall score is as follows:
[0201] P1: Overall score 0.81;
[0202] P7: Overall score 0.84;
[0203] In further environmental response testing, P7 maintained a median score of ≥77 under low-temperature conditions in the marginal environment, which met the requirements. Therefore, the target variety T*=P7 was finally selected.
[0204] Beneficial Effects and Analysis: The multi-dimensional data fusion and evaluation model significantly improves discrimination ability: among 15 varieties, only P7 achieved a stable score of ≥79 under all environments; traditional single-environment screening methods cannot simultaneously meet this level of adaptability. The multi-stage screening mechanism improves universality: the screening path from 15→5→2→1 gradually eliminates varieties with poor environmental adaptability or low stability, ensuring a traceable and highly transparent process. The scoring method is robust and scalable: the model's score drift in extreme environments is <5 points, and it maintains high stability after simulation; indicating that this method has strong cross-regional application capabilities.
[0205] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for screening flaxseed varieties based on comprehensive data analysis, characterized in that: include: S100. Collect phenotypic characteristic data, grain physicochemical index data and environmental meteorological data of multiple flaxseed varieties under multiple growth environments, and construct a multidimensional dataset D; S200. Based on the multidimensional dataset D, perform multi-scale feature recombination and variable enhancement to construct a high-dimensional feature vector T of flaxseed varieties under various environments. The feature vector is generated by a feature compression mechanism with a principal component retention rate of not less than 90%. S300. Based on the feature vector T of each variety and the corresponding growth environment feature distribution, an adaptive weighted scoring model W is established, in which the scoring weight coefficient is adjusted according to the performance volatility and stability of the variety under extreme environment, and includes a phenotypic robustness discrimination submodule. S400. Using model W, dynamic scoring is performed on each flaxseed variety to obtain a variety score set S. Multiquantile distribution analysis is performed on the score data in S to identify the candidate variety set C with the best fit in different target application scenarios. S500 outputs the flaxseed variety with the best overall performance in candidate variety set C as the target screening result.
2. The flaxseed variety screening method based on comprehensive data analysis according to claim 1, characterized in that: S100 specifically includes: the phenotypic characteristic data including the dynamic change curve of the growth cycle, leaf area index and biomass change rate; the physicochemical index data including α-linolenic acid content, lignan accumulation rate and grain water metabolism parameters; and the environmental meteorological data including temperature, sunshine, precipitation frequency and climate anomaly factors.
3. The flaxseed variety screening method based on comprehensive data analysis according to claim 1, characterized in that: S200 specifically includes: The structure of different types of data in the multidimensional dataset D is defined, and the scale levels are divided according to the consistency of time series and the correlation of variables. Phenotypic-physicochemical dual-channel feature paths and meteorological influencing factor coupling channels are constructed respectively. Nonlinear mapping and coupling transformation are performed on variables at different scales to output candidate feature groups, and the first round of candidate feature matrix T′ is generated based on the importance score of the variables; The feature matrix T′ is subjected to dimension-wise principal component analysis to extract the principal component vectors with a cumulative contribution rate of not less than 90%, forming a high-dimensional compressed feature vector T, while simultaneously retaining the cross-scale collaborative change trend factors; The high-dimensional feature vector T is mapped into the variety feature space and used as the input feature set for the subsequent scoring model.
4. The flaxseed variety screening method based on comprehensive data analysis according to claim 1, characterized in that: The S300 specifically includes: Based on the high-dimensional feature vector T corresponding to each flaxseed variety and the feature distribution density of its growth environment, an environment-specific response factor matrix E is constructed to characterize the adaptability surface of the variety to different environmental variables. The stability of each principal component in T under typical and extreme environments is simulated by interval perturbation, the performance volatility index V is calculated, and this index is used as a dynamic adjustment factor for the scoring weight. A multi-core fusion scoring model W with a phenotypic robust discrimination submodule is constructed, in which the submodule adopts a bidirectional feature stability filter to identify and suppress phenotypic outliers and short-period fluctuation signals; By combining E, V, and robust output parameters, the weights of each scoring channel in model W are iteratively updated using a Bayesian optimization algorithm to form a weighted scoring structure with adaptive adjustment capabilities, thereby achieving accurate evaluation of flaxseed varieties under different climatic conditions.
5. The flaxseed variety screening method based on comprehensive data analysis according to claim 4, characterized in that: The scoring model W further includes a graph-structured attention enhancement module GAM-R and a robust two-layer dynamically tuned network DRN, specifically including: The feature vectors T among flaxseed varieties are constructed into a weighted graph structure using Euclidean distance and environmental similarity index. Nodes represent the features of each variety, and edge weights reflect the degree of environmental interaction. An improved graph attention mechanism is introduced to redistribute the weights of different varieties in heterogeneous environments to generate an attention-enhanced feature matrix A_T. A_T is input into a robust two-layer dynamic parameter tuning network DRN. The first layer is used to extract the nonlinear mismatch pattern between phenotype and environmental variables. The second layer uses the error propagation path as the adjustment benchmark to dynamically adjust the weight coefficients in W and optimize the robustness of the scoring model. Finally, the optimized parameter set output by DRN is applied to the scoring model W.
6. The flaxseed variety screening method based on comprehensive data analysis according to claim 1, characterized in that: The S400 specifically includes: The scoring model W is used to dynamically score each flaxseed variety under different growth environments, generating a scoring time series matrix S; A multiquantile distribution fitting analysis was performed on the scoring matrix S to construct a quantile-environmental factor response model Q, which was used to identify the stable performance range of each variety under a set environmental combination. Based on the output of the Q model, a variety-scenario matching function F is designed. The function F uses the median score, the upper and lower quartiles, and the target application environment weight factor as parameters to calculate the variety suitability score. Based on the fitness scores, the set of varieties C with the best overall performance under the target environment and target requirements is selected from the scoring matrix S, forming a set of highly adaptable candidate varieties.
7. The flaxseed variety screening method based on comprehensive data analysis according to claim 6, characterized in that: Further improvements to the adaptability of candidate variety set C include: An improved density-sensitive clustering algorithm, DS-DBC, is applied to a subset of samples with a highly stable rating structure in the rating matrix S. The algorithm constructs a gradient field of the rating structure based on the rate of change of the rating quantile distribution density, and identifies cluster center varieties with potential multi-objective adaptability. An environmental heterogeneity variation simulator is introduced into each cluster to reconstruct the target scenario of the scoring trajectory of candidate varieties. The fitting scoring response of candidate varieties in a non-sampling environment is generated by perturbation inference. Based on the reconstruction results, the generalization fitness index G of each candidate variety is calculated, and the varieties in C are screened a second time, retaining only the subset that performs best in multi-scenario simulations, forming the final set of candidate varieties for promotion level C'.
8. The flaxseed variety screening method based on comprehensive data analysis according to claim 1, characterized in that: The S500 specifically includes: Based on the score distribution characteristics of each variety in the candidate variety set C, an improved hierarchical multi-objective decision model IM-HMOD is constructed. The model integrates phenotypic stability index, physicochemical characteristic index and environmental adaptability score as the input vector of the decision layer. A confidence interval is constructed for the score value of each variety, and a stability weight function is introduced to correct the range of score fluctuations and exclude samples with high volatility and low confidence. The IM-HMOD model was used to assign weighted scores to each variety in C, and the environmental heterogeneity response ability of the top 5% of the varieties was tested. Finally, the target flaxseed variety T* with the best overall performance and strong cross-environment versatility was selected. Output the target flaxseed variety T* as the screening endpoint.