Big Data-Based Evaluation System for Germplasm Resources of Chinese Medicinal Herbs
By constructing a big data-based evaluation system for the germplasm resources of Chinese medicinal herbs, the problem of stable production caused by neglecting nonlinear physiological mechanisms and long-tail environmental risks in traditional methods has been solved. This system enables accurate grading and efficacy accumulation prediction under extreme climate conditions, thereby improving the stable production guarantee of Chinese medicinal herb germplasm resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 赤峰市农牧科学院
- Filing Date
- 2026-03-09
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods for evaluating germplasm resources of Chinese medicinal herbs neglect nonlinear physiological mechanisms and long-tail environmental risks, resulting in a lack of stable yield guarantees for selected varieties under sudden disasters. Furthermore, the mismatch between environmental data and growth stages in the evaluation leads to errors, making it impossible to accurately predict the accumulation of medicinal effects.
A big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources was constructed, including multi-source heterogeneous data fusion and feature mapping, nonlinear metabolic kinetic modeling, global linearization processing, environmental uncertainty construction, and robust grading determination modules. Through dynamic time warping, Koopman operator matrix updating, and recursive least squares method, the system achieves accurate grading of germplasm resources.
It improves the stable production problem caused by neglecting nonlinear physiological mechanisms and long-tail environmental risks in traditional evaluation methods, eliminates the mismatch between environment and growth stage, and improves the evaluation accuracy and adaptability to long-term monitoring under extreme climate conditions.
Smart Images

Figure CN122134200A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of evaluation technology for germplasm resources of Chinese medicinal materials, and in particular to a grading evaluation system for germplasm resources of Chinese medicinal materials based on big data. Background Technology
[0002] The grading evaluation of Chinese medicinal herb germplasm resources is the foundation for achieving "high quality, high price" and the selection of authentic varieties of Chinese medicinal herbs. With the development of big data technology, existing germplasm evaluation methods have gradually shifted from traditional field observation to data-driven predictive models. This involves using historical meteorological data, soil data, and genotype data, and constructing environment-yield prediction models through machine learning algorithms to assess the potential quality grade of specific germplasm in target planting areas.
[0003] However, traditional evaluations mostly use static statistics based on the mean, which ignore nonlinear physiological mechanisms and long-tail environmental risks, resulting in a lack of stable production guarantees for bred varieties under sudden disasters. Summary of the Invention
[0004] To overcome the above shortcomings, this invention provides a big data-based evaluation system for the germplasm resources of Chinese medicinal materials. It aims to improve the problem that traditional evaluation methods mostly use static statistics based on the mean, which ignore nonlinear physiological mechanisms and environmental long-tail risks, resulting in a lack of stable yield guarantee for selected varieties under sudden disasters.
[0005] This invention provides the following technical solution: a big data-based system for evaluating the grading of germplasm resources of Chinese medicinal herbs, comprising the following modules:
[0006] The multi-source heterogeneous data fusion and feature mapping module is used to align environmental time-series data and phenotypic data with phenological periods using dynamic time warping algorithm, and to embed genomic SNP site data to generate high-dimensional feature vectors.
[0007] The nonlinear metabolic kinetics modeling module is used to construct reaction-diffusion partial differential equations that include environmentally driven convection-diffusion terms and gene-regulated reaction source terms, characterizing the spatiotemporal evolution of effective components.
[0008] The global linearization module is used to solve the Koopman operator matrix using the extended dynamic mode decomposition algorithm, and to convert the partial differential equation into a linear discrete state equation in the observation function space.
[0009] An environmental uncertainty building module is used to map the environmental distribution to a regenerating kernel Hilbert space using kernel mean embedding, and to construct a sub-Bruker fuzzy sphere based on the maximum average difference distance;
[0010] The robustness rating module is used to construct a two-layer optimization model, solve for the minimum quality index under the worst distribution of the sub-Bruker fuzzy sphere through semidefinite programming relaxation, and then determine the rating.
[0011] The phenotypic trajectory tracking and model evolution correction module is used to monitor the actual growth trajectory residuals and to perform online incremental updates of the Koopman operator matrix using the recursive least squares method.
[0012] By adopting the above technical solutions, a metabolic mechanism model is constructed and combined with the optimization solution of the minimum quality by the Bloomberg bar, so as to achieve accurate grading under extreme climate. This improves the problem that traditional evaluation mostly adopts static statistics based on the mean, which ignores nonlinear physiological mechanisms and environmental long-tail risks, resulting in the lack of stable production guarantee for selected varieties under sudden disasters.
[0013] Preferably, the multi-source heterogeneous data fusion and feature mapping module includes:
[0014] The complete growth cycle sequence of historical high-yield and high-quality germplasm was selected as the benchmark reference template, and the real-time monitoring sequence of the germplasm to be evaluated was used as the query sequence.
[0015] Calculate the Euclidean distance between each time point in the query sequence and each time point in the benchmark template, and construct the cumulative distance matrix;
[0016] Search for a regular path from the starting point to the ending point in the cumulative distance matrix such that the cumulative Euclidean distance on the path is minimized, and then perform nonlinear stretching or compression on the environmental time series data based on the regular path.
[0017] Single nucleotide polymorphism sites in genomic data are encoded one-hot to generate static genotype vectors;
[0018] The aligned dynamic environment time-series vector is concatenated and spliced with the static genotype vector to generate a hybrid high-dimensional feature vector containing time-varying and time-invariant features.
[0019] Preferably, the nonlinear metabolic kinetics modeling module includes:
[0020] Based on the principles of fluid mechanics, the convective heat transfer coefficient is calculated using the Nusselt number correlation, which includes the Reynolds number and Prandtl number. A nonlinear convective diffusion term driven by real-time wind speed and temperature fields is constructed to characterize the material transport flux of effective components in plant tissues.
[0021] The rate of enzyme-catalyzed reaction is described by the Michaelis-Menten kinetic equation, and the genotype feature vector is mapped to the maximum reaction rate parameter and the Michaelis constant to construct a gene-regulated metabolic synthesis nonlinear source term.
[0022] A first-order kinetic decay coefficient, corrected by the environmental stress intensity index, is introduced to construct a reaction sink term characterizing the natural decomposition process of metabolites;
[0023] By superimposing the nonlinear convection-diffusion term, the nonlinear metabolic synthesis source term, and the reaction sink term, a partial differential equation is established to describe the changes in the concentration field of the effective component with time and spatial location along the plant growth axis.
[0024] Preferably, the global linearization processing module includes:
[0025] Construct an observation function dictionary containing the original state variables, polynomial combinations of the original state variables, and nonlinear functions based on radial basis kernel functions;
[0026] The low-dimensional state vector after discretization of the reaction-diffusion partial differential equation is input into the observation function dictionary to calculate the high-dimensional enhanced observation vector with a dimension higher than that of the original state space.
[0027] The completeness of the high-dimensional enhanced observation vector in the Hilbert space is verified, so that the evolution trajectory of the system dynamics in this space can be approximated by linear operators.
[0028] Preferably, the global linearization processing module further includes:
[0029] Construct an input data matrix containing the current time step boosted observation vector, and an output data matrix containing the next time step boosted observation vector;
[0030] Construct a least-squares optimization model with the objective function being the left multiplication of the Koopman operator matrix by the Frobenius norm, which is the difference between the input and output data matrices.
[0031] The Moore-Penrose pseudo-inverse matrix of the input data matrix is calculated using the singular value decomposition method.
[0032] Multiply the output data matrix by the Moore-Penrose pseudo-inverse matrix to obtain the optimal approximation finite-dimensional Koupman operator matrix, and establish an explicit linear discrete state-space equation.
[0033] Preferably, the environmental uncertainty construction module includes:
[0034] We select the Gaussian radial basis function as the positive definite kernel function and define the inner product operation rules for the reproducing kernel Hilbert space;
[0035] For each historical environmental data sample, feature mapping is used to transform it into elements in a high-dimensional feature space;
[0036] Calculate the arithmetic mean of the corresponding elements of all historical environmental data samples in the feature space to obtain the kernel mean embedding vector that represents the distribution of historical experience;
[0037] By utilizing the norm definition of the reproducing kernel Hilbert space, the maximum average difference distance between two kernel mean embedding vectors with different probability distributions is quantified.
[0038] Preferably, the environmental uncertainty construction module further includes:
[0039] The kernel mean embedding vector of the calculated historical experience distribution is used as the center of the sphere;
[0040] Based on the preset confidence level and the number of historical samples, the statistical deviation threshold required to cover the true distribution is calculated using the central inequality, and this threshold is set as the radius of the fuzzy sphere.
[0041] Construct a set of probability distributions containing all fuzzy spheres whose maximum average difference distance from the sphere center is less than or equal to the radius of the fuzzy sphere, and use this set as the environmental stress sub-Blu-ray fuzzy sphere.
[0042] Preferably, the robustness level determination module includes:
[0043] Define the problem as maximizing the expected value of the final cumulative amount of effective components as the objective function;
[0044] In the inner layer optimization, the worst environment distribution that minimizes the objective function value is searched using the fuzzy sphere of the blobs as the feasible region.
[0045] In the outer layer optimization, the quality index bound that the global linear state-space model can achieve under the worst environmental distribution conditions is calculated;
[0046] The linear evolution equations and control input constraints of the global linear state-space model are used as all the physical constraints of the two-layer optimization model.
[0047] Preferably, the robustness level determination module further includes:
[0048] Construct a symmetric semidefinite matrix variable defined by the outer product of the lifting observation vector and its transpose;
[0049] By utilizing the properties of matrix trace operations, the nonlinear objective function and constraints in the original optimization problem are transformed into linear functions of the symmetric semidefinite matrix variables;
[0050] The non-convex rank constraint of the symmetric semidefinite matrix variable is relaxed, thereby transforming the original problem into a convex optimization problem on a semidefinite cone.
[0051] The semidefinite programming problem is solved using the original dual interior point method to obtain the global optimal matrix, and the minimum quality index is restored from the global optimal matrix based on the spectral decomposition technique.
[0052] Preferably, the phenotypic trajectory tracking and model evolution correction module includes:
[0053] Set a forgetting factor to exponentially decay the weights of historical data;
[0054] At each growth monitoring time, the prediction error vector between the actual observed vector and the predicted vector based on the current Koopman operator matrix is calculated.
[0055] Calculate the gain matrix based on the covariance matrix of the previous time step and the current observation vector;
[0056] The current Koopman operator matrix is recursively updated using the gain matrix and the prediction error vector.
[0057] The covariance matrix is updated for calculation in the next time step, completing the online adaptive correction of the operator matrix.
[0058] The present invention has the following beneficial effects:
[0059] 1. In this invention, by constructing a metabolic mechanism model and combining it with the optimization solution of the minimum quality, the precise grading under extreme climates can be achieved. This improves the problem that traditional evaluations mostly use static statistics based on the mean, which ignore nonlinear physiological mechanisms and environmental long-tail risks, resulting in a lack of stable production guarantee for selected varieties under sudden disasters.
[0060] 2. In this invention, the phenological periods are aligned by using a dynamic time warping algorithm through a multi-source heterogeneous data fusion module, thereby eliminating spatiotemporal misalignment. This improves upon the traditional evaluation method, which mostly uses absolute calendar time analysis and ignores the asynchronous nature of biological developmental rhythms, resulting in the failure of physical matching between environmental factors and growth stages and the distortion of feature extraction.
[0061] 3. In this invention, a nonlinear metabolic kinetic model is constructed and linearized using the Koopman operator, thereby balancing biological mechanism characterization and computational efficiency. This improves upon the problem that traditional evaluations mostly use purely data-driven black-box models, which lack descriptions of the synthesis and transport mechanisms of active ingredients, resulting in low accuracy in predicting drug efficacy accumulation under complex environmental interactions.
[0062] 4. In this invention, the recursive least squares method is used to perform online incremental updates of the operator matrix, thereby endowing the model with adaptive ability to resist concept drift. This improves the problem that traditional models mostly adopt one-time offline training, which cannot perceive the cumulative error caused by the evolution of soil microenvironment, resulting in a significant decrease in the accuracy of evaluation in the later stage of long-term monitoring. Attached Figure Description
[0063] Figure 1This is a module architecture diagram of the big data-based germplasm resource grading evaluation system for Chinese medicinal materials proposed in this invention.
[0064] Figure 2 This is a flowchart of the multi-source heterogeneous data fusion and feature mapping process for the big data-based germplasm resource grading evaluation system for Chinese medicinal materials proposed in this invention.
[0065] Figure 3 This is a flowchart of the global linearization process for the big data-based evaluation system for the germplasm resources of Chinese medicinal materials proposed in this invention.
[0066] Figure 4 This is a flowchart of the robustness rating determination and SDP solution of the big data-based Chinese medicinal herb germplasm resource grading system proposed in this invention.
[0067] Figure 5 This is a flowchart of the online evolution correction process for the model of the big data-based evaluation system for germplasm resources of Chinese medicinal materials proposed in this invention. Detailed Implementation
[0068] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Example 1: In the first embodiment of the present invention, the present invention provides a big data-based system for evaluating the grading of germplasm resources of Chinese medicinal materials, such as... Figures 1-5 As shown, it includes the following modules:
[0070] The multi-source heterogeneous data fusion and feature mapping module is used to align environmental time-series data and phenotypic data with phenological periods using dynamic time warping algorithm, and to embed genomic SNP site data to generate high-dimensional feature vectors.
[0071] Furthermore, the multi-source heterogeneous data fusion and feature mapping module includes:
[0072] The complete growth cycle sequence of historical high-yield and high-quality germplasm was selected as the benchmark reference template, and the real-time monitoring sequence of the germplasm to be evaluated was used as the query sequence.
[0073] Calculate the Euclidean distance between each time point in the query sequence and each time point in the benchmark template, and construct the cumulative distance matrix;
[0074] Search for a regular path from the starting point to the ending point in the cumulative distance matrix such that the cumulative Euclidean distance on the path is minimized, and then perform nonlinear stretching or compression on the environmental time series data based on the regular path.
[0075] Single nucleotide polymorphism sites in genomic data are encoded one-hot to generate static genotype vectors;
[0076] The aligned dynamic environment time-series vector and the static genotype vector are concatenated and spliced to generate a hybrid high-dimensional feature vector containing both time-varying and time-invariant features.
[0077] Specifically, the multi-source heterogeneous data fusion and feature mapping module mainly addresses the problem of phenological period misalignment caused by climate differences in the growth cycle of Chinese medicinal materials, as well as the problem of heterogeneous data not being uniformly input into the model.
[0078] The system first collects two types of basic data. The first type is environmental time-series data, including temperature, humidity, light intensity, and soil trace element content; this data is a dynamic sequence that changes over time. The second type is genomic data, specifically single nucleotide polymorphism (SNP) site data obtained through high-throughput sequencing; this data represents static characteristics that do not change over time.
[0079] For environmental time-series data, due to climatic differences across years or production areas, phenological nodes such as the greening period and flowering period of the same medicinal herb may shift in calendar time. Directly inputting absolute time will lead to misaligned feature extraction. This module selects complete growth cycle data of historically selected high-yielding and high-quality germplasm as the benchmark reference template sequence. The real-time monitoring data of the germplasm to be evaluated will be used as the query sequence. .
[0080] Set a reference template sequence The length is , represented as Query sequence The length is , represented as .in and All are feature vectors containing multidimensional environmental factors.
[0081] The module first calculates the Euclidean distance between each time point in the query sequence and each time point in the baseline reference template. The distance metric formula is as follows: ;in Indicates the query sequence number The first time point and benchmark reference template Local distance between points in time The number of dimensions of environmental factors. and The corresponding time points are the first The numerical values of each environmental factor.
[0082] Subsequently, build Cumulative distance matrix of dimension Any point in the matrix Cumulative distance value The result was obtained through recursive calculation using dynamic programming:
[0083] ;
[0084] in Represents the extension of the timeline. Represents compression of the timeline. This represents the alignment of the timeline. This is a minimum value function. The boundary conditions are set as follows: .
[0085] Module in cumulative distance matrix Search for a line starting from the origin To the finish line Regular path This path satisfies the condition of minimizing the cumulative Euclidean distance along the path. Based on this path... Query sequence Environmental data is mapped to a benchmark reference template. On the timeline. This process nonlinearly maps the physical timeline to a biological phenological axis, eliminating differences in growth rate and generating an aligned dynamic environmental timeline vector. .
[0086] For genomic SNP data, which consists of the base characters A, T, C, and G, direct numerical calculations are not possible. This module performs one-hot encoding on key SNP sites. (Assuming...) The key SNP site, for the first If the base type of a site is A, then it is encoded as: If it is T, then the encoding is And so on. All of them. The coding vectors of each locus are concatenated to generate a static genotype vector. .
[0087] The module performs a feature concatenation operation. This aligns the dynamic environment time-series vectors. With static genotype vector Concatenate the features along the feature dimension to construct a hybrid high-dimensional feature vector. :
[0088] ;
[0089] in This represents a vector concatenation operation. The resulting high-dimensional feature vector is... It includes both historical environmental stress information that has been adjusted for time bias and genotypic information that determines the genetic potential of germplasm, serving as standard input data for subsequent nonlinear metabolic kinetic modeling modules. Through this processing, spatiotemporal and dimensional unification of heterogeneous data is achieved, resolving the evaluation error problem caused by the mismatch between environmental data and growth stage in traditional methods.
[0090] The nonlinear metabolic kinetics modeling module is used to construct reaction-diffusion partial differential equations that include environmentally driven convection-diffusion terms and gene-regulated reaction source terms, characterizing the spatiotemporal evolution of effective components.
[0091] Furthermore, the nonlinear metabolic kinetics modeling module includes:
[0092] Based on the principles of fluid mechanics, the convective heat transfer coefficient is calculated using the Nusselt number correlation, which includes the Reynolds number and Prandtl number. A nonlinear convective diffusion term driven by real-time wind speed and temperature fields is constructed to characterize the material transport flux of effective components in plant tissues.
[0093] The rate of enzyme-catalyzed reaction is described by the Michaelis-Menten kinetic equation, and the genotype feature vector is mapped to the maximum reaction rate parameter and the Michaelis constant to construct a gene-regulated metabolic synthesis nonlinear source term.
[0094] A first-order kinetic decay coefficient, corrected by the environmental stress intensity index, is introduced to construct a reaction sink term characterizing the natural decomposition process of metabolites;
[0095] By superimposing the nonlinear convection-diffusion term, the nonlinear source term of metabolic synthesis, and the reaction sink term, a partial differential equation describing the change of the concentration field of the effective component with time and spatial location along the plant growth axis is established.
[0096] Specifically, the nonlinear metabolic kinetics modeling module receives a mixed high-dimensional feature vector output from the previous module, which includes aligned environmental time-series components and static genotype components. The module treats the plant as a bioreactor influenced by environmental fields, establishing a model along the plant's growth axis. and time Changing effective ingredient concentration field The evolution equation.
[0097] The module models the transport flux of active ingredients within plant tissues based on fluid dynamics principles. The material flow within the plant is driven by the external environment, particularly wind and temperature fields, exhibiting fluid-like convective heat transfer characteristics. The module utilizes Reynolds number... With Prandtl number Calculate the Nusel number This is used to determine the convective heat transfer coefficient driven by the environment. .
[0098] The Nusselt number correlation is set as follows: ;in , , The Reynolds number is an empirical fluid dynamics constant determined based on the geometry of the plant stem. Defined as ,in For fluid density, For real-time wind speed, The characteristic diameter of the plant. For fluid dynamic viscosity. Prandtl number. Defined as ,in For specific heat capacity, is the thermal conductivity.
[0099] The calculated Nusel number Directly determines the convective heat transfer coefficient This is then mapped to the equivalent material transport rate within the plant. Constructing a nonlinear convection-diffusion term as follows:
[0100] ;
[0101] in For the effective diffusion coefficient, The convection velocity is driven by the real-time environmental field. This term represents the spatial gradient operator along the growth axis. It characterizes the passive transport and diffusion process of the active ingredient under the influence of the environmental physical field.
[0102] For the biosynthesis of active ingredients, the module uses the Michaelis-Menten kinetic equation to describe the rate of enzyme-catalyzed reactions. Genotype data determines the activity and abundance of the relevant synthases. The module establishes a mapping function to convert the input static genotype feature vector... Converted to the maximum reaction rate parameter in the Michaelis-Menten equation With Michaelis constant .
[0103] Constructing nonlinear source terms for metabolic synthesis : ;in The maximum synthesis rate is determined by genotype. The affinity constant between the enzyme and its substrate, determined by genotype. For precursor substrate in The concentration at specific times. This parameter characterizes the rate at which plants actively synthesize active ingredients under the control of a specific genotype, reflecting the positive contribution of internal biological mechanisms to ingredient accumulation.
[0104] To address the natural decomposition and degradation by environmental stressors of the active ingredients, the module constructs reaction sinks. Considering that environmental stress accelerates the degradation or transformation of secondary metabolites, the module introduces an environmental stress intensity index. For the first-order kinetic attenuation coefficient Make corrections.
[0105] Constructing a reaction pool to characterize metabolite breakdown processes as follows:
[0106] ;
[0107] in Based on the natural attenuation coefficient, Environmental sensitivity coefficient, This is a dimensionless environmental stress intensity index calculated based on anomalies in temperature and humidity. This represents the current concentration of the active ingredient. This value indicates the loss and metabolic depletion of the active ingredient during the growth process.
[0108] The module linearly superimposes the aforementioned nonlinear convection-diffusion terms, metabolic-synthetic nonlinear source terms, and reaction sink terms, and establishes a field describing the concentration of the effective component based on the law of conservation of mass. Response-diffusion partial differential equations as a function of time and spatial position along the plant growth axis:
[0109] ;
[0110] in This represents the rate of change of the concentration of the effective component over time. This partial differential equation forms the input object for the subsequent global linearization module, providing a mathematical model based on biophysical mechanisms for the dynamic quality evaluation of germplasm resources.
[0111] The global linearization module is used to solve the Koopman operator matrix using the extended dynamic mode decomposition algorithm, and to transform the partial differential equations into linear discrete state equations in the observation function space.
[0112] Furthermore, the global linearization processing module includes:
[0113] Construct an observation function dictionary containing the original state variables, polynomial combinations of the original state variables, and nonlinear functions based on radial basis kernel functions;
[0114] The low-dimensional state vector after discretization of the reaction-diffusion partial differential equation is input into the observation function dictionary to calculate the high-dimensional enhanced observation vector with a dimension higher than that of the original state space.
[0115] Verify the completeness of the high-dimensional lifting observation vector in the Hilbert space, so that the evolution trajectory of the system dynamics in this space can be approximated by linear operators.
[0116] The global linearization processing module also includes:
[0117] Construct an input data matrix containing the current time step boosted observation vector, and an output data matrix containing the next time step boosted observation vector;
[0118] Construct a least-squares optimization model with the objective function being the left multiplication of the Koopman operator matrix by the Frobenius norm, which is the difference between the input and output data matrices.
[0119] The Moore-Penrose pseudo-inverse matrix of the input data matrix is calculated using the singular value decomposition method.
[0120] Multiply the output data matrix by the Moore-Penrose pseudo-inverse matrix to obtain the optimal approximation finite-dimensional Koupman operator matrix, and establish an explicit linear discrete state-space equation.
[0121] Specifically, the global linearization module receives the nonlinear reaction-diffusion partial differential equations output from the previous module. Because these equations contain complex nonlinear coupling terms, direct solution is computationally intensive and rarely yields the global optimum. This module utilizes Koopman operator theory to elevate the finite-dimensional nonlinear dynamic system to an infinite-dimensional Hilbert space, where linear operators are sought to approximate the evolution trajectory of the original system.
[0122] The module first performs spatial discretization on the input reaction-diffusion partial differential equation, for example, by using the finite difference method to obtain low-dimensional state vectors at discrete time points. This vector It includes the concentration values of effective components at each spatial node of the plant.
[0123] Subsequently, an observation function dictionary is constructed. This dictionary consists of a set of linearly independent basis functions, designed to capture the nonlinear characteristics of the system. The dictionary of observation functions is defined as follows:
[0124] ;
[0125] in For the dictionary vector of observation functions, For the enhanced spatial dimension, and It is much larger than the dimension of the original state space. These are specific basis functions. In this embodiment, the basis functions include the original state variables themselves. Polynomial combination of the original state variables and radial basis kernel function The introduction of the radial basis kernel function ensures the completeness of the high-dimensional enhanced observation vector in the Hilbert space, enabling the evolution trajectory of the system dynamics in this function space to be effectively approximated by linear operators.
[0126] The module will use the low-dimensional state vector Input the dictionary of observation functions, and calculate High-dimensional boosted observation vector at time step : .
[0127] Based on snapshot sequences generated from historical growth data or simulation data, the module constructs an input data matrix. With output data matrix Assume it exists. Snapshot data at each time step:
[0128] ;
[0129] in Including from time 1 to time 2 The observation vector is increased in dimension at time step. Including from time 2 to time 3 A high-dimensional enhanced observation vector at each time point. Each column vector represents a snapshot of the system's full state at a given time point.
[0130] In order to find the optimal linear evolution operator , making The module constructs a least-squares optimization model. The objective function is set as minimizing the Frobenius norm of the linear mapping error:
[0131] ;
[0132] in To optimize the objective function, Let be the finite-dimensional Koupman operator matrix to be solved. This represents the Frobenius norm of the matrix, which is the square root of the sum of the squares of all its elements. The objective function characterizes the linear prediction value. Compared with the true value The overall deviation between them.
[0133] To solve the above optimization problem, the module uses the Singular Value Decomposition (SVD) method to calculate the input data matrix. Moore-Penrose pseudoinverse matrix For the matrix Perform singular value decomposition: ;in It is a left singular vector matrix. It is a diagonal singular value matrix. This is the transpose of the right singular vector matrix. The pseudo-inverse matrix is calculated based on this. : ;in This is a diagonal matrix formed by taking the reciprocals of the singular values. For singular values close to zero, truncation is used to enhance numerical stability.
[0134] The module will output a data matrix. The calculated pseudo-inverse matrix Multiplying them yields the optimal approximation of the Koopman operator matrix. : Based on the solved matrix The module establishes explicit linear discrete state-space equations: ;in This is the high-dimensional state vector for the next time step. The Koopman operator matrix The state evolution part in the text, The Koopman operator matrix The control input section in the middle, The environmental control variables are used. This linear equation transforms the originally complex nonlinear PDE evolution into a simple matrix multiplication operation, serving as the standard input model for subsequent robustness level determination modules and semidefinite programming solutions.
[0135] An environmental uncertainty building module is used to map the environmental distribution to a regenerating kernel Hilbert space using kernel mean embedding, and to construct a sub-Bruker fuzzy sphere based on the maximum average difference distance;
[0136] Furthermore, the environmental uncertainty building block includes:
[0137] We select the Gaussian radial basis function as the positive definite kernel function and define the inner product operation rules for the reproducing kernel Hilbert space;
[0138] For each historical environmental data sample, feature mapping is used to transform it into elements in a high-dimensional feature space;
[0139] Calculate the arithmetic mean of the corresponding elements of all historical environmental data samples in the feature space to obtain the kernel mean embedding vector that represents the distribution of historical experience;
[0140] By utilizing the norm definition of the reproducing kernel Hilbert space, the maximum average difference distance between two kernel mean embedding vectors with different probability distributions is quantified.
[0141] The environmental uncertainty building block also includes:
[0142] The kernel mean embedding vector of the calculated historical experience distribution is used as the center of the sphere;
[0143] Based on the preset confidence level and the number of historical samples, the statistical deviation threshold required to cover the true distribution is calculated using the central inequality, and this threshold is set as the radius of the fuzzy sphere.
[0144] Construct a set of probability distributions containing all fuzzy spheres whose maximum average difference distance from the sphere center is less than or equal to the radius of the fuzzy sphere, and use this set as the environmental stress sub-Bruker fuzzy sphere.
[0145] Specifically, the environmental uncertainty construction module receives historical environmental big data, including long-term rainfall, sunshine duration, and accumulated temperature data recorded by weather stations. Since environmental data often does not follow a normal distribution and has extreme long-tail distribution characteristics, the module uses the Regenerative Kernel Hilbert Space (RKHS) tool to map the probability distribution to geometric points, constructing a fuzzy set containing the true distribution.
[0146] The module first defines the regenerating kernel Hilbert space. The geometric structure. A Gaussian radial basis function is chosen as the positive definite kernel function. The kernel is used to measure the similarity between two environment state vectors. The kernel function is defined as: ;in and Let be any two environment state vectors. For the Euclidean norm, This is the bandwidth parameter of the kernel function, used to control the local sensitivity of the feature map. The feature map is defined through this kernel function. , making This mapping transforms the originally low-dimensional and non-linear environmental data space Mapping to an infinite-dimensional high-dimensional feature space .
[0147] For the collected Historical environmental data sample These samples constitute the historical experience distribution. The module calculates the arithmetic mean of the corresponding elements of all samples in the feature space, obtaining the kernel mean embedding vector representing the historical experience distribution. : ;in The total number of samples. For the first Feature mapping of historical samples. This embedding vector It contains all the statistical moments of the original distribution and is the unique representation of the original probability distribution in the RKHS space.
[0148] The module utilizes the norm definition of the reproducing kernel Hilbert space to quantify two different probability distributions. and The distance between them. The maximum mean difference distance (MMD) is defined as the Hilbert spatial distance between the embedding vectors corresponding to two distributions: ;in For distribution Kernel mean embedding, For distribution Kernel mean embedding, This is the RKHS space norm. This distance index is effective in measuring the differences between non-Gaussian distributions.
[0149] To ensure that the constructed set can cover unknown real-world environment distributions The module needs to determine the radius of the fuzzy sphere. Based on a preset confidence level and the number of historical samples The statistical bias threshold is calculated using McDiermead's concentration inequality: ;in This is an upper bound constant for the kernel function, typically the Gaussian kernel. . The allowed probability of default, for example, is 0.05. This is the sample size. This threshold... It represents the maximum possible statistical deviation between the empirical distribution and the true distribution given a sample size.
[0150] The module uses the calculated historical experience distribution embedding vector Using the sphere as the center, and the calculated statistical deviation threshold... Using the radius as the basis, construct a fuzzy sphere for environmental stress analysis. :
[0151] ;
[0152] in This represents the set of all probability measures in the environment state space. The fuzzy sphere... Includes all historical experience distributions that are less than or equal to the distance in MMD. The probability distribution of the set. This set serves as the input to the subsequent robustness rating module, forcing the evaluation model to find the worst-case scenario among all possible distributions within the set, thereby ensuring the adaptability of the evaluation results to extreme climate change.
[0153] The robustness rating module is used to construct a two-layer optimization model, solve for the minimum quality index under the worst distribution of the sub-Bruker fuzzy sphere through semidefinite programming relaxation, and then rate it.
[0154] Furthermore, the robustness level determination module includes:
[0155] Define the problem as maximizing the expected value of the final cumulative amount of effective components as the objective function;
[0156] In the inner layer optimization, the feasible region is the fuzzy sphere of the blobs, and the worst environment distribution that minimizes the objective function value is searched.
[0157] In the outer layer optimization, the quality index bound that the global linear state-space model can achieve under the worst environmental distribution conditions is calculated;
[0158] The linear evolution equations and control input constraints of the global linear state-space model are used as all the physical constraints of the bi-level optimization model.
[0159] The robustness rating module also includes:
[0160] Construct a symmetric semidefinite matrix variable defined by the outer product of the lifting observation vector and its transpose;
[0161] By utilizing the properties of matrix trace operations, the nonlinear objective function and constraints in the original optimization problem are transformed into linear functions with respect to symmetric semidefinite matrix variables;
[0162] By relaxing the non-convex rank constraints of the symmetric semidefinite matrix variables, the original problem is transformed into a convex optimization problem on a semidefinite cone.
[0163] The semidefinite programming problem is solved using the original-dual interior-point method to obtain the global optimal matrix, and the minimum quality index is restored from the global optimal matrix based on the spectral decomposition technique.
[0164] Specifically, the robustness rating module receives the linear discrete state-space equations output by the global linearization processing module and the environmental stress fuzzy sphere output by the environmental uncertainty construction module. The core task of the module is to calculate the lower bound of the cumulative amount of effective components of germplasm resources in all possible environmental distributions defined by the fuzzy sphere, and use this as the basis for rating.
[0165] The module first constructs a two-layer optimization model. The outer layer optimization aims to find the minimum quality of germplasm resources. The inner layer optimization, on the other hand, optimizes the model within a given set of parameters. Within, find the worst environmental distribution that minimizes quality indicators. .
[0166] Define the objective function Let be the mathematical expectation of the final cumulative amount of the effective component. This is the final moment of the growth cycle. This is the boosted state vector at the final moment. This is the selection vector for extracting the effective component concentration from the state vector. The two-level optimization model is expressed as follows: ;in Indicates the probability distribution The expected value of the following mathematical expression To select the transpose of a vector. These are state variables constrained by system dynamics. Physical constraints include the linear evolution equations generated by the global linearization module:
[0167] ;
[0168] in The state evolution matrix is... To control the input matrix, The environmental input variable is used. This constraint ensures that, when searching for the worst distribution, the evolution of the system's state always follows the objective laws of plant physiology.
[0169] The aforementioned bilevel optimization problem involves searching for a probability distribution, and the state variables... Environmental input The existence of coupling causes the original problem to be a non-convex optimization problem. The module performs convexification through variable substitution. An extended state vector is defined. Construct a symmetric semidefinite matrix variable defined by the outer product of the lifting observation vector and its transpose. : ;in It is a symmetric positive definite matrix with rank 1. This matrix contains the second-order moment information of the state variables and the cross-correlation information between the state and the control input.
[0170] Utilizing the properties of matrix trace operations The module transforms the quadratic or nonlinear terms in the original optimization problem into terms about matrix variables. A linear function. For example, a quadratic term in the original objective function or constraints. It can be rewritten as: ;in This is the expanded coefficient matrix. After transformation, the original linear evolution equation... Rewritten as about and Linear equality constraints: ;in and For the reason and The constructed coefficient matrix.
[0171] Original matrix variables The non-convex rank constraint must be satisfied. The module relaxes this constraint, discarding the rank constraint and retaining only the semi-qualitative constraint. At this point, the original non-convex optimization problem is transformed into a convex optimization problem SDP on a semi-fixed cone: The module uses the primal-dual interior-point method to iteratively solve this semidefinite programming problem. This algorithm, by approximating the optimal solution along the central path of the feasible region, can obtain the globally optimal matrix sequence in polynomial time. .
[0172] Obtain the globally optimal matrix Then, the module reconstructs the physical state from the spectral decomposition technique. Perform eigenvalue decomposition: ;in For eigenvalues, The eigenvectors are eigenvalues. Since the relaxed solution may have a rank other than 1, the module selects the eigenvector corresponding to the largest eigenvalue as the optimal state vector. The best approximation. Calculate the minimum quality index. This indicator represents the minimum effective component content achievable by this germplasm resource under all extremely harsh environmental distributions covered by the fuzzy sphere. The module will... Compare with the preset national standard level threshold, if If it is, then it is determined to be a primary source; if If it is, then it is determined to be a secondary source. The final determination result is output.
[0173] The phenotypic trajectory tracking and model evolution correction module is used to monitor the actual growth trajectory residuals and to perform online incremental updates of the Koopman operator matrix using the recursive least squares method.
[0174] Furthermore, the phenotypic trajectory tracking and model evolution correction module includes:
[0175] Set a forgetting factor to exponentially decay the weights of historical data;
[0176] At each growth monitoring time, the prediction error vector between the actual observed vector and the predicted vector based on the current Koopman operator matrix is calculated.
[0177] Calculate the gain matrix based on the covariance matrix of the previous time step and the current observation vector;
[0178] The current Koopman operator matrix is recursively updated using the gain matrix and the prediction error vector;
[0179] The covariance matrix is updated for calculation in the next time step, completing the online adaptive correction of the operator matrix.
[0180] Specifically, the phenotypic trajectory tracking and model evolution correction module receives the initial Koopman operator matrix from the global linearization processing module. In addition, it receives real-time observation data from the multi-source heterogeneous data fusion module. The module executes the following iterative calculation process within each growth monitoring time step to achieve real-time correction of the plant growth kinetics model.
[0181] The module first sets the forgetting factor. Its value range is The forgetting factor is used to control the weight of historical data on the current parameter estimation. The smaller the value, the faster the system forgets old data and the higher its sensitivity to new data, making it suitable for periods of rapid growth and mutation. A value closer to 1 indicates stronger system memory, making it suitable for periods of stable growth. Initialize the covariance matrix. It is usually set to a large positive number multiplied by the identity matrix. ,in It is a pre-defined large positive number used to characterize the uncertainty of the initial estimate.
[0182] During growth monitoring The sensor collects the actual phenotypic observation vectors. The module retrieves the previous time step. Improved observation vector and the estimated value of the Koopman operator matrix stored at the current time. Calculate the prediction error vector. : ;in These are the high-dimensional feature vectors that are actually observed. This is the estimated state value at the current time obtained by the model based on the state at the previous time step. It characterizes the degree of deviation of the current model from the actual growth trajectory.
[0183] The module uses the covariance matrix updated in the previous time step. and the observation vector of the previous time step Calculate the Kalman gain matrix : ;in This is the transpose of a vector. The denominator part... This is the scalar normalization factor. Gain matrix. This determines the weighting of prediction errors in parameter updates.
[0184] Using the calculated gain matrix and prediction error vector The current Koopman operator matrix is corrected to obtain a new estimate. : ;in This is the transpose of the gain matrix. This step utilizes the principle of orthogonal projection to project the prediction error onto the parameter space, thereby minimizing the prediction error of the corrected operator matrix at the current data point.
[0185] In order to proceed with the next iteration, the module updates the covariance matrix. :
[0186] ;
[0187] in This is used to prevent the covariance matrix from approaching zero over time, thus maintaining the algorithm's continuous tracking capability. Updated This reflects the change in the accuracy of parameter estimation and propagates to the next time step. Used to calculate the new gain.
[0188] Through the above steps, the module completed the process from... arrive The model parameters evolve at each time step. The updated Koopman operator matrix. The results are immediately fed back to the robustness rating module for the next round of quality prediction and grading. This mechanism eliminates the cumulative error of static models in long-term monitoring, ensuring that the evaluation system can dynamically adapt to the individual developmental specificity of germplasm resources.
[0189] Example 2: Application in the selection of stress-resistant Panax notoginseng germplasm in climate-unstable regions. Panax notoginseng has a long growth cycle of up to three years and is extremely sensitive to light, temperature, water, and soil conditions. When introduced to new production areas, it often faces unstable climate shocks such as late spring frosts or extreme droughts. In this scenario, existing technologies face four core challenges: First, the "spatiotemporal misalignment" of multi-source heterogeneous data leads to the failure of matching environmental factors with phenological periods based on absolute calendar time analysis, resulting in distorted feature extraction. Second, traditional black-box models cannot capture the nonlinear "toxic stimulant effect" of secondary metabolites under stress, making it difficult to predict the accumulation of medicinal effects under complex environmental interactions. Third, mean-based evaluation based on the assumption of normal distribution masks the long-tail risk of extreme climates, resulting in the selected "superior varieties" lacking robustness against sudden disasters. Fourth, static models cannot adapt to the "conceptual drift" of plant sensitivity to the environment during long growth cycles, leading to a significant decrease in the accuracy of later evaluations. To solve these problems, this invention provides a big data-based evaluation system for the grading of medicinal herb germplasm resources, the structure of which is as follows: Figure 1 As shown. The specific implementation process of this system is as follows:
[0190] The system first utilizes a dynamic time warping algorithm to address the spatiotemporal misalignment of phenological periods in multi-source heterogeneous data, and integrates genomic and environmental characteristics. Then, it constructs reaction-diffusion partial differential equations to characterize the nonlinear metabolic mechanism of active ingredients within the plant. By employing Koopman operator theory, this nonlinear model is transformed into a global linear discrete state equation, significantly reducing computational complexity. Simultaneously, a fuzzy sphere covering the long-tail risk of extreme climates is constructed using the regeneration kernel Hilbert space and the maximum average difference distance. Based on this, a two-layer optimization model is established, and a semidefinite programming relaxation technique is used to solve for the minimum quality index of germplasm under the most severe environmental distribution for robust grading. Finally, a recursive least squares method is introduced to incrementally correct the model parameters online. This invention effectively overcomes the technical challenges of data drift, mechanism black box, and lack of stress resistance risk assessment in traditional evaluation methods, achieving accurate dynamic evaluation throughout the entire life cycle.
[0191] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A big data-based evaluation system for the grading of germplasm resources of Chinese medicinal herbs, characterized in that, Includes the following modules: The multi-source heterogeneous data fusion and feature mapping module is used to align environmental time-series data and phenotypic data with phenological periods using dynamic time warping algorithm, and to embed genomic SNP site data to generate high-dimensional feature vectors. The nonlinear metabolic kinetics modeling module is used to construct reaction-diffusion partial differential equations that include environmentally driven convection-diffusion terms and gene-regulated reaction source terms, characterizing the spatiotemporal evolution of effective components. The global linearization module is used to solve the Koopman operator matrix using the extended dynamic mode decomposition algorithm, and to convert the partial differential equation into a linear discrete state equation in the observation function space. An environmental uncertainty building module is used to map the environmental distribution to a regenerating kernel Hilbert space using kernel mean embedding, and to construct a sub-Bruker fuzzy sphere based on the maximum average difference distance; The robustness rating module is used to construct a two-layer optimization model, solve for the minimum quality index under the worst distribution of the sub-Bruker fuzzy sphere through semidefinite programming relaxation, and then determine the rating. The phenotypic trajectory tracking and model evolution correction module is used to monitor the actual growth trajectory residuals and to perform online incremental updates of the Koopman operator matrix using the recursive least squares method.
2. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The multi-source heterogeneous data fusion and feature mapping module includes: The complete growth cycle sequence of historical high-yield and high-quality germplasm was selected as the benchmark reference template, and the real-time monitoring sequence of the germplasm to be evaluated was used as the query sequence. Calculate the Euclidean distance between each time point in the query sequence and each time point in the benchmark template, and construct the cumulative distance matrix; Search for a regular path from the starting point to the ending point in the cumulative distance matrix such that the cumulative Euclidean distance on the path is minimized, and then perform nonlinear stretching or compression on the environmental time series data based on the regular path. Single nucleotide polymorphism sites in genomic data are encoded one-hot to generate static genotype vectors; The aligned dynamic environment time-series vector is concatenated and spliced with the static genotype vector to generate a hybrid high-dimensional feature vector containing time-varying and time-invariant features.
3. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The nonlinear metabolic kinetics modeling module includes: Based on the principles of fluid mechanics, the convective heat transfer coefficient is calculated using the Nusselt number correlation, which includes the Reynolds number and Prandtl number. A nonlinear convective diffusion term driven by real-time wind speed and temperature fields is constructed to characterize the material transport flux of effective components in plant tissues. The rate of enzyme-catalyzed reaction is described by the Michaelis-Menten kinetic equation, and the genotype feature vector is mapped to the maximum reaction rate parameter and the Michaelis constant to construct a gene-regulated metabolic synthesis nonlinear source term. A first-order kinetic decay coefficient, corrected by the environmental stress intensity index, is introduced to construct a reaction sink term characterizing the natural decomposition process of metabolites; By superimposing the nonlinear convection-diffusion term, the nonlinear metabolic synthesis source term, and the reaction sink term, a partial differential equation is established to describe the changes in the concentration field of the effective component with time and spatial location along the plant growth axis.
4. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The global linearization processing module includes: Construct an observation function dictionary containing the original state variables, polynomial combinations of the original state variables, and nonlinear functions based on radial basis kernel functions; The low-dimensional state vector after discretization of the reaction-diffusion partial differential equation is input into the observation function dictionary to calculate the high-dimensional enhanced observation vector with a dimension higher than that of the original state space. The completeness of the high-dimensional enhanced observation vector in the Hilbert space is verified, so that the evolution trajectory of the system dynamics in this space can be approximated by linear operators.
5. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The global linearization processing module also includes: Construct an input data matrix containing the current time step boosted observation vector, and an output data matrix containing the next time step boosted observation vector; Construct a least-squares optimization model with the objective function being the left multiplication of the Koopman operator matrix by the Frobenius norm, which is the difference between the input and output data matrices. The Moore-Penrose pseudo-inverse matrix of the input data matrix is calculated using the singular value decomposition method. Multiply the output data matrix by the Moore-Penrose pseudo-inverse matrix to obtain the optimal approximation finite-dimensional Koupman operator matrix, and establish an explicit linear discrete state-space equation.
6. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The environmental uncertainty construction module includes: We select the Gaussian radial basis function as the positive definite kernel function and define the inner product operation rules for the reproducing kernel Hilbert space; For each historical environmental data sample, feature mapping is used to transform it into elements in a high-dimensional feature space; Calculate the arithmetic mean of the corresponding elements of all historical environmental data samples in the feature space to obtain the kernel mean embedding vector that represents the distribution of historical experience; By utilizing the norm definition of the reproducing kernel Hilbert space, the maximum average difference distance between two kernel mean embedding vectors with different probability distributions is quantified.
7. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The environmental uncertainty construction module also includes: The kernel mean embedding vector of the calculated historical experience distribution is used as the center of the sphere; Based on the preset confidence level and the number of historical samples, the statistical deviation threshold required to cover the true distribution is calculated using the central inequality, and this threshold is set as the radius of the fuzzy sphere. Construct a set of probability distributions containing all fuzzy spheres whose maximum average difference distance from the sphere center is less than or equal to the radius of the fuzzy sphere, and use this set as the environmental stress sub-Blu-ray fuzzy sphere.
8. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The robustness level determination module includes: Define the problem as maximizing the expected value of the final cumulative amount of effective components as the objective function; In the inner layer optimization, the worst environment distribution that minimizes the objective function value is searched using the fuzzy sphere of the blobs as the feasible region. In the outer layer optimization, the quality index bound that the global linear state-space model can achieve under the worst environmental distribution conditions is calculated; The linear evolution equations and control input constraints of the global linear state-space model are used as all the physical constraints of the two-layer optimization model.
9. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The robustness level determination module also includes: Construct a symmetric semidefinite matrix variable defined by the outer product of the lifting observation vector and its transpose; By utilizing the properties of matrix trace operations, the nonlinear objective function and constraints in the original optimization problem are transformed into linear functions of the symmetric semidefinite matrix variables; The non-convex rank constraint of the symmetric semidefinite matrix variable is relaxed, thereby transforming the original problem into a convex optimization problem on a semidefinite cone. The semidefinite programming problem is solved using the original dual interior point method to obtain the global optimal matrix, and the minimum quality index is restored from the global optimal matrix based on the spectral decomposition technique.
10. The big data-based evaluation system for the grading of Chinese medicinal herb germplasm resources according to claim 1, characterized in that, The phenotypic trajectory tracking and model evolution correction module includes: Set a forgetting factor to exponentially decay the weights of historical data; At each growth monitoring time, the prediction error vector between the actual observed vector and the predicted vector based on the current Koopman operator matrix is calculated. Calculate the gain matrix based on the covariance matrix of the previous time step and the current observation vector; The current Koopman operator matrix is recursively updated using the gain matrix and the prediction error vector. The covariance matrix is updated for calculation in the next time step, completing the online adaptive correction of the operator matrix.