Dynamic identification method and system for biotechnology product risks based on multi-index fusion

By constructing a dynamic data set through multi-indicator fusion and machine learning algorithms, the dynamic and complexity of the risk behavior of biotechnology products is solved, the accurate identification and management of risks within the life cycle of genetically modified crops is achieved, and the accuracy of risk assessment and the targetedness of control measures are improved.

CN120611983BActive Publication Date: 2025-10-03RICE RES ISTITUTE ANHUI ACAD OF AGRI SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511117740.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-03
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively identify and manage the dynamics and complexity of the risk behavior of biotechnology products throughout their life cycle, especially when there is an interactive impact of risk characteristics between different types of products, resulting in insufficient accuracy and applicability of assessment results.

Method used

A multi-indicator fusion method is used to construct a dynamic data set through machine learning algorithms. Combined with time series analysis and multiple algorithms (such as random forest, principal component analysis, support vector machine, etc.), the mutation frequency, functional abnormality fluctuations and ecological impact factor trends of biotechnology products are extracted, the association between nucleotide substitution patterns and functional abnormalities is identified, a dynamic risk feature portrait is generated and risk control measures are optimized.

Benefits of technology

It has achieved accurate identification and dynamic management of risks of biotechnology products, improved the accuracy of risk prediction and the safety of environmental release, and provided highly targeted risk control measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611983B_ABST
    Figure CN120611983B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for dynamically identifying risks of biotechnology products based on multi-indicator fusion, which relates to the field of biosafety management technology, including: S1, obtaining experimental data and environmental data from various stages of the life cycle of genetically modified crops, constructing a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, and extracting the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators and dynamic change trends of ecological impact factors; S2, based on the mutation frequency of molecular structure sequences and abnormal fluctuations in functional performance indicators, extracting nucleotide substitution patterns and functional abnormality data sets of gene editing products in the research and development stage and the application stage, and determining the correlation pattern between nucleotide substitution patterns and functional abnormalities; this method and system for dynamically identifying risks of biotechnology products based on multi-indicator fusion realizes accurate prediction and dynamic management of ecological risks of genetically modified crops, and significantly improves the safety of environmental release.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biosafety management technology, and in particular to a method and system for dynamically identifying risks of biotechnology products based on multi-indicator fusion. Background Art

[0002] Biotechnology products are increasingly used in healthcare, agriculture, and industry. Their safety and risk management are directly related to human health, ecological balance, and industrial development. However, due to the complexity and diversity of biotechnology products, accurately identifying and managing their risk profiles has become a key challenge.

[0003] Currently, risk assessments rely heavily on traditional statistical methods and expert experience. These methods often struggle to capture the dynamics and underlying patterns of biotechnology products, given their diverse characteristics. In particular, traditional methods struggle to account for the heterogeneity of risk behavior patterns throughout the product lifecycle and the interplay of risk profiles across different product types, resulting in inaccurate and inapplicable assessment results. A key challenge lies in effectively analyzing the complex patterns of risk behavior in biotechnology products. First, the dynamic nature of risk behavior makes single-point assessments incapable of reflecting the potential risks of a product at different stages. For example, the risk of a gene-edited product may manifest as uncertainty about gene mutations during the R&D phase, but may shift to uncertainty about ecological impacts during the application phase. Second, this dynamic nature further exacerbates the complexity of interactions between risk profiles. For example, changes in the molecular structure of a gene-edited product may lead to functional abnormalities, which in turn may amplify ecological risks following release. This interactive complexity makes it difficult for traditional methods to construct a comprehensive risk profile, especially when tailoring risk management strategies for different product types, often leading to assessment bias or strategy mismatch. For example, a genetically modified crop may have enhanced insect resistance due to adjustments in its molecular structure, but at the same time may cause unexpected ecological impacts. Traditional assessments make it difficult to accurately predict such interaction risks. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for dynamic identification of biotechnology product risks based on multi-indicator fusion, which effectively captures the dynamic patterns of biotechnology product risk behavior through machine learning algorithms and analyzes the complex interactive relationships between their characteristics.

[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a method for dynamic identification of risks of biotechnology products based on multi-indicator fusion, comprising: S1, obtaining experimental data and environmental data from each stage of the life cycle of genetically modified crops, constructing a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, extracting the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators and dynamic change trends of ecological impact factors; S2, extracting nucleotide substitution patterns and functional abnormality data sets of gene editing products in the research and development stage and the application stage based on the mutation frequency of molecular structure sequences and abnormal fluctuations in functional performance indicators, and determining the association pattern between nucleotide substitution patterns and functional abnormalities; S3, extracting the potential effect characteristics of ecological impact factors after the release of functional abnormalities on the environment from the association pattern of nucleotide substitution patterns and functional abnormalities, and obtaining the interaction feature set between functional abnormalities and ecological impact factors; S4, extracting the genetic abnormalities from the interaction feature set. The nucleotide substitution patterns and ecological impact factor data related to the insect resistance of genetically modified crops are used to determine the probability of unexpected ecological impacts caused by enhanced insect resistance; S5. If the probability of unexpected ecological impacts caused by enhanced insect resistance exceeds the preset threshold, the key factors of the ecological impact factors after environmental release are extracted from the interaction feature set to construct a dynamic risk feature portrait data set including mutation frequency, abnormal fluctuations and ecological factor interactions; S6. Based on the dynamic risk feature portrait data set, the similarity of nucleotide substitution patterns and ecological factor interactions is calculated to obtain the risk classification results of insect-resistant genetically modified crops; S7. The nucleotide substitution patterns and ecological factor interaction characteristics are extracted from the risk classification results to determine the priority of risk control measures for mutation frequency and abnormal fluctuations; S8. Based on the priority of risk control measures, combined with the mutation frequency and abnormal fluctuation data of the life cycle stage, a dynamically adjusted risk control plan including nucleotide substitution constraints and ecological factor monitoring is generated.

[0006] Preferably, S1 includes obtaining experimental data and environmental data from various stages of the transgenic crop life cycle, constructing a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, storing the data in a time series format, and obtaining an initial data set; using a time series analysis method to process the molecular structure sequences in the initial data set, calculating the mutation frequency at each time point, and obtaining a mutation frequency sequence; using a statistical analysis method to perform anomaly detection on the mutation frequency sequence, and if the mutation frequency exceeds a preset threshold, marking it as an abnormal point to obtain an abnormal mutation point set; using a sliding window method to detect the fluctuation of the indicator value based on the time series data of the functional performance indicator, and if the fluctuation amplitude exceeds a preset threshold, marking it as an abnormal fluctuation to obtain an abnormal fluctuation point set; obtaining time series data of the ecological impact factor, using a trend analysis method to extract the changing trend of the factor value, and obtaining an ecological impact trend sequence; using an association analysis method to fuse the abnormal mutation point set, the abnormal fluctuation point set and the ecological impact trend sequence, and construct a dynamic relationship model between the three to obtain a comprehensive dynamic relationship data set; using a machine learning algorithm random forest to perform classification prediction on the comprehensive dynamic relationship data set, determine the potential connection between the molecular structure mutation and the abnormal functional performance and the ecological impact trend, and obtain a classification result.

[0007] Preferably, S2 includes: obtaining nucleotide sequence data and functional performance data from gene-edited products of transgenic crops, constructing an initial data set in a time series format, including nucleotide substitution patterns and functional abnormality records, and obtaining an initial time series data set; using a time series analysis method to process the nucleotide sequences in the initial time series data set, calculating the nucleotide substitution frequencies in the R&D stage and the application stage, and obtaining a substitution frequency sequence; using a sliding window method to analyze the fluctuation characteristics of the substitution frequency sequence, and if the fluctuation amplitude exceeds a preset threshold, marking it as a high-frequency substitution point, and obtaining a high-frequency substitution point set; for the time series of functional performance data, Statistical analysis methods are used to detect functional abnormality points. If the functional value deviates from the preset threshold, it is marked as an abnormal point to obtain a functional abnormality point set; the Gini index analysis method is used to fuse the high-frequency substitution point set and the functional abnormality point set, and the association strength between nucleotide substitution frequency and functional abnormality is calculated to obtain an association strength data set; the random forest algorithm is used to classify the association strength data set, determine the potential connection between nucleotide substitution patterns and functional abnormalities, and obtain a classification result data set; the trend analysis method is used to extract trends in the time dimension of the classification result data set, analyze the dynamic change relationship between substitution patterns and functional abnormalities, and obtain a dynamic trend data set.

[0008] Preferably, S3 includes: obtaining the original data set of functional abnormality records and ecological impact factors from the association data of the nucleotide substitution pattern and functional abnormality of the gene editing product, including the nucleotide substitution pattern, functional abnormality mark and ecological impact factor value, to obtain an initial interaction data set; using a standardization processing method, normalizing the ecological impact factor values ​​in the initial interaction data set, eliminating dimensional differences, and obtaining a standardized interaction data set; using a principal component analysis method, performing dimensionality reduction processing on the functional abnormality records and ecological impact factor values ​​in the standardized interaction data set, extracting the main interaction features, and obtaining a principal component feature set; if the feature contribution rate in the principal component feature set exceeds a preset threshold, it is marked as a high correlation feature to obtain a high correlation feature set; based on the high correlation feature set, using a cluster analysis method, grouping the interaction patterns between functional abnormalities and ecological impact factors to obtain an interaction pattern set; using a time series analysis method, performing dynamic change detection on the interaction pattern set, analyzing the long-term effect trend of functional abnormalities on ecological impact factors, and obtaining a dynamic effect trend set; using a visualization processing method, graphically displaying the dynamic effect trend set to determine the interaction law between functional abnormalities and ecological impact factors.

[0009] Preferably, the S4 includes: obtaining an original data set including nucleotide substitution sites, insect resistance markers and ecological impact factor values ​​from the nucleotide substitution patterns and ecological impact factor data of transgenic crops, normalizing the ecological impact factor values ​​using the Z-score normalization method to obtain a standardized data set; based on the standardized data set, using the support vector machine algorithm, optimizing the classification boundaries through the radial basis function kernel and the regularization parameter, calculating the probability of unexpected ecological impact caused by enhanced insect resistance, and obtaining a classification probability set; if a probability value in the classification probability set exceeds a preset threshold, it is marked as a high-risk impact to obtain a high-risk impact set; if the probability value is lower than the threshold, it is marked as a low-risk impact to obtain a low-risk impact set; based on the According to the high-risk impact set, the K-means clustering algorithm was used to group nucleotide substitution patterns and ecological influencing factors, extract potential interaction patterns, and obtain the interaction pattern set; through the time series analysis method, the dynamic change detection of the nucleotide substitution patterns of high-risk impacts in the interaction pattern set was carried out, and its long-term correlation trend with the ecological influencing factors was analyzed to obtain the dynamic correlation trend set; the visualization processing method was used to graphically display the dynamic correlation trend set, determine the long-term interaction law between enhanced insect resistance and ecological influencing factors, and obtain the interaction law set; highly correlated nucleotide substitution patterns were extracted from the interaction law set, and the decision tree algorithm was used to construct a prediction model for enhanced insect resistance and unexpected ecological impacts to obtain the prediction model set.

[0010] Preferably, S5 includes: extracting ecological impact factors related to enhanced insect resistance from the interaction feature set, using the principal component analysis method to perform dimensionality reduction processing on the ecological impact factors to obtain a key factor set; if the feature contribution rate of the key factor set exceeds a preset threshold, using the cluster analysis method to group the key factor set to obtain an ecological factor interaction set; based on the ecological factor interaction set, calculating the statistical characteristics of mutation frequency and abnormal fluctuations to generate a dynamic risk feature set; using the dynamic risk feature set, using the random forest algorithm to classify the dynamic changes of ecological factor interactions to obtain a risk classification result; based on the risk classification result, performing time series decomposition on the dynamic risk feature set to extract abnormal fluctuation trends; based on the abnormal fluctuation trends, constructing a dynamic risk feature portrait data set to generate a feature portrait set; using the feature portrait set, using a visualization processing method to generate a dynamic risk distribution map to determine the dynamic interaction law between enhanced insect resistance and unexpected ecological impacts.

[0011] Preferably, the S6 includes: extracting the mutation frequency and abnormal fluctuation pattern of genetically modified crops from the dynamic risk feature portrait data set, using the K-means cluster analysis method, and calculating the similarity matrix of the interaction between the nucleotide substitution pattern and the ecological factor based on the Euclidean distance to obtain a preliminary risk classification result; based on the preliminary risk classification result, extracting the mutation frequency and abnormal fluctuation pattern corresponding to the high-risk category, using the time series decomposition method to calculate the periodic change trend of the mutation frequency, and obtaining a periodic fluctuation feature set; based on the periodic fluctuation feature set, calculating the interaction strength between the ecological factor and the mutation frequency, using the Pearson correlation coefficient analysis method to generate an interaction strength matrix, and obtaining a set of key influencing factors for the interaction of ecological factors. ; If the interaction intensity of the key influencing factor set exceeds the preset threshold, a multidimensional scaling analysis is performed on the mutation frequency and abnormal fluctuation pattern to generate a two-dimensional mapping of the risk distribution and obtain a risk distribution feature set; based on the risk distribution feature set, the support vector machine algorithm is used to classify the risk distribution characteristics, generate the boundary division of high-risk and low-risk categories, and obtain the classification boundary parameters; based on the classification boundary parameters, the nucleotide substitution pattern corresponding to the risk category is extracted, and the interaction similarity matrix between the nucleotide substitution pattern and the ecological factor is calculated to generate a dynamic interaction feature set; based on the dynamic interaction feature set, a visualization processing method is used to generate a dynamic change diagram of the risk distribution of insect-resistant crops to determine the dynamic evolution law of the risk distribution.

[0012] Preferably, S7 includes: obtaining nucleotide substitution patterns and ecological factor interaction characteristics from risk categories, and using a feature extraction method to generate a feature data set containing mutation frequency and abnormal fluctuations; based on the feature data set, using a decision tree algorithm, calculating the splitting path of each feature through the Gini index, and generating an initial decision tree model; if the splitting path complexity of the initial decision tree model exceeds a preset threshold, then using a post-pruning strategy to remove splitting nodes with low contribution to generate an optimized decision tree model; based on the optimized decision tree model, extracting the splitting path characteristics corresponding to the mutation frequency and abnormal fluctuations, and determining the preliminary priority of the risk control measures; obtaining the nucleotide substitution pattern corresponding to the priority measures with a priority higher than the preset threshold from the preliminary priority, using the Pearson correlation coefficient analysis method to calculate the correlation matrix of the interaction with the ecological factors, and generating an interaction feature set; based on the interaction feature set, using a visualization processing method to generate a dynamic distribution map of the interaction between mutation frequency and ecological factors, and determine the dynamic priority of the risk control measures; if the priority measures in the dynamic priority do not match the preset threshold, then performing a secondary analysis on the interaction feature set, using the principal component analysis method to extract key risk features, and determine the final risk control measure priority.

[0013] Preferably, S8 includes: obtaining mutation frequency characteristics and abnormal fluctuation characteristics from priority risk control measures that are higher than a preset threshold, and generating an initial feature data set in combination with the life cycle stage; analyzing the correlation between the initial feature data set and the life cycle stage through a logistic regression algorithm to obtain a preliminary dynamic adjustment plan framework; if the complexity of the plan framework exceeds the preset threshold, the principal component analysis method is used to extract the key components in the initial feature data set to generate a feature data set; based on the feature data set, the correlation matrix between nucleotide substitution and ecological factor monitoring is calculated to obtain an interactive feature set; the cluster analysis method is used to group the interactive feature set to generate dynamic distribution characteristics of the interaction between mutation frequency and ecological factors to obtain a dynamic distribution pattern; if the priority measures in the dynamic distribution pattern do not match the preset threshold, the interactive feature set is processed secondary, and the weighted average method is used to calculate the weight of each feature to obtain the final dynamic adjustment plan framework; based on the final dynamic adjustment plan framework, a dynamic adjustment risk control plan including nucleotide substitution constraints and ecological factor monitoring is generated to determine the priority ranking of the final plan.

[0014] A biotechnology product risk dynamic identification system based on multi-indicator fusion is used to implement the steps of the biotechnology product risk dynamic identification method based on multi-indicator fusion, the system comprising: a data acquisition module, which acquires experimental data and environmental data from each stage of the life cycle of genetically modified crops, constructs a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, extracts the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators and dynamic change trends of ecological impact factors; a stage feature analysis module, which extracts nucleotide substitution patterns and functional abnormality data sets of gene editing products in the research and development stage and the application stage according to the mutation frequency of molecular structure sequences and abnormal fluctuations in functional performance indicators, and determines the association pattern between nucleotide substitution patterns and functional abnormalities; an interactive feature extraction module, which extracts the potential effect characteristics of ecological impact factors after the release of functional abnormalities on the environment from the association pattern of nucleotide substitution patterns and functional abnormalities, and obtains an interactive feature set between functional abnormalities and ecological impact factors; an ecological risk identification module, which extracts the potential effect characteristics of ecological impact factors after the release of functional abnormalities from the interaction pattern between nucleotide substitution patterns and functional abnormalities, and obtains an interactive feature set between functional abnormalities and ecological impact factors; and an ecological risk identification module, which extracts the potential effect characteristics of ecological impact factors after the release of functional abnormalities from the interaction pattern between nucleotide substitution patterns and functional abnormalities. The mutual feature set extracts nucleotide substitution patterns and ecological impact factor data related to the insect resistance of transgenic crops to determine the probability of unexpected ecological impact caused by enhanced insect resistance; the dynamic portrait construction module extracts key factors of ecological impact factors after environmental release from the interactive feature set if the probability of unexpected ecological impact caused by enhanced insect resistance exceeds the preset threshold, and constructs a dynamic risk feature portrait data set including mutation frequency, abnormal fluctuation and ecological factor interaction; the risk classification module calculates the similarity of nucleotide substitution patterns and ecological factor interactions based on the dynamic risk feature portrait data set to obtain the risk classification results of insect-resistant transgenic crops; the strategy priority determination module extracts nucleotide substitution patterns and ecological factor interaction characteristics from the risk classification results to determine the priority of risk control measures for mutation frequency and abnormal fluctuation; the risk control generation module generates a dynamically adjusted risk control plan including nucleotide substitution constraints and ecological factor monitoring based on the risk control measure priority and combined with the mutation frequency and abnormal fluctuation data of the life cycle stage.

[0015] It can be seen from the above technical solution that the present invention has the following beneficial effects:

[0016] This method and system for dynamic identification of risks of biotechnology products based on multi-index fusion solves the core problem of how to accurately identify and control the unexpected ecological impacts caused by enhanced insect resistance during the development and environmental release of transgenic crops. The present invention constructs a dynamic data set, integrates molecular structure sequences, functional performance indicators and ecological impact factors, and uses time series analysis to extract mutation frequencies, abnormal fluctuations and ecological trends; uses random forest and principal component analysis to explore the association between nucleotide substitution patterns and functional abnormalities and ecological impact characteristics; optimizes classification boundaries through support vector machines to evaluate the ecological risk probability of enhanced insect resistance; for high-risk scenarios, cluster analysis and decision tree algorithms further identify risk patterns and optimize the priority of control measures, and finally generate a dynamically adjusted risk control plan. The present invention achieves accurate prediction and dynamic management of the ecological risks of transgenic crops through multi-algorithm collaboration and full life cycle data integration, significantly improving the safety of environmental release. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Flow chart of the method of the present invention;

[0018] Figure 2 This is a system connection diagram of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] like Figure 1As shown, the present invention provides a technical solution: a method for dynamic identification of biotechnology product risks based on multi-indicator fusion, comprising: S1, obtaining experimental data and environmental data from each stage of the life cycle of genetically modified crops, constructing a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, extracting the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators and dynamic change trends of ecological impact factors; S2, extracting nucleotide substitution patterns and functional abnormality data sets of gene editing products in the research and development stage and the application stage based on the mutation frequency of molecular structure sequences and abnormal fluctuations in functional performance indicators, and determining the association pattern between nucleotide substitution patterns and functional abnormalities; S3, extracting the potential effect characteristics of ecological impact factors after the release of functional abnormalities on the environment from the association pattern of nucleotide substitution patterns and functional abnormalities, and obtaining the interaction feature set between functional abnormalities and ecological impact factors; S4, extracting the genetically modified crops from the interaction feature set. The nucleotide substitution patterns and ecological impact factor data related to the insect resistance of crops are used to determine the probability of unexpected ecological impact caused by enhanced insect resistance; S5. If the probability of unexpected ecological impact caused by enhanced insect resistance exceeds the preset threshold, the key factors of the ecological impact factors after environmental release are extracted from the interaction feature set to construct a dynamic risk feature portrait data set including mutation frequency, abnormal fluctuation and ecological factor interaction; S6. Based on the dynamic risk feature portrait data set, the similarity of nucleotide substitution patterns and ecological factor interactions is calculated to obtain the risk classification results of insect-resistant transgenic crops; S7. The nucleotide substitution patterns and ecological factor interaction characteristics are extracted from the risk classification results to determine the priority of risk control measures for mutation frequency and abnormal fluctuation; S8. Based on the priority of risk control measures, combined with the mutation frequency and abnormal fluctuation data of the life cycle stage, a dynamically adjusted risk control plan including nucleotide substitution constraints and ecological factor monitoring is generated.

[0021] This method constructs a dynamic dataset integrating multi-source data to systematically analyze the risk evolution of transgenic crops at different lifecycle stages. The method first collects data on molecular structure sequences, functional performance indicators, and ecological and environmental impact factors from each stage of transgenic crop development to environmental release. A time series data model is then established to quantify mutation frequencies, fluctuations in functional abnormalities, and trends in ecological impact. Subsequently, pattern recognition and data mining techniques are used to extract correlations between nucleotide substitution patterns and functional abnormalities. The potential impacts of functional abnormalities on environmental ecological factors are further analyzed to form an interactive feature set. This feature set reveals how functional imbalances caused by specific gene mutations affect key ecosystem factors (such as non-target insect abundance and microbial diversity). Based on this, dynamic risk feature profiles are constructed and similarity calculated to classify and identify potential ecological risks of transgenic crops. Finally, based on the risk classification results and the importance ranking of interactive features, the corresponding risk control measures are prioritized. Combined with lifecycle dynamics information, a dynamically adjusted risk control plan, including nucleotide substitution restrictions and ecological monitoring strategies, is generated, thus enabling intelligent risk management and precise regulation of insect-resistant transgenic crops throughout their lifecycle.

[0022] Compared with existing static and phased risk assessment methods, the method of the present invention has significant technical advantages and practical application value. First, by introducing time series modeling and multi-index fusion mechanism, dynamic collection and evolutionary analysis of genetically modified crop risk information is realized, overcoming the limitation that traditional methods cannot reflect the changes of risks with the life cycle. Secondly, based on the interactive modeling between nucleotide substitution patterns, functional abnormalities and ecological impact factors, it can identify potential unexpected ecological impact paths and improve the accuracy and scientificity of risk prediction. Thirdly, by constructing dynamic risk feature portraits and similarity analysis mechanisms, risk grading and classification of different insect-resistant genetically modified crops are realized, making risk management measures targeted and operational. Finally, the dynamically adjusted risk control plan output by this method not only takes into account the mutation frequency and functional fluctuations of the current state, but also integrates life cycle trend information. It is forward-looking and real-time, which helps to establish an accurate and efficient biotechnology product safety supervision system.

[0023] S1 includes obtaining experimental data and environmental data from various stages of the genetically modified crop life cycle, constructing a dynamic data set containing molecular structure sequences, functional performance indicators and ecological impact factors, storing them in a time series format, and obtaining an initial data set; using a time series analysis method to process the molecular structure sequences in the initial data set, calculating the mutation frequency at each time point, and obtaining a mutation frequency sequence; using a statistical analysis method to detect anomalies in the mutation frequency sequence, and if the mutation frequency exceeds a preset threshold, it is marked as an abnormal point, and an abnormal mutation point set is obtained; based on the time series data of the functional performance indicators, a sliding window method is used to detect the fluctuation of the indicator values, and if the fluctuation amplitude exceeds a preset threshold, it is marked as an abnormal fluctuation, and an abnormal fluctuation point set is obtained; obtaining the time series data of the ecological impact factors, using a trend analysis method to extract the changing trend of the factor values, and obtaining an ecological impact trend sequence; using an association analysis method, the abnormal mutation point set, the abnormal fluctuation point set and the ecological impact trend sequence are integrated to construct a dynamic relationship model between the three, and obtain a comprehensive dynamic relationship data set; using the machine learning algorithm random forest, classification and prediction are performed on the comprehensive dynamic relationship data set, judging the potential connection between molecular structure mutations and functional performance anomalies and ecological impact trends, and obtaining classification results.

[0024] In this embodiment, the life cycle of a genetically modified crop is first divided into four stages: research and development, field trials, small-scale promotion, and commercial cultivation. Within each stage, three types of data are collected simultaneously on a quarterly basis: molecular structure sequence data, functional performance indicator data, and ecological impact factor data. All collected data are archived using timestamps as key values ​​and constructed into an initial time series dataset containing a time field. For molecular structure sequence data, high-throughput sequencing is used to obtain the nucleotide sequence of a specific gene segment. At each time point, the sequence at the current time point is compared one-to-one with the sequence at the previous time point. The number of base changes is counted and recorded. The number of base changes is divided by the total number of bases in the sequence to obtain the mutation frequency at the current time point. The mutation frequencies at all time points are used to form a mutation frequency sequence. Next, the mean and standard deviation of the mutation frequency sequence are calculated. The mean is the sum of the mutation frequencies at all time points, divided by the total number of time points. The standard deviation is the sum of the squares of the differences between the mutation frequencies at each time point and the mean, divided by the total number of time points, and then the square root is taken. The abnormal mutation frequency threshold is set to the mean plus 2 times the standard deviation. For example, if the mean is 0.3% and the standard deviation is 0.15%, the threshold is 0.6%. Time points exceeding this threshold are marked as abnormal mutation points. The set of all time points that meet this condition is the abnormal mutation point set.

[0025] For functional performance indicator data, such as protein expression or field insecticide mortality, a time series is constructed. A sliding window processing method is used, with the window width set to 7 days. Each window slides back 1 day, and the difference between the maximum and minimum values ​​of the indicator value in each window is calculated to obtain the fluctuation amplitude of the window. After traversing all time periods, the fluctuation amplitudes of all windows are summarized and the 95th percentile is taken as the abnormal fluctuation threshold. For example, if the fluctuation amplitudes of all windows are arranged in ascending order and the 95th percentile corresponds to 10, then when the fluctuation amplitude of a window is greater than 10, the middle point of the window is marked as an abnormal fluctuation point. Repeat this process to form a set of all abnormal fluctuation points, that is, the abnormal fluctuation point set.

[0026] For ecological impact factor data, such as non-target insect populations or microbial diversity indices, data are collected monthly and a time series is constructed. The first-order difference method is used to calculate the difference between the current and previous time points. A positive difference indicates an upward trend, a negative difference indicates a downward trend, and a zero difference indicates a stable trend. Ultimately, an ecological trend series corresponding to each time point is formed.

[0027] The three datasets were time-aligned to ensure that the data for each time point in the three series were complete and corresponded to each other. A comprehensive dynamic relationship dataset was then constructed. Each record in the dataset contained information about whether the time point was an anomalous mutation, an anomalous fluctuation, and the corresponding ecological trend type. This dataset was then used as input for training a random forest classification model. During training, time points with significant ecological anomalies were first classified as positive based on manual annotation or historical expert assessment, while the remaining time points were classified as negative. A random forest classifier consisting of 100 decision trees was then constructed. Each tree was trained on 60% of the original sample set. At each node split, two features were randomly selected from the total number of features to determine the best split. After model training, the test sample was fed into the model, and the model output was a classification result for each time point, indicating whether there was a potential ecological risk. A positive classification result indicated that the mutation or functional fluctuation at that time point was likely to have an ecological impact, and a score was also output for the importance of each feature in the classification decision. Finally, by analyzing the importance of features and classification results, we can determine whether the molecular mutation frequency and functional fluctuations are significant and can lead to changes in ecosystem trends, providing basic support for subsequent risk level assessment and risk response strategies.

[0028] S2 includes: obtaining nucleotide sequence data and functional performance data from gene-edited products of genetically modified crops, constructing an initial data set in a time series format, including nucleotide substitution patterns and functional abnormality records, to obtain an initial time series data set; using a time series analysis method to process the nucleotide sequences in the initial time series data set, calculating the nucleotide substitution frequencies in the research and development stage and the application stage, to obtain a substitution frequency series; using a sliding window method to analyze the fluctuation characteristics of the substitution frequency series, and if the fluctuation amplitude exceeds a preset threshold, it is marked as a high-frequency substitution point to obtain a high-frequency substitution point set; using a statistical analysis method to detect functional abnormality points in the time series of functional performance data, and if the functional value deviates from the preset threshold, it is marked as an abnormal point to obtain a functional abnormality point set; using a Gini index analysis method to fuse the high-frequency substitution point set and the functional abnormality point set, calculate the correlation strength between the nucleotide substitution frequency and the functional abnormality, to obtain a correlation strength data set; using a random forest algorithm to classify the correlation strength data set, determine the potential connection between the nucleotide substitution pattern and the functional abnormality, to obtain a classification result data set; using a trend analysis method to extract trends in the time dimension of the classification result data set, analyze the dynamic change relationship between the substitution pattern and the functional abnormality, to obtain a dynamic trend data set.

[0029] In this embodiment, nucleotide sequence data and functional performance data are first obtained from the gene-edited products of genetically modified crops in chronological order to construct an initial time series data set. The nucleotide sequence data are obtained by high-throughput sequencing technology to obtain the complete sequence of a specific site, and the functional performance data include the expression level of insect-resistant protein, growth rate, mortality rate, etc. Each data point corresponds to a specific time point, for example, sampling is performed every 7 days to form a structured data set with a timestamp. When processing the nucleotide sequence data, the sequences of two adjacent time points are first compared base by base, and the number of base change positions is counted. For example, if the sequence length is 3000 bases and 60 sites are replaced, the replacement frequency is 60 divided by 3000, which is 2%. The replacement frequencies of all time points are calculated in sequence to form a replacement frequency sequence. Subsequently, the sliding window analysis method is used to process the replacement frequency sequence. The sliding window width is set to 5 time points, and it slides forward 1 time point each time. The maximum replacement frequency minus the minimum replacement frequency in the window is calculated to obtain the fluctuation amplitude of each window. Arrange the fluctuation amplitudes of all windows in ascending order, and select the 95th percentile as the fluctuation amplitude threshold. For example, if the 95th percentile of all fluctuation amplitudes is 1.2%, then set 1.2% as the high-frequency replacement point determination threshold. When the fluctuation amplitude of a window exceeds this threshold, the time point corresponding to the center of the window is marked as a high-frequency replacement point, and all such time points constitute the high-frequency replacement point set.

[0030] Next, statistical analysis methods were used to process the functional performance data. First, the overall mean and standard deviation of the functional indicator sequence were calculated. For example, if the mean of a protein expression sequence across 10 time points is 220 and the standard deviation is 15, then the upper threshold for functional abnormality is 220 plus 2 times 15, which equals 250, and the lower threshold for functional abnormality is 220 minus 2 times 15, which equals 190. Time points with functional values ​​exceeding 250 or falling below 190 are labeled as functional abnormalities, and all such points constitute the functional abnormality point set. Next, the set of high-frequency substitution points and the set of functional abnormalities are fused, and the strength of the association between them is calculated using the Gini index. Specifically, a sample set is constructed based on time points, with each record labeled as a high-frequency substitution point and a functional abnormality point. The sample is then partitioned based on these two Boolean features, and the Gini index is calculated for each subset. A smaller Gini index value indicates a stronger co-occurrence of the two features, indicating a higher correlation between nucleotide substitutions and functional abnormalities. The Gini index value at each time point constructed by the above operation is the association strength at that point, and all time points form an association strength dataset.

[0031] After completing the calculation of the association strength, the random forest algorithm was used to classify the dataset, and the model was trained to determine which substitution patterns were significantly associated with functional abnormalities. The model inputs were features such as the substitution frequency, functional value, and whether it was abnormal at each time point. The labels were assigned by experts in the existing data to indicate whether there was a risk. A forest model was constructed using 100 decision trees, and the classification results for each time point were finally obtained, output as a classification result dataset. Finally, a trend analysis was performed on the classification result dataset, and the sliding window method was used to count the frequency of time points classified as high-risk in continuous time periods to identify changes in risk trends. For example, if 4 out of 5 consecutive time points were high-risk, it was determined to be a rising risk interval, and this time period was included in the dynamic trend dataset. Through the above steps, the temporal correlation and risk dynamic trend between nucleotide substitution patterns and functional abnormalities were fully clarified, providing a solid data foundation and model support for subsequent risk prediction and management.

[0032] S3 includes: obtaining the original data set of functional abnormality records and ecological impact factors from the correlation data of nucleotide substitution patterns and functional abnormalities of gene editing products, including nucleotide substitution patterns, functional abnormality markers and ecological impact factor values, to obtain an initial interaction data set; using a standardization processing method, normalizing the ecological impact factor values ​​in the initial interaction data set to eliminate dimensional differences and obtain a standardized interaction data set; using a principal component analysis method, performing dimensionality reduction processing on the functional abnormality records and ecological impact factor values ​​in the standardized interaction data set, extracting the main interaction features, and obtaining a principal component feature set; if the feature contribution rate in the principal component feature set exceeds a preset threshold, it is marked as a high-correlation feature to obtain a high-correlation feature set; based on the high-correlation feature set, using a cluster analysis method, grouping the interaction patterns between functional abnormalities and ecological impact factors to obtain an interaction pattern set; using a time series analysis method, performing dynamic change detection on the interaction pattern set, analyzing the long-term effect trend of functional abnormalities on ecological impact factors, and obtaining a dynamic effect trend set; using a visualization processing method, graphically displaying the dynamic effect trend set to determine the interaction law between functional abnormalities and ecological impact factors.

[0033] The implementation of S3 is divided into the following seven steps: The first step is to extract nucleotide substitution patterns, functional abnormality markers, and ecological impact factor values ​​from the gene editing product dataset and simultaneously integrate these three types of information in chronological order to form an initial interactive dataset. Nucleotide substitution patterns are represented by the type of base change at each site, functional abnormality markers are represented by the presence or absence of abnormal functional states, and ecological impact factor values ​​are quantified environmental parameters such as non-target insect populations and soil microbial indices. Each time point constitutes a data record.

[0034] The second step is to standardize the ecological impact factor values. Specifically, the maximum and minimum values ​​of each ecological factor at all time points are counted, with the normalization interval set between 0 and 1. For each ecological factor value at each sample point, the current value minus the minimum value of the factor is subtracted, and then divided by the difference between the maximum and minimum values. This ensures that the values ​​of different factors are compared on a unified scale, avoiding dimensional inconsistencies that could affect subsequent analysis.

[0035] The third step is to perform principal component analysis on the standardized data. This analysis first constructs the covariance matrix of the functional abnormality markers and the ecological factor values, then calculates the eigenvalues ​​of the covariance matrix and their corresponding eigenvectors, sorts them from large to small by eigenvalue, and calculates the proportion of each eigenvalue to the sum of all eigenvalues. If the proportion of a certain feature exceeds the set threshold, it is considered to be the principal component. In this embodiment, the threshold is set to 0.8, that is, the principal component must explain more than 80% of the data variance to ensure that the selected features are representative. All principal components that meet this condition are retained to form a highly correlated feature set.

[0036] The fourth step is cluster analysis based on the highly correlated feature set. This analysis uses the K-means clustering algorithm. First, the initial number of clusters is set to 5 based on the data volume, and 5 initial center points are selected. Then, all sample points are traversed, and the Euclidean distance between each center point and each sample point is calculated. The samples are assigned to the cluster with the closest center point. The center point of each cluster is then updated to the mean of all samples within it. The assignment and update process is repeated until all center points converge. Finally, 5 different interaction patterns are obtained, which constitute the interaction pattern set.

[0037] The fifth step is to conduct time series analysis on the interaction pattern set. This analysis method uses quarterly units as the unit and counts the frequency of each pattern within the corresponding time period. Specifically, the number of times a pattern appears in a quarter is divided by the total number of samples in that quarter to form a frequency series. This frequency series is then subjected to differential analysis. If a pattern appears at a frequency of 0.1, 0.2, and 0.3 for three consecutive quarters, and the frequency trend is increasing, the pattern is considered to have a positive dynamic effect trend. Similarly, a decreasing trend is considered a negative trend. All trend records constitute a dynamic effect trend set.

[0038] The sixth step is to graphically represent the dynamic interaction trend set. A line graph is used to represent the temporal trajectory of each interaction pattern, with the horizontal axis representing the time series and the vertical axis representing the frequency of the pattern in each quarter. This line graph can intuitively demonstrate the long-term impact of dysfunction on ecological factors.

[0039] The seventh step is to analyze the graphical results and record the interaction patterns between functional abnormalities and ecological factors. Based on the temporal evolution of different interaction patterns, we can identify which functional abnormalities are significantly correlated with which ecological factors and their impact direction, providing direct data support for subsequent ecological risk assessment and decision-making.

[0040] S4 includes: obtaining the original data set containing nucleotide substitution sites, insect resistance markers and ecological impact factor values ​​from the nucleotide substitution patterns and ecological impact factor data of genetically modified crops, normalizing the ecological impact factor values ​​using the Z-score normalization method to obtain a standardized data set; based on the standardized data set, using the support vector machine algorithm, through the radial basis function kernel and regularization parameters, optimizing the classification boundaries, calculating the probability of unexpected ecological impact caused by enhanced insect resistance, and obtaining a classification probability set; if a probability value in the classification probability set exceeds a preset threshold, it is marked as a high-risk impact, and a high-risk impact set is obtained; if the probability value is lower than the threshold, it is marked as a low-risk impact, and a low-risk impact set is obtained; according to the high-risk The K-means clustering algorithm was used to group nucleotide substitution patterns and ecological impact factors, and potential interaction patterns were extracted to obtain the interaction pattern set. The time series analysis method was used to detect dynamic changes in nucleotide substitution patterns with high risk impacts in the interaction pattern set, and their long-term correlation trends with ecological impact factors were analyzed to obtain the dynamic correlation trend set. The visualization processing method was used to graphically display the dynamic correlation trend set, and the long-term interaction rules between enhanced insect resistance and ecological impact factors were determined to obtain the interaction rule set. Nucleotide substitution patterns with high correlation were extracted from the interaction rule set, and the decision tree algorithm was used to construct a prediction model for enhanced insect resistance and unexpected ecological impacts to obtain the prediction model set.

[0041] First, a raw data set containing nucleotide substitution sites, insect resistance markers, and ecological impact factor values ​​was collected from transgenic crops. Nucleotide substitution sites refer to locations where base substitutions occur within the target gene sequence, and these data were determined through gene sequencing alignment. The insect resistance marker is a binary variable, with a value of 1 indicating enhanced insect resistance and a value of 0 indicating no enhancement. Ecological impact factor values ​​include multiple environmental indicators, such as non-target insect populations, soil nitrogen and phosphorus content, and heavy metal concentrations in water. Each factor is regularly recorded using field testing instruments or environmental monitoring platforms, uniformly numbered by sample, and constructed into a structured data table. To ensure data comparability, all ecological impact factor values ​​were normalized using a Z-score. This involves calculating the mean and standard deviation of each ecological factor separately. The mean of each factor is then subtracted from the factor value and divided by its standard deviation. The resulting data set is standardized, with zero mean and unit standard deviation, effectively eliminating differences in the original dimensions of each factor.

[0042] Subsequently, a support vector machine algorithm was used to construct a classification model to determine whether nucleotide substitutions lead to ecological risks. The radial basis function was selected as the kernel function because it is suitable for dealing with nonlinear problems. The regularization parameter is used to adjust the model's tolerance to training errors to avoid overfitting. The value of the regularization parameter was selected from 0.01, 0.1, 1, and 10 using the grid search method. The model classification accuracy was tested in a 5-fold cross-validation, and the parameter combination with the highest classification accuracy was selected as the final model setting. In the process of training the model, the nucleotide substitution sites and ecological factor values ​​of the standardized data set were used as input features, and the insect resistance label was used as the classification target. After the model training was completed, all samples were predicted, and the classification probability of each sample being "enhanced insect resistance causing unexpected ecological impacts" was output to obtain a classification probability set.

[0043] The classification probability threshold was set at 0.7. When the classification probability value of a sample is greater than 0.7, it is considered to have a high probability of causing ecological risk and is marked as a high-risk impact. If it is less than or equal to 0.7, it is marked as a low-risk impact. This threshold was set through expert evaluation and historical case comparison to ensure a balance between sensitivity and specificity in risk identification. These were then aggregated to form a high-risk impact set and a low-risk impact set.

[0044] Next, nucleotide substitution sites and their corresponding ecological factor values ​​were extracted from the high-risk impact set as analysis samples. K-means clustering was used to identify potential interaction patterns. A K value of 5 was used, and the optimal number of clusters was determined based on the elbow rule when the silhouette coefficient was maximized. Initially, five points were randomly selected from the sample as initial cluster centers. Samples were then classified based on the Euclidean distance between the sample and the cluster center. The mean of all samples in each cluster was calculated and used as the new cluster center. This process was repeated until all sample categories stabilized and no longer changed, ultimately resulting in a set of interaction patterns.

[0045] Nucleotide substitution sites with high-risk effects in the interaction pattern set were analyzed for dynamic change detection. The temporal distribution sequence of each substitution site was extracted, and its frequency of occurrence over different time periods was analyzed using a moving average. A sliding window size of 5 was set, representing a window of five time units. Frequency trends were identified by comparing the difference in frequency between windows. If a substitution site showed a significant increase in frequency (an increase exceeding 30% of the initial frequency) over three consecutive time windows, it was considered to have a dynamic change trend and was included in the dynamic association trend set.

[0046] In the dynamic correlation trend set, the trends of substitution sites and their associated ecological factors are graphically displayed, using visualization tools to create trend charts. The horizontal axis represents time, while the vertical axis shows the trajectories of nucleotide substitution frequency and ecological factor values, respectively. This allows for the observation of temporal correlations between substitution patterns and ecological risk, ultimately forming a set of interaction patterns.

[0047] Finally, substitution patterns highly correlated with ecological factors were screened from the interaction patterns, specifically nucleotide substitution sites with correlation coefficients exceeding 0.8. Using these as input features and the ecological impact level as the target variable, a decision tree algorithm was used to train a prediction model. During the training process, the optimal partitioning nodes were selected based on the information gain criterion, and the sample set was continuously split to form a tree structure. Ultimately, a set of prediction models was obtained that could be used to predict whether enhanced insect resistance would lead to unintended ecological impacts.

[0048] S5 includes: extracting ecological impact factors related to enhanced insect resistance from the interaction feature set, using the principal component analysis method to reduce the dimensionality of the ecological impact factors to obtain a key factor set; if the feature contribution rate of the key factor set exceeds the preset threshold, the cluster analysis method is used to group the key factor set to obtain an ecological factor interaction set; based on the ecological factor interaction set, the statistical characteristics of the mutation frequency and abnormal fluctuations are calculated to generate a dynamic risk feature set; through the dynamic risk feature set, the random forest algorithm is used to classify the dynamic changes of the ecological factor interaction to obtain the risk classification results; based on the risk classification results, the dynamic risk feature set is decomposed into time series to extract the abnormal fluctuation trend; based on the abnormal fluctuation trend, a dynamic risk feature portrait data set is constructed to generate a feature portrait set; through the feature portrait set, a visualization processing method is used to generate a dynamic risk distribution map to determine the dynamic interaction law between enhanced insect resistance and unexpected ecological impacts.

[0049] First, ecological factors directly associated with enhanced insect resistance were extracted from the obtained interactive feature set. Selection criteria included indicators with significant statistical correlations with non-target insect populations, ecosystem stability, water quality changes, and biodiversity. The raw values ​​of these ecological factors were organized chronologically into multidimensional time series data, and principal component analysis (PCA) was applied for dimensionality reduction. PCA first calculated the covariance matrix between all ecological factor variables. The eigenvalues ​​and corresponding eigenvectors of this matrix were then calculated and sorted from largest to smallest eigenvalue. The variance explained by each principal component in the original data was then calculated, and the cumulative contributions were accumulated. If the cumulative contribution of a principal component exceeded 80%, it was considered significantly representative and included in the key factor set. The 80% threshold was determined based on a preset explanatory power requirement to ensure that the extracted principal components retained a significant portion of the original data variation.

[0050] Next, cluster analysis was performed based on the key factor set obtained above, using the K-means algorithm for grouping. During the clustering process, the number of clusters was initially set to 3. This value was determined by the elbow rule, which evaluates the location of the "elbow" point in the error sum-squared trend graph for different numbers of clusters. During the initialization phase, three samples were randomly selected as initial cluster centers. The Euclidean distances of all samples to these three centers were then calculated, and each sample was assigned to the nearest center. The new cluster center for each group of samples was then recalculated. This process was repeated until the centers remained unchanged. Ultimately, several sets of ecological factor interactions representing different ecological behavior patterns were obtained.

[0051] Next, for each set of ecological factor interactions, we analyze its time series trends and count the frequency of mutations and the number of abnormal fluctuations for each factor within a specified time window. Mutation frequency is defined as the number of times a factor value changes by one standard deviation from its historical mean per unit time. Abnormal fluctuations are counted using a sliding window method: using a window width of 5 time units and moving forward one unit at a time, we compare the mean change rate of each factor within the current window with that within the previous window. If the change rate exceeds a set threshold of 20%, it is recorded as an abnormal fluctuation point. All mutation frequencies and the number of abnormal fluctuation points are counted to generate a dynamic risk signature set.

[0052] Subsequently, a random forest algorithm was used to classify the dynamic risk feature set. Specifically, a random forest model consisting of 100 decision trees was constructed. Each tree was trained using a different sample and feature subset. The training dataset was obtained through repeated sampling with replacement. Input features included the mutation frequency and number of abnormal fluctuations of ecological impact factors, and the output was a high-risk or low-risk classification label. After model training was complete, all samples were classified, and the distribution of high-risk samples was statistically analyzed.

[0053] For samples classified as high-risk, we perform time series decomposition in chronological order, using a sliding window method to extract their temporal trend components. Each window is set to 5 time units wide, with a step size of 1 unit. The change in the proportion of high-risk samples within each window is calculated. If the rate of change in risk ratio between two adjacent windows exceeds 20%, it is marked as an abnormal fluctuation trend point. Ultimately, these abnormal trends are combined into a time series, forming an abnormal fluctuation trend set.

[0054] Next, based on the aforementioned abnormal fluctuation trend set, we aggregate the corresponding mutation frequency, number of abnormal fluctuations, and risk classification results at each time point to construct a time series of risk signature records. These records are then combined to form a dynamic risk signature dataset. Each record contains a timestamp, ecological factor value, mutation frequency, abnormal fluctuation mark, and risk label.

[0055] Finally, the dataset was input into the visualization module, and a multidimensional time change graph was generated graphically, where the horizontal axis was time, the vertical axis was the ecological impact factor value and risk level, and different colors were used to identify the risk level. The final dynamic risk distribution map was generated to clarify the long-term interaction between enhanced insect resistance and unexpected ecological impacts, and present it graphically to provide a basis for risk prediction and intervention.

[0056] S6 includes: extracting the mutation frequency and abnormal fluctuation pattern of genetically modified crops from the dynamic risk feature portrait data set, using the K-means cluster analysis method, and calculating the similarity matrix of the interaction between nucleotide substitution pattern and ecological factors based on Euclidean distance to obtain preliminary risk classification results; based on the preliminary risk classification results, extracting the mutation frequency and abnormal fluctuation pattern corresponding to the high-risk category, using the time series decomposition method to calculate the periodic change trend of the mutation frequency, and obtaining a periodic fluctuation feature set; based on the periodic fluctuation feature set, calculating the interaction strength between ecological factors and mutation frequency, using the Pearson correlation coefficient analysis method to generate an interaction strength matrix, and obtaining a set of key influencing factors for the interaction of ecological factors; if relevant If the interaction intensity of the key influencing factor set exceeds the preset threshold, a multidimensional scaling analysis is performed on the mutation frequency and abnormal fluctuation pattern to generate a two-dimensional mapping of the risk distribution and obtain a risk distribution feature set; based on the risk distribution feature set, the support vector machine algorithm is used to classify the risk distribution characteristics, generate the boundary division of high-risk and low-risk categories, and obtain the classification boundary parameters; based on the classification boundary parameters, the nucleotide substitution patterns corresponding to the risk categories are extracted, and the interaction similarity matrix between the nucleotide substitution patterns and the ecological factors is calculated to generate a dynamic interaction feature set; based on the dynamic interaction feature set, a visualization processing method is used to generate a dynamic change diagram of the risk distribution of insect-resistant crops and determine the dynamic evolution law of the risk distribution.

[0057] First, mutation frequency data and abnormal fluctuation patterns corresponding to each sample were extracted from the dynamic risk profile dataset. Mutation frequency refers to the number of nucleotide substitutions detected within a specific time window (e.g., monthly). This data is obtained by comparing the changes in the genetic sequences of genetically modified crops at different time points. Abnormal fluctuation patterns are analyzed using a sliding window method to analyze the fluctuations in mutation frequency. If the change in mutation frequency exceeds 30% of the historical mean within three consecutive time points, an abnormal fluctuation pattern is considered established. After extracting these data, all samples were clustered using the K-means clustering method. The number of clusters, K, was set to three using the elbow method. This method calculates the total within-cluster squared error as the number of clusters changes from one to ten. The optimal cluster number corresponds to the inflection point where the error decreases gradually. During the clustering process, Euclidean distance was used as a measure of similarity between samples. The nucleotide substitution pattern and ecological factor value of each sample were constructed into a feature vector, and the distance between these vectors was used to determine the cluster category to which it belonged. Finally, a preliminary risk classification result was output, labeling each sample as high-risk, medium-risk, or low-risk.

[0058] After obtaining preliminary classification results, we extract all samples belonging to the high-risk category and re-obtain their corresponding mutation frequency series and abnormal fluctuation pattern data. We then use a time series decomposition method to conduct a periodic analysis. This method breaks down the time series into three components: a trend term, a seasonal term, and a random term. The trend term uses a sliding average method to calculate the mean of the two periods before and after each time point as the trend value. The seasonal term is calculated by calculating the average difference between the values ​​at the same position within the period. Ultimately, we identify cyclical trends and record the periodic peaks and troughs of all mutation frequencies, forming a set of cyclical fluctuation features.

[0059] Based on the set of periodic fluctuation characteristics, the Pearson correlation coefficient between each ecological factor and the mutation frequency series was calculated. This coefficient was calculated by first calculating the mean and standard deviation of the mutation frequency and ecological factor, then finding the covariance between the two, and finally dividing the covariance by the product of the standard deviations to obtain the correlation coefficient. If the Pearson correlation coefficient of an ecological factor is greater than 0.7, it is considered to have a significant correlation with the mutation frequency and is included in the set of key influencing factors. The threshold of 0.7 is a common standard for measuring strong correlations in statistical analysis.

[0060] Next, if the interaction strength value of any factor in the key influencing factor set exceeds 0.7, a multidimensional scaling analysis method is initiated. Specifically, the Euclidean distance matrix between all pairs of samples in the high-dimensional feature space is constructed. These samples are then mapped into a two-dimensional space using a stress function minimization method while maintaining the original distance relationship between samples as much as possible. This ultimately forms a two-dimensional mapping diagram, i.e., the risk distribution feature set.

[0061] The risk distribution feature set is then classified using a support vector machine algorithm. This algorithm employs a linear kernel function and uses the maximum margin principle to find the optimal classification boundary. The regularization parameter is set to unity to prevent overfitting. During model training, labeled samples are fed into the model, and weight coefficients and intercepts are determined through iterative optimization to determine the boundary between high-risk and low-risk categories. The classification boundary parameters are then output and used to determine the risk category of subsequent samples.

[0062] Subsequently, based on the obtained classification boundary parameters, the nucleotide substitution patterns corresponding to each category were screened and their interaction similarities with key ecological factors were calculated. The similarity calculation continued using Euclidean distance to obtain an interaction similarity matrix between the nucleotide substitution patterns and ecological factors, forming a dynamic interaction feature set.

[0063] Finally, using visualization methods, each data set in the dynamic interactive feature set is mapped into a graphical dynamic risk distribution map. Colors represent risk levels, and curves depict the evolution of risk over time. Combining the timeline and risk distribution allows for a visual display of the dynamic evolution of risk distribution for insect-resistant crops, enabling proactive identification of potential unintended ecological impacts and providing a basis for intervention.

[0064] S7 includes: obtaining nucleotide substitution patterns and ecological factor interaction characteristics from risk categories, using feature extraction methods to generate a feature data set containing mutation frequency and abnormal fluctuations; based on the feature data set, using a decision tree algorithm, calculating the splitting path of each feature through the Gini index, and generating an initial decision tree model; if the splitting path complexity of the initial decision tree model exceeds the preset threshold, then using a post-pruning strategy to remove splitting nodes with low contribution to generate an optimized decision tree model; based on the optimized decision tree model, extracting the splitting path characteristics corresponding to the mutation frequency and abnormal fluctuations, and determining the preliminary priority of the risk control measures; obtaining the nucleotide substitution patterns corresponding to the priority measures with a priority higher than the preset threshold from the preliminary priority, using the Pearson correlation coefficient analysis method to calculate the correlation matrix of the interaction with the ecological factors, and generating an interaction feature set; based on the interaction feature set, using a visualization processing method to generate a dynamic distribution map of the interaction between mutation frequency and ecological factors, and determining the dynamic priority of the risk control measures; if the priority measures in the dynamic priority do not match the preset threshold, then performing a secondary analysis on the interaction feature set, using the principal component analysis method to extract key risk characteristics, and determine the final priority of the risk control measures.

[0065] In this embodiment, the interaction characteristics of nucleotide substitution patterns and ecological factors are first extracted from the data set that was determined to be a high-risk category in the previous step. This process includes screening out records marked as high-risk from the dynamic risk identification results, and extracting the mutation frequency data, abnormal fluctuation labels, and ecological factor values ​​collected at the same time. The feature extraction method refers to calculating the correlation and change rate between each dimension based on data standardization, retaining features that have a significant impact on risk identification, and forming a structured feature data set containing mutation frequency and abnormal fluctuations. Each record in this data set clearly corresponds to the observation value of a biotechnology product at a time point.

[0066] Afterwards, the model was constructed using a decision tree algorithm. During the modeling process, the feature dataset was first input, the dataset was partitioned according to each feature, the Gini index was calculated, and the feature with the smallest Gini index was selected as the splitting feature for the current node. The Gini index is calculated by counting the purity of each category of samples under that split. The formula is the sum of the squared probabilities of all categories and the complement. After all features are calculated, the feature with the largest information gain is selected for partitioning, and the process is recursively executed until the termination condition is met. In terms of parameter settings, the maximum tree depth is set to 10 layers to avoid model overfitting; the minimum number of split samples is set to 5 to ensure that each split sample is representative; and the information gain threshold is set to 0.01. Node splitting is only allowed when the information gain is above this threshold. These parameters are tuned using cross-validation on the training and validation sets.

[0067] After the decision tree model is initially established, the model's split path complexity is calculated. Complexity is defined as the product of the total number of leaf nodes and the average split depth. If this complexity index exceeds the preset threshold of 80, it indicates that the model is overly complex, affecting its generalization ability. At this time, post-pruning is performed. The pruning steps are to evaluate the contribution of parent node splits to the overall model accuracy from leaf nodes upward. If the model's accuracy on the validation set drops by less than 2 percentage points after removing the node, the node is considered a low-contribution node and is deleted. The final result is an optimized decision tree model with a streamlined structure but an accuracy above 95% of the original model.

[0068] Next, based on the optimized model, the combined characteristics of mutation frequency and abnormal fluctuations corresponding to each high-risk predicted path are extracted and mapped to the preliminary priority of risk control measures. If the mutation frequency in a split path continuously exceeds 1.5 times the historical average and abnormal fluctuations are frequent (marked as abnormal more than three times in a continuous time window), the path is assigned a "high" priority. Priorities are divided into three levels: "high," "medium," and "low," each level representing the urgency of the recommended intervention.

[0069] Subsequently, associated nucleotide substitution patterns were extracted from all high-priority pathways and their correlations with ecological factors were assessed using Pearson correlation coefficient analysis. For each pair of variables, the mean and standard deviation were calculated, and the standardized values ​​were then calculated and covariance was divided by the product of the standard deviations to obtain the correlation coefficient. If the absolute value of the Pearson correlation coefficient between a substitution pattern and an ecological factor was greater than 0.7, it was considered to be significantly interactive, and the pattern and factor combination was included in the interaction feature set.

[0070] After establishing the interactive feature set, a visualization method was used to generate a dynamic distribution diagram of the interaction between mutation frequency and ecological factors. The horizontal axis of the diagram represents the time series, while the vertical axis represents the standardized mutation frequency and the standardized ecological factor value, respectively. A two-color line chart depicts their changing trajectories, and areas of abnormal fluctuation are marked with different background colors. This is used to visually demonstrate the synchronization or lag between the two, further revealing the correlation trend.

[0071] Finally, if some priority measures in the dynamic distribution diagram are found to be inconsistent with the preset intervention level, that is, if they are assigned a "high" priority but their interaction characteristics do not show significant fluctuations in the diagram, a secondary analysis of the interaction feature set is performed. This step uses principal component analysis (PCA). First, the covariance matrix of all interaction features is calculated, and the principal components with a cumulative variance contribution exceeding 85% are extracted to form a key risk feature set. The priority of the risk measures is then adjusted based on the principal component loadings to obtain the final priority of the risk control measures. This ensures that each measure has a clear calculation basis and visual support, improving the scientific nature and operability of the risk response.

[0072] After obtaining the interactive feature set, a visualization method was used to generate a dynamic distribution diagram of the interaction between mutation frequency and ecological factors. This method involves mapping each pair of nucleotide substitution patterns and ecological factors in the interactive feature set to a two-dimensional coordinate system, with mutation frequency as the vertical axis and ecological factor values ​​as the horizontal axis, forming a time series curve. Each time point represents an observation, and all time points are plotted as a line graph to form a dynamic distribution diagram. The color depth indicates the strength of the mutation frequency, which is quantified using a color scale. The color range is from light blue to dark red, representing five frequency intervals from low to high. Specifically, 0 to 0.2 is light blue, 0.2 to 0.4 is blue, 0.4 to 0.6 is dark blue, 0.6 to 0.8 is orange, and 0.8 to 1 is red. This diagram is used to identify trends in the impact of changes in ecological factors on mutation patterns over time, serving as a basis for subsequent dynamic priority adjustments. If some priority measures in the dynamic prioritization do not match the preset thresholds—that is, the actual priority is inconsistent with the response level set by the policy—and the sudden fluctuation frequency of high-priority measures does not meet the warning standard (for example, a volatility below 0.3), this indicates that the current features contain redundant or misleading information. Therefore, a secondary analysis of the interaction feature set is performed. This secondary analysis utilizes principal component analysis (PCA). This involves calculating the covariance matrix of all variables in the interaction feature set and performing eigenvalue decomposition on the matrix to obtain a set of principal components and their corresponding explained variance proportions. The top principal components with a cumulative explained variance proportion exceeding 85% are retained as key features, and dimensions with lower explanatory power are removed to construct a new risk feature space. In this embodiment, if the explained variance proportion of a principal component is less than 5%, it is not retained. Retaining 3 to 5 principal components typically meets the accuracy requirements of subsequent analysis. Finally, based on the key risk feature set extracted by PCA, the final priority of the risk control measures is reassessed and determined. This evaluation process calculates a weighted score for each observation under the key principal component based on the loading value of each principal component, that is, the weight of the original variable in the principal component. If the weighted score is higher than the set threshold of 0.7, the observation point is classified as a high-priority risk event, and the corresponding risk control measures are marked as a first-level response measure. Scores between 0.4 and 0.7 are medium priority, and scores below 0.4 are low priority responses. All priority classification results are manually reviewed and confirmed to form a final priority list of risk control measures, which serves as the basis for risk prevention and control strategies.

[0073] The parameters are described as follows: the mutation frequency is obtained by dividing the number of replacement events per unit time window by the total number of observed sites; the abnormal fluctuation is obtained by counting the number of times exceeding the set standard deviation threshold, and the standard deviation threshold is set to the historical mean plus 1.5 times the standard deviation; the splitting path complexity is calculated by multiplying the number of leaf nodes by the average path depth; the Gini index is obtained by summing the squares of the probabilities of various samples and taking the complement, which is used to evaluate the classification purity; the Pearson correlation coefficient calculates the linear correlation between variables, ranging from negative 1 to positive 1, and an absolute value greater than 0.7 indicates a strong correlation; the retention criterion for principal component analysis is that the cumulative explained variance ratio is not less than 85%, and variables with principal component loading values ​​exceeding 0.6 are considered key influencing factors; the priority threshold is obtained by regression analysis of historical emergency response samples to ensure the accuracy and feasibility of risk grading.

[0074] S8 includes: obtaining mutation frequency characteristics and abnormal fluctuation characteristics from priority risk control measures that are higher than the preset threshold, and combining them with the life cycle stage to generate an initial feature data set; using the logistic regression algorithm, analyzing the correlation between the initial feature data set and the life cycle stage to obtain a preliminary dynamic adjustment plan framework; if the complexity of the plan framework exceeds the preset threshold, the principal component analysis method is used to extract the key components in the initial feature data set to generate a feature data set; based on the feature data set, the correlation matrix between nucleotide substitution and ecological factor monitoring is calculated to obtain an interactive feature set; the cluster analysis method is used to group the interactive feature set to generate dynamic distribution characteristics of the interaction between mutation frequency and ecological factors to obtain a dynamic distribution pattern; if the priority measures in the dynamic distribution pattern do not match the preset threshold, the interactive feature set is processed secondary, and the weighted average method is used to calculate the weight of each feature to obtain the final dynamic adjustment plan framework; based on the final dynamic adjustment plan framework, a dynamic adjustment risk control plan including nucleotide substitution constraints and ecological factor monitoring is generated to determine the priority ranking of the final plan.

[0075] First, the corresponding mutation frequency characteristics and abnormal fluctuation characteristics are extracted from the identified priority risk control measures that are higher than the preset threshold. The mutation frequency characteristic refers to the number of substitutions that occur in a unit length nucleotide sequence within a certain time window, and the unit is times per kilobase; the abnormal fluctuation characteristic refers to the relative change of the mutation frequency in multiple consecutive time periods exceeding 30% of the historical average. Then, combined with the life cycle stage of the current sample, which is divided into the research and development stage, the trial stage, the commercialization stage, etc., it is used to supplement the environmental background information and form an initial feature data set containing life cycle labels, mutation frequency values, and fluctuation marks. The data set uses each time point as a sample unit to form a structured data table.

[0076] Next, the logistic regression algorithm is used to model the initial feature data set. The independent variables of the model include mutation frequency values, abnormal fluctuation marks, and life cycle stages, and the dependent variable is whether a risk control response is triggered. In the calculation of the logistic regression model, the independent variables are first standardized, and the mean and standard deviation of each feature are calculated. The standardized value is obtained by subtracting the mean and dividing it by the standard deviation. The maximum likelihood estimation method is then used to estimate the coefficient of each feature, and the model outputs the probability of each record triggering a risk. When the probability exceeds 0.5, it is considered to be a high-risk state. The regularization parameter uses L2 regularization to prevent overfitting, and its value is selected between 0.01, 0.1, 1, and 10 through five-fold cross-validation to minimize the logarithmic loss on the validation set.

[0077] If the model's structural complexity exceeds a preset threshold, further optimization is required. Complexity is defined as the product of the number of variables and the number of nonzero coefficients. The threshold is set at 20. If this threshold is exceeded, principal component analysis (PCA) is used to extract key features. The method first calculates the covariance matrix of the initial feature data, then calculates its eigenvalues ​​and eigenvectors. These are sorted in descending order by eigenvalue, and principal components with a cumulative contribution exceeding 80% are selected for dimensionality reduction. These principal components form a new feature dataset.

[0078] Subsequently, a correlation matrix between nucleotide substitutions and ecological factor monitoring values ​​was calculated based on the reduced feature dataset. This correlation was calculated using the Pearson correlation coefficient. The steps involved calculating the mean, standard deviation, and covariance of each feature time series. Finally, the correlation coefficient was calculated by dividing the covariance by the product of the two feature standard deviations. If the absolute value of the correlation coefficient between a feature and an ecological factor was greater than 0.7, it was considered a strong correlation and included in the interactive feature set.

[0079] Cluster analysis was performed on the interaction feature set using the K-means clustering method. The K value was chosen based on the elbow rule, determined when the decrease in the squared error within a cluster slowed significantly. Here, it was set to three clusters. Euclidean distance was used as the similarity metric for clustering, and features were assigned to the nearest cluster center based on the minimum distance. The average change trend of each interaction feature type across lifecycle stages was determined, forming a dynamic distribution pattern.

[0080] If the generated dynamic distribution pattern does not match the priorities in the risk control measures, that is, if there is no significant concentration of medium- and high-priority features in a certain category, the interactive feature set is subjected to another weighted average. The specific method is: for each feature, the average risk contribution is calculated based on its frequency and risk probability at each stage of the lifecycle, which is used as its weight. The weighted average of all features forms a new risk representation, forming the final dynamic adjustment plan framework.

[0081] Finally, a dynamically adjusted risk management plan is generated based on this framework. This plan clearly lists the mutation frequency thresholds and corresponding ecological factor trends that should be focused on at each lifecycle stage, as well as nucleotide substitution constraints and ecological factor monitoring elements. Each measure is ranked from high to low based on its average risk contribution, generating a final priority list. This completes the entire implementation process of this step.

[0082] When reprocessing the interactive feature set, a weighted average method is used to calculate the weight of each feature. The specific process is as follows: First, the frequency of occurrence of each interactive feature in each lifecycle stage is counted. The frequency is the ratio of the number of samples marked as high risk for that feature in a given stage to the total number of samples in that stage. Second, the average correlation coefficient of each feature across all stages is calculated, using the absolute value to enhance the ability to identify strongly negatively correlated features. Third, the frequency value is multiplied by the correlation coefficient to obtain the weight of each feature. To maintain the total weight of 1, all feature weights are normalized. That is, the weight of each feature is divided by the sum of all feature weights to form the final feature weight set. Based on this weighted feature set, the risk impact indicators are recombined to construct the final dynamic adjustment solution framework.

[0083] Subsequently, based on the final dynamic adjustment plan framework, mutation frequency data and ecological factor monitoring information were integrated to formulate a dynamic risk control plan. This plan includes two key components: first, nucleotide substitution constraints, which clarify which substitution patterns should be restricted or subject to increased supervision under certain ecological factors; second, ecological factor monitoring requirements, which specify which ecological factors should be monitored at specific mutation frequencies. Finally, all possible control measures were ranked according to the weight values ​​of their corresponding characteristics, with higher values ​​indicating higher priority. This prioritized ranking of the final plan, prioritizing high-weighted measures to control risk spread and mutation anomalies.

[0084] like Figure 2As shown, a biotechnology product risk dynamic identification system based on multi-indicator fusion is also provided, which is used to implement the steps of the biotechnology product risk dynamic identification method based on multi-indicator fusion, and the system includes: a data acquisition module, which obtains experimental data and environmental data from each stage of the genetically modified crop life cycle, constructs a dynamic data set including molecular structure sequences, functional performance indicators and ecological impact factors, and extracts the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators and dynamic change trends of ecological impact factors; a stage feature analysis module, which extracts the nucleotide substitution pattern and functional abnormality data set of gene editing products in the research and development stage and the application stage according to the mutation frequency of molecular structure sequences and abnormal fluctuations in functional performance indicators, and determines the association pattern between nucleotide substitution pattern and functional abnormality; an interactive feature extraction module, which extracts the potential effect characteristics of ecological impact factors after functional abnormality is released into the environment from the association pattern between nucleotide substitution pattern and functional abnormality, and obtains the interactive feature set between functional abnormality and ecological impact factors; an ecological risk identification module , extract the nucleotide substitution pattern and ecological impact factor data related to the insect resistance of transgenic crops from the interaction feature set, and judge the probability of unexpected ecological impact caused by enhanced insect resistance; dynamic portrait construction module, if the probability of unexpected ecological impact caused by enhanced insect resistance exceeds the preset threshold, extract the key factors of the ecological impact factors after environmental release from the interaction feature set, and construct a dynamic risk feature portrait data set including mutation frequency, abnormal fluctuation and ecological factor interaction; risk classification module, based on the dynamic risk feature portrait data set, calculates the similarity of nucleotide substitution pattern and ecological factor interaction, and obtains the risk classification result of insect-resistant transgenic crops; strategy priority determination module, extracts nucleotide substitution pattern and ecological factor interaction characteristics from the risk classification result, and determines the priority of risk control measures for mutation frequency and abnormal fluctuation; risk control generation module, based on the risk control measure priority, combines the mutation frequency and abnormal fluctuation data of the life cycle stage, and generates a dynamically adjusted risk control plan including nucleotide substitution constraints and ecological factor monitoring.

[0085] This system integrates multiple modules to dynamically identify risks throughout the life cycle of biotechnology products, particularly genetically modified crops. First, the data acquisition module continuously collects experimental and environmental data from the R&D, testing, promotion, and commercialization stages of genetically modified crops, constructing a dynamic dataset encompassing molecular structure sequences, functional performance indicators, and ecological impact factors. It also extracts information on nucleotide substitution mutation frequencies, abnormal fluctuations in functional performance, and dynamic trends in ecological factors from the time series. Subsequently, the stage feature analysis module divides the collected data into life cycle stages and, through statistical analysis, extracts nucleotide substitution patterns and functional abnormalities from each stage, establishing significant associations between nucleotide substitution patterns and functional abnormalities. Based on these results, the interactive feature extraction module further identifies the potential effects of functional abnormalities on ecological factors after release, extracting the resulting interactive feature set for subsequent analysis. Next, the ecological risk identification module extracts nucleotide substitution patterns and ecological factor data associated with enhanced insect resistance from the interactive feature set and assesses the probability of these patterns causing unintended ecological impacts. If this probability exceeds a preset threshold, the dynamic profile construction module further extracts key ecological factors after release from the interactive features, constructing a dynamic risk profile dataset encompassing the interaction between mutation frequencies, abnormal fluctuations, and ecological factors. Based on this profile dataset, the risk classification module calculates the similarity between nucleotide substitution patterns and interactions with ecological factors, and uses this information to classify the risks of insect-resistant genetically modified crops. Furthermore, the strategy prioritization module analyzes the interaction characteristics of nucleotide substitution patterns and ecological factors extracted from the classification results, and determines the priority of control measures for different risk categories based on mutation frequency and abnormal fluctuation. Finally, the risk control generation module combines the mutation frequency and abnormal fluctuation data for each risk type during the life cycle stage to generate a dynamically adjusted risk control plan that includes nucleotide substitution constraints and ecological factor monitoring requirements. This plan clarifies the intervention priorities and implementation paths for different risk levels, thereby achieving dynamic, scientific, and systematic risk identification and response for biotechnology products.

[0086] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A dynamic identification method for biotechnology product risks based on multi-index fusion, characterized by: include: S1. Acquire experimental and environmental data from each stage of the genetically modified crop life cycle to construct a dynamic dataset containing molecular structure sequences, functional performance indicators, and ecological impact factors. Extract the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators, and dynamic trends in ecological impact factors. S2. Based on the frequency of molecular structure sequence mutations and abnormal fluctuations in functional performance indicators, extract the nucleotide substitution patterns and functional abnormality datasets of gene editing products during the R&D and application stages, and determine the correlation pattern between nucleotide substitution patterns and functional abnormalities; S3. Extract the potential effect characteristics of the functional abnormality on the ecological impact factors after release from the nucleotide substitution pattern and the functional abnormality association pattern, and obtain the interaction feature set between the functional abnormality and the ecological impact factors; S4. Extract nucleotide substitution patterns and ecological impact factor data related to insect resistance of transgenic crops from the interactive feature set to determine the probability of unintended ecological impacts caused by enhanced insect resistance; S5. If the probability of unexpected ecological impact caused by enhanced insect resistance exceeds the preset threshold, extract the key factors of ecological impact factors after environmental release from the interactive feature set, and construct a dynamic risk feature portrait dataset that includes mutation frequency, abnormal fluctuations, and interactions between ecological factors; S6. Based on the dynamic risk profile dataset, calculate the similarity of nucleotide substitution patterns and ecological factor interactions to obtain risk classification results for insect-resistant transgenic crops; S7. Extract the interaction characteristics of nucleotide substitution patterns and ecological factors from the risk classification results to determine the priority of risk control measures for mutation frequency and abnormal fluctuations; S8. Based on the priority of risk control measures, combined with the mutation frequency and abnormal fluctuation data of the life cycle stage, a dynamically adjusted risk control plan including nucleotide substitution constraints and ecological factor monitoring is generated.

2. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1 is characterized by: Said S1 comprises: Acquire experimental and environmental data from each stage of the genetically modified crop life cycle, construct a dynamic dataset containing molecular structure sequences, functional performance indicators, and ecological impact factors, and store them in a time series format to obtain an initial dataset; The time series analysis method is used to process the molecular structure sequence in the initial data set, calculate the mutation frequency at each time point, and obtain the mutation frequency sequence; Through statistical analysis methods, the mutation frequency sequence is detected for anomalies. If the mutation frequency exceeds the preset threshold, it is marked as an abnormal point, and the abnormal mutation point set is obtained; Based on the time series data of functional performance indicators, a sliding window method is used to detect the fluctuation of the indicator value. If the fluctuation amplitude exceeds the preset threshold, it is marked as abnormal fluctuation, and the abnormal fluctuation point set is obtained; Obtain the time series data of ecological impact factors, use trend analysis methods to extract the changing trends of factor values, and obtain the ecological impact trend series; By using the correlation analysis method, the abnormal mutation point set, abnormal fluctuation point set and ecological impact trend sequence are integrated to build a dynamic relationship model among the three, and obtain a comprehensive dynamic relationship data set. The machine learning algorithm random forest was used to perform classification prediction on the comprehensive dynamic relationship data set, determine the potential connection between molecular structure mutations and abnormal functional performance and ecological impact trends, and obtain classification results.

3. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1, characterized in that: The S2 includes: Acquire nucleotide sequence data and functional performance data from gene-edited products of genetically modified crops, construct an initial dataset in time series format, including nucleotide substitution patterns and functional abnormality records, and obtain an initial time series dataset; The time series analysis method was used to process the nucleotide sequences in the initial time series data set, and the nucleotide substitution frequencies in the R&D stage and the application stage were calculated to obtain the substitution frequency series. The fluctuation characteristics of the replacement frequency sequence are analyzed by the sliding window method. If the fluctuation amplitude exceeds the preset threshold, it is marked as a high-frequency replacement point, and a high-frequency replacement point set is obtained; For the time series of functional performance data, statistical analysis methods are used to detect functional abnormalities. If the functional value deviates from the preset threshold, it is marked as an abnormal point, and a set of functional abnormality points is obtained; By using the Gini index analysis method, the high-frequency substitution point set and the functional abnormality point set were integrated to calculate the correlation strength between the nucleotide substitution frequency and functional abnormality, and obtain the correlation strength data set; The random forest algorithm was used to classify the association strength dataset to determine the potential relationship between nucleotide substitution patterns and functional abnormalities, and a classification result dataset was obtained. Through the trend analysis method, the trend of the classification result data set in the time dimension is extracted, the dynamic change relationship between the replacement pattern and functional abnormality is analyzed, and a dynamic trend data set is obtained.

4. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1, characterized in that: The S3 includes: From the association data between nucleotide substitution patterns and functional abnormalities of gene-edited products, we obtain the original dataset of functional abnormality records and ecological impact factors, including nucleotide substitution patterns, functional abnormality markers, and ecological impact factor values, to obtain the initial interaction dataset. The ecological impact factor values ​​in the initial interaction dataset were normalized using a standardization method to eliminate dimensional differences and obtain a standardized interaction dataset. The principal component analysis method was used to reduce the dimensionality of the functional abnormality records and ecological impact factor values ​​in the standardized interaction dataset, extract the main interaction features, and obtain the principal component feature set. If the feature contribution rate in the principal component feature set exceeds the preset threshold, it is marked as a high-correlation feature and a high-correlation feature set is obtained; Based on the highly correlated feature set, cluster analysis was used to group the interaction patterns between functional abnormalities and ecological impact factors, and the interaction pattern set was obtained. Through time series analysis, the dynamic change detection of the interaction pattern set is carried out, the long-term trend of the functional abnormality on the ecological impact factors is analyzed, and the dynamic effect trend set is obtained; A visualization processing method is used to graphically display the dynamic action trend set and determine the interaction rules between functional abnormalities and ecological influencing factors.

5. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1 is characterized by: The S4 includes: From the nucleotide substitution patterns and ecological impact factor data of transgenic crops, an original data set containing nucleotide substitution sites, insect resistance markers, and ecological impact factor values ​​was obtained. The ecological impact factor values ​​were normalized using the Z-score normalization method to obtain a standardized data set. Based on the standardized data set, the support vector machine algorithm was used to optimize the classification boundaries through the radial basis function kernel and regularization parameters, and the probability of unexpected ecological impact caused by enhanced insect resistance was calculated to obtain the classification probability set. If a probability value in the classification probability set exceeds a preset threshold, it is marked as a high-risk impact, and a high-risk impact set is obtained. If the probability value is lower than the threshold, it is marked as a low-risk impact, and a low-risk impact set is obtained. According to the high-risk impact set, the K-means clustering algorithm was used to group the nucleotide substitution patterns and ecological impact factors, extract potential interaction patterns, and obtain the interaction pattern set; By using time series analysis methods, dynamic changes in the nucleotide substitution patterns with high risk impacts in the interaction pattern set were detected, and their long-term correlation trends with ecological impact factors were analyzed to obtain a dynamic correlation trend set. Using visualization methods, the dynamic correlation trend set is graphically displayed to determine the long-term interaction between insect resistance enhancement and ecological influencing factors, and the interaction rule set is obtained; Highly correlated nucleotide substitution patterns were extracted from the interaction regularity set, and a decision tree algorithm was used to construct a prediction model for enhanced insect resistance and unexpected ecological impacts to obtain a prediction model set.

6. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1, characterized in that: The S5 includes: The ecological impact factors related to insect resistance enhancement were extracted from the interactive feature set, and the principal component analysis method was used to reduce the dimensionality of the ecological impact factors to obtain the key factor set. If the characteristic contribution rate of the key factor set exceeds the preset threshold, the key factor set is grouped using the cluster analysis method to obtain the ecological factor interaction set; Based on the interaction set of ecological factors, the statistical characteristics of mutation frequency and abnormal fluctuation are calculated to generate a dynamic risk feature set; Through the dynamic risk feature set, the random forest algorithm is used to classify the dynamic changes of ecological factor interactions and obtain the risk classification results; Based on the risk classification results, the dynamic risk feature set is decomposed into time series to extract abnormal fluctuation trends; According to the abnormal fluctuation trend, a dynamic risk feature portrait data set is constructed to generate a feature portrait set; Through the feature portrait set and the use of visualization processing methods, a dynamic risk distribution map is generated to determine the dynamic interaction between enhanced insect resistance and unexpected ecological impacts.

7. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1 is characterized by: The S6 includes: The mutation frequencies and abnormal fluctuation patterns of genetically modified crops were extracted from the dynamic risk profile dataset. K-means cluster analysis was used to calculate the similarity matrix of the interaction between nucleotide substitution patterns and ecological factors based on Euclidean distance to obtain preliminary risk classification results. Based on the preliminary risk classification results, the mutation frequency and abnormal fluctuation pattern corresponding to the high-risk category are extracted. The periodic change trend of the mutation frequency is calculated using the time series decomposition method to obtain the periodic fluctuation feature set. Based on the periodic fluctuation feature set, the interaction strength between ecological factors and mutation frequency was calculated. The Pearson correlation coefficient analysis method was used to generate the interaction strength matrix and obtain the key influencing factor set of ecological factor interaction. If the interaction intensity of the key influencing factor set exceeds the preset threshold, a multidimensional scaling analysis is performed on the mutation frequency and abnormal fluctuation pattern to generate a two-dimensional mapping of the risk distribution and obtain a risk distribution feature set; Based on the risk distribution feature set, the support vector machine algorithm is used to classify the risk distribution features, generate the boundary division of high-risk and low-risk categories, and obtain the classification boundary parameters; According to the classification boundary parameters, the nucleotide substitution patterns corresponding to the risk categories are extracted, the interaction similarity matrix between the nucleotide substitution patterns and ecological factors is calculated, and a dynamic interaction feature set is generated; Based on the dynamic interactive feature set, a visualization processing method is used to generate a dynamic change map of the risk distribution of insect-resistant crops and determine the dynamic evolution law of the risk distribution.

8. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1 is characterized by: The S7 includes; The interaction characteristics of nucleotide substitution patterns and ecological factors are obtained from risk categories, and feature extraction methods are used to generate a feature dataset containing mutation frequencies and abnormal fluctuations. Based on the feature data set, the decision tree algorithm is used to calculate the splitting path of each feature through the Gini index to generate the initial decision tree model; If the splitting path complexity of the initial decision tree model exceeds the preset threshold, a post-pruning strategy is adopted to remove the splitting nodes with low contribution and generate an optimized decision tree model; Based on the optimized decision tree model, the splitting path characteristics corresponding to the mutation frequency and abnormal fluctuations are extracted to determine the initial priority of risk control measures; The nucleotide substitution patterns corresponding to the priority measures with priorities higher than the preset threshold were obtained from the preliminary priorities, and the correlation matrix of the interactions with the ecological factors was calculated using the Pearson correlation coefficient analysis method to generate the interaction feature set; Based on the interactive feature set, a visualization method is used to generate a dynamic distribution map of the interaction between mutation frequency and ecological factors, and to determine the dynamic priority of risk control measures; If the priority measures in the dynamic priority do not match the preset threshold, a secondary analysis of the interactive feature set is performed using the principal component analysis method to extract key risk features and determine the final priority of the risk control measures.

9. The method for dynamic identification of biotechnology product risks based on multi-index fusion according to claim 1, characterized in that: The S8 includes: Obtain mutation frequency characteristics and abnormal fluctuation characteristics from priority risk control measures above the preset threshold, and generate an initial feature dataset based on the life cycle stage; By using a logistic regression algorithm, we analyze the correlation between the initial feature dataset and the life cycle stages, and obtain a preliminary dynamic adjustment scheme framework. If the complexity of the solution framework exceeds the preset threshold, the principal component analysis method is used to extract the key components in the initial feature data set to generate a feature data set; Based on the feature data set, the correlation matrix between nucleotide substitution and ecological factor monitoring was calculated to obtain the interaction feature set; Cluster analysis method is used to group the interaction feature sets, generate dynamic distribution features of the interaction between mutation frequency and ecological factors, and obtain dynamic distribution patterns; If the priority measures in the dynamic distribution pattern do not match the preset threshold, the interactive feature set is processed twice, and the weighted average method is used to calculate the weight of each feature to obtain the final dynamic adjustment plan framework; Based on the final dynamic adjustment plan framework, a dynamic adjustment risk management plan including nucleotide substitution constraints and ecological factor monitoring is generated to determine the priority ranking of the final plan.

10. A system for dynamically identifying risks of biotechnology products based on multi-indicator fusion, used to implement the steps of the method for dynamically identifying risks of biotechnology products based on multi-indicator fusion according to any one of claims 1 to 9, characterized in that: The system comprises: The data acquisition module acquires experimental and environmental data from each stage of the genetically modified crop life cycle, constructs a dynamic data set containing molecular structure sequences, functional performance indicators, and ecological impact factors, and extracts the mutation frequency of molecular structure sequences over time, abnormal fluctuations in functional performance indicators, and dynamic trends in ecological impact factors; The stage feature analysis module extracts nucleotide substitution patterns and functional abnormality datasets of gene editing products during the R&D and application stages based on the frequency of molecular structure sequence mutations and abnormal fluctuations in functional performance indicators, and determines the correlation pattern between nucleotide substitution patterns and functional abnormalities; The interactive feature extraction module extracts the potential effects of functional abnormalities on ecological impact factors after release from the nucleotide substitution pattern and the functional abnormality association pattern, and obtains the interactive feature set between functional abnormalities and ecological impact factors; The ecological risk identification module extracts nucleotide substitution patterns and ecological impact factor data related to insect resistance of transgenic crops from the interactive feature set to determine the probability of unexpected ecological impacts caused by enhanced insect resistance; A dynamic profile construction module: if the probability of unexpected ecological impact caused by enhanced insect resistance exceeds a preset threshold, key factors of ecological impact factors after environmental release are extracted from the interactive feature set to construct a dynamic risk feature profile dataset that includes mutation frequency, abnormal fluctuations, and interactions between ecological factors; The risk classification module calculates the similarity between nucleotide substitution patterns and ecological factor interactions based on the dynamic risk profile dataset to obtain risk classification results for insect-resistant transgenic crops; The strategy priority determination module extracts nucleotide substitution patterns and ecological factor interaction characteristics from risk classification results to determine the priority of risk control measures for mutation frequency and abnormal fluctuations; The risk control generation module generates a dynamically adjusted risk control plan that includes nucleotide substitution constraints and ecological factor monitoring based on the priority of risk control measures and combined with the mutation frequency and abnormal fluctuation data of the life cycle stage.

Citation Information

Patent Citations

  • Nucleic acid constructs

    CN101155597A

  • Nuclear radiation index early warning system based on multi-algorithm fusion

    CN120234772A