A data-driven pathogenic microorganism target analysis method and system

By constructing a pathogen microbial target analysis system and integrating the entire life cycle characteristic data of pathogens, the one-sided nature of target analysis and environmental response lag problems in the existing technology are solved, and multi-dimensional dynamic recognition and accurate modeling of target recognition are achieved, which improves the comprehensiveness and accuracy of target screening.

CN120256901BActive Publication Date: 2025-08-26CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510748076.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-26
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, pathogenic microbial target analysis lacks systematic research on the correlation of characteristics and the impact of environmental factors on different growth stages, resulting in target screening that cannot reflect the integrity of the pathogen's life cycle, and the induction effect of environmental factors lacks dynamic modeling, making it difficult to achieve accurate target prediction and risk warning.

Method used

By building a data-driven pathogenic microbial target analysis system, integrating the entire life cycle characteristic data of the pathogen, establishing a time-dimensional correlation model, realizing early and late feature synergistic quantification and dynamic recognition of environmental induction factors, including target sample filing, driving feature generation, feature matrix analysis and potential target classification processing.

Benefits of technology

It realizes multi-dimensional data-driven target dynamic recognition, improves the comprehensiveness and accuracy of target recognition, reduces the cost of manual analysis, avoids subjective judgment errors, and provides accurate target screening and environmentally responsive target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256901B_ABST
    Figure CN120256901B_ABST
Patent Text Reader

Abstract

The present invention discloses a data-driven pathogen target analysis method and system, which belongs to the field of data analysis technology. The method includes: establishing a pathogen target experimental sample library, separating pathogen type, environmental induction factor, and early and late feature appearance time data clusters; configuring parameters and generating driving feature states; constructing an early and late target feature matrix, analyzing feature correlation; iteratively evaluating to generate a potential target matrix and marking environmental induction factors. The system includes target sample archiving, driving feature generation, feature matrix analysis, and potential target classification processing modules. The present invention quantifies the temporal correlation of pathogen early and late features through data-driven modeling, realizes dynamic target identification and environmental factor correlation analysis, and is conducive to improving target screening efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a data-driven pathogenic microorganism target analysis method and system. Background Art

[0002] Pathogen target analysis is a core component of antimicrobial drug development and pathogen detection. Existing technologies primarily rely on static analysis of single features, lacking systematic research into the correlations between pathogen characteristics at different growth stages (e.g., early metabolism, late pathogenicity) and the impact of environmental factors. For example, the dynamic correlations between pathogen metabolites expressed early and late in the pathogen's virulence factors have not been effectively quantified, resulting in target screening that fails to reflect the entire pathogen life cycle. Furthermore, the induction effects of environmental factors (e.g., temperature, nutrient concentration) on target characteristics lack dynamic modeling, making accurate target prediction and risk warning difficult. Summary of the Invention

[0003] In response to the one-sidedness of target analysis and the lag in environmental response in the existing technology, the present invention proposes a data-driven pathogen target analysis method and system. By integrating the characteristic data of the entire life cycle of pathogens, a time dimension correlation model is constructed to achieve the quantification of the synergistic characteristics of early and late stages and the dynamic identification of environmental induction factors, providing a new paradigm for the precise screening of pathogen targets.

[0004] The present invention provides the following technical solutions:

[0005] A data-driven pathogenic microorganism target analysis system includes: a target sample archiving module, a driving feature generation module, a feature matrix analysis module, and a potential target classification processing module;

[0006] The target sample archiving module is used to establish a pathogenic microorganism target experimental sample library, store target experimental sample data and generate data clusters;

[0007] The driving feature generation module is used to configure data cluster parameters, establish time dimension relationships, and capture and generate driving feature states;

[0008] The feature matrix analysis module is used to lock samples and construct a target feature matrix to analyze the correlation between early and late features;

[0009] The potential target classification processing module is used to iteratively evaluate feature relevance, generate a potential target matrix and mark environmental induction factors.

[0010] Furthermore, the target sample archiving module includes a sample storage unit and a data cluster generation unit;

[0011] The sample storage unit is used to store uniquely coded pathogenic microorganism target experimental samples and record pathogen types, environmental inducing factors, and the time of appearance of early and late stage characteristics;

[0012] The data cluster generating unit separates data clusters based on indicator category information, including pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters, and late feature appearance time data clusters.

[0013] Furthermore, the driving feature generation module includes a parameter configuration unit and a driving capture unit;

[0014] The parameter configuration unit is used to configure the data cluster parameters and establish a time delay scale correspondence between the early feature appearance time and the late feature appearance time;

[0015] The driving capture unit is used to generate a driving characteristic state by taking the pathogen type and the environmental induction factor as data driving objects.

[0016] Furthermore, the feature matrix analysis module includes a matrix construction unit and a correlation analysis unit;

[0017] The matrix construction unit is used to construct early and late target feature matrices using the driving feature states as matrix elements;

[0018] The correlation analysis unit quantifies the early and late feature synergy and calculates feature correlation based on Boolean matrix intersection and union operations.

[0019] Furthermore, the potential target classification processing module includes an iterative evaluation unit and a potential target marking unit;

[0020] The iterative evaluation unit iteratively adjusts the time delay scale by presetting the feature correlation threshold to screen the target feature matrix with high correlation;

[0021] The potential target marking unit is used to accumulate the screened matrix to generate a potential target matrix, extract row sequences and mark potential environmental induction factors corresponding to pathogen types.

[0022] A data-driven pathogenic microorganism target analysis method comprises the following steps:

[0023] Step S1: Establish a pathogenic microorganism target experiment sample library, store target experiment sample data, record pathogen type, environmental induction factor, early feature appearance time and late feature appearance time, and separate corresponding data clusters;

[0024] Step S2: Parameter configuration is performed on each data cluster, and a time dimension relationship between the appearance time of early features and the appearance time of late features is established; pathogen type and environmental induction factor are used as data driving objects to capture and generate driving feature states;

[0025] Step S3: Based on the time dimension relationship, the pathogenic microorganism target experimental samples at the early feature appearance time or the late feature appearance time are locked respectively to form the target feature matrix at the early feature appearance time and the late feature appearance time, and the feature correlation analysis between the pathogenic microorganism target experimental samples at the early and late stages is performed;

[0026] Step S4: Iteratively evaluate the feature correlation to form a potential target matrix; based on the different row sequences in the potential target matrix corresponding to the pathogen type, mark the potential environmental induction factors and output them to the experimenter port.

[0027] Furthermore, the specific implementation process of step S1 includes:

[0028] Establishing a pathogenic microorganism target experimental sample library, wherein the pathogenic microorganism target experimental sample library records different pathogenic microorganism target experimental samples, and each pathogenic microorganism target experimental sample is accompanied by different indicator category information, wherein the indicator category information includes pathogen type, environmental induction factor, early characteristic appearance time, and late characteristic appearance time;

[0029] Based on the indicator category information, the pathogenic microorganism target experimental samples are separated to obtain data clusters with different indicator category information attributes, including pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters and late feature appearance time data clusters.

[0030] Furthermore, the specific implementation process of step S2 includes:

[0031] Configuring pathogen type data cluster , Environmental Induced Factor Data Cluster , early feature appearance time data cluster and late feature appearance time data clusters ,in, represents the hth pathogenic microorganism target experimental sample, H represents the total number of pathogenic microorganism target experimental samples, 、 and Respectively represent the experimental samples of pathogenic microorganism targets Pathogen type obtained during isolation , environmental induction factors and early feature appearance time , n, i and r are the coding numbers of pathogen type, environmental induction factor and early characteristic appearance time respectively, N, I and R are the maximum coding numbers of pathogen type, environmental induction factor and early characteristic appearance time respectively, and 、 and , g is the time delay scale, the early feature appearance time and late feature appearance time Having a time dimension correspondence with the time delay scale g;

[0032] Capturing data driven objects: pathogen type and environmental induction factors , generate the driving characteristic state , in pathogenic microorganism target experimental samples If the pathogen type is captured at the same time and environmental induction factors , then let the driving characteristic state If the pathogen type is not captured at the same time and environmental induction factors , then let the driving characteristic state .

[0033] Furthermore, the specific implementation process of step S3 includes:

[0034] The driving feature state is used as the matrix element with row number n and column number i to constitute the early feature appearance time and late feature appearance time The target feature matrix under and ;

[0035] Evaluate pathogenic microorganism target experimental samples based on target feature matrix The correlation between the early and late features , where Represents the target feature matrix and The number of matrix elements with a value of 1 after the Boolean intersection operation between them, Represents the target feature matrix and The number of matrix elements with a value of 1 after the Boolean union operation between them;

[0036] In the above method, the physical principle of feature correlation is: feature correlation is equal to "the number of feature-environment combinations that appear simultaneously in the early and late stages" divided by "the number of feature-environment combinations that appear at least once in the early and late stages". The closer the value of feature correlation is to 1, the stronger the synergy of early and late features is, and the greater the target association risk is.

[0037] Furthermore, the specific implementation process of step S4 includes:

[0038] Preset feature-related thresholds and set Perform iterative evaluation of feature relevance, and When the iterative evaluation stops, G is the time node number when the target experiment ends. When the iterative evaluation stops, the feature correlation greater than or equal to the feature correlation threshold is screened out. Corresponding target feature matrix ;

[0039] The selected target feature matrix Perform matrix accumulation and summation operations, and record the matrix after accumulation and summation operations as the potential target matrix , extract the potential target matrix The matrix elements of the nth row in , to form a row sequence , and mark the pathogen type Potential environmental induction factors ;

[0040] In the above method, low-correlation data are eliminated through dynamic iteration, high-risk target combinations are focused on, and dynamic correlation mining between environmental factors and pathogen characteristics is achieved.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] Dynamic target identification driven by multi-dimensional data: By establishing a pathogen target experimental sample library, the system systematically integrates multi-dimensional data such as pathogen type, environmental inducing factors, and early and late characteristic times, breaking through the limitations of traditional single-feature analysis. For example, in bacterial target analysis, the time series data of "Escherichia coli-high temperature-early metabolites-late endotoxins" can be simultaneously correlated to quantify the synergy between characteristics of different growth stages and environmental factors, thereby improving the comprehensiveness of target identification.

[0043] Accurate modeling of temporal associations: By constructing early and late feature matrices and calculating temporal correlation coefficients, the system can quantify the strength of associations between pathogenic features at different time points. For example, in virus research, by analyzing the temporal correlation coefficients between early antigen expression and late pathogenic proteins, it is possible to precisely locate strongly associated target combinations such as "dengue virus-37°C-NS1 protein-E protein," providing a quantitative basis for viral life cycle research.

[0044] Dynamic labeling of environmental inducible factors: Through iterative evaluation and potential target matrix analysis, the system can automatically label key environmental inducible factors for specific pathogen types. For example, in fungal experiments, by filtering highly correlated matrices through cumulative summation operations, the environmental-feature combination of "Candida albicans - acidic pH - hyphal morphology - invasion enzyme" can be identified, providing precise environmentally responsive targets for antifungal drug development.

[0045] Improved experimental efficiency and accuracy: The system significantly reduces manual analysis costs through automated processes such as data cluster separation, matrix modeling, and iterative evaluation. It also avoids subjective judgment errors through quantitative evaluation of time correlation coefficients, improving the accuracy and repeatability of target screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0047] Figure 1 It is a schematic diagram of the steps of a data-driven pathogenic microorganism target analysis method of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] In the first embodiment of the present invention, a data-driven pathogenic microorganism target analysis system is provided, which includes: a target sample archiving module, a driving feature generation module, a feature matrix analysis module, and a potential target classification processing module;

[0050] The target sample archiving module is used to establish a pathogenic microorganism target experimental sample library, store target experimental sample data and generate data clusters;

[0051] Among them, the target sample archiving module includes a sample storage unit and a data cluster generation unit;

[0052] Sample storage unit, used to store uniquely coded pathogenic microorganism target experimental samples, record pathogen type, environmental induction factors and the time of appearance of early and late characteristics;

[0053] A data cluster generation unit separates data clusters based on indicator category information, including pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters, and late feature appearance time data clusters;

[0054] The driving feature generation module is used to configure data cluster parameters, establish time dimension relationships, and capture and generate driving feature states;

[0055] Among them, the driving feature generation module includes a parameter configuration unit and a driving capture unit;

[0056] A parameter configuration unit, used to configure data cluster parameters and establish a time delay scale correspondence between the early feature appearance time and the late feature appearance time;

[0057] A driving capture unit is used to generate a driving characteristic state by taking pathogen types and environmental induction factors as data driving objects;

[0058] Feature matrix analysis module, used to lock samples and construct target feature matrix, and analyze the correlation between early and late features;

[0059] Among them, the feature matrix analysis module includes a matrix construction unit and a correlation analysis unit;

[0060] A matrix construction unit, used to construct early and late target feature matrices using driver feature states as matrix elements;

[0061] The correlation analysis unit quantifies the synergy of early and late features and calculates feature correlation based on Boolean matrix intersection and union operations;

[0062] Potential target classification processing module, used to iteratively evaluate feature relevance, generate a potential target matrix and mark environmental induction factors;

[0063] Among them, the potential target classification processing module includes an iterative evaluation unit and a potential target marking unit;

[0064] The iterative evaluation unit, through presetting feature correlation thresholds, iteratively adjusts the time delay scale to screen target feature matrices with high correlation;

[0065] The potential target marking unit is used to accumulate the screened matrix to generate a potential target matrix, extract the row sequence and mark the potential environmental induction factors corresponding to the pathogen type.

[0066] See also Figure 1 In the second embodiment, a data-driven pathogenic microorganism target analysis method is provided to adapt to the above-mentioned first embodiment. The method includes the following steps:

[0067] Step S1: Establish a pathogenic microorganism target experiment sample library, store target experiment sample data, record pathogen type, environmental induction factor, early feature appearance time and late feature appearance time, and separate corresponding data clusters;

[0068] Exemplarily, a pathogenic microorganism target experimental sample library is established, in which different pathogenic microorganism target experimental samples are recorded, and each pathogenic microorganism target experimental sample is accompanied by different indicator category information, and the indicator category information includes pathogen type, environmental induction factor, early feature appearance time, and late feature appearance time;

[0069] Based on the indicator category information, the pathogenic microorganism target experimental samples were separated to obtain data clusters with different indicator category information attributes. The data clusters included pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters, and late feature appearance time data clusters.

[0070] Step S2: Parameter configuration is performed on each data cluster, and a time dimension relationship between the appearance time of early features and the appearance time of late features is established; pathogen type and environmental induction factor are used as data driving objects to capture and generate driving feature states;

[0071] For example, configure the pathogen type data cluster , Environmental Induced Factor Data Cluster , early feature appearance time data cluster and late feature appearance time data clusters ,in, represents the hth pathogenic microorganism target experimental sample, H represents the total number of pathogenic microorganism target experimental samples, 、 and Respectively represent the experimental samples of pathogenic microorganism targets Pathogen type obtained during isolation , environmental induction factors and early feature appearance time , n, i and r are the coding numbers of pathogen type, environmental induction factor and early characteristic appearance time respectively, N, I and R are the maximum coding numbers of pathogen type, environmental induction factor and early characteristic appearance time respectively, and 、 and , g is the time delay scale, the early feature appearance time and late feature appearance time Having a time dimension correspondence with the time delay scale g;

[0072] Capturing data driven objects: pathogen type and environmental induction factors , generate the driving characteristic state , in pathogenic microorganism target experimental samples If the pathogen type is captured at the same time and environmental induction factors , then let the driving characteristic state If the pathogen type is not captured at the same time and environmental induction factors , then let the driving characteristic state .

[0073] Step S3: Based on the time dimension relationship, the pathogenic microorganism target experimental samples at the early feature appearance time or the late feature appearance time are locked respectively to form the target feature matrix at the early feature appearance time and the late feature appearance time, and the feature correlation analysis between the pathogenic microorganism target experimental samples at the early and late stages is performed;

[0074] For example, the driving feature state is taken as a matrix element with row number n and column number i to respectively constitute the early feature appearance time and late feature appearance time The target feature matrix under and ;

[0075] Evaluate pathogenic microorganism target experimental samples based on target feature matrix The correlation between the early and late features , where Represents the target feature matrix and The number of matrix elements with a value of 1 after the Boolean intersection operation between them, Represents the target feature matrix and The number of matrix elements with a value of 1 after the Boolean union operation between them.

[0076] Step S4: Iteratively evaluate the feature correlation to form a potential target matrix; based on the different row sequences in the potential target matrix corresponding to the pathogen type, mark the potential environmental inducing factors and output them to the experimenter port;

[0077] For example, a feature-related threshold is preset, and Perform iterative evaluation of feature relevance, and When the iterative evaluation stops, G is the time node number when the target experiment ends. When the iterative evaluation stops, the feature correlation greater than or equal to the feature correlation threshold is screened out. Corresponding target feature matrix ;

[0078] The selected target feature matrix Perform matrix accumulation and summation operations, and record the matrix after accumulation and summation operations as the potential target matrix , extract the potential target matrix The matrix elements of the nth row in , to form a row sequence , and mark the pathogen type Potential environmental induction factors ;

[0079] For example, in the bacterial target analysis scenario, the experimental object is Pseudomonas aeruginosa, and the main environmental factor studied is iron deficiency concentration. , mainly in the form of siderophores, in the late , forming a biofilm, and generating a potential target matrix after iterative evaluation, marking "iron deficiency" as the key environmental induction factor. The iron carrier is strongly correlated with the biofilm in an iron-deficient environment, and can be used as a combined antibacterial target; in the virus target analysis scenario, the experimental object is influenza virus, and the main environmental factor studied is the host cell temperature. The hemagglutinin HA is observed in the early stage, and the neuraminidase NA is observed in the late stage. After iterative evaluation, generating a potential target matrix, marking "37℃" as the key environmental factor, HA and NA are synergistic targets, and the HA and NA expressions of influenza virus at 37℃ are strongly correlated, which can be used as a dual-target combination of antiviral drugs.

[0080] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0081] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A data-driven pathogenic microorganism target analysis method, characterized in that: The method comprises the following steps: Step S1: Establish a pathogenic microorganism target experiment sample library, store target experiment sample data, record pathogen type, environmental induction factor, early feature appearance time and late feature appearance time, and separate corresponding data clusters; Step S2: Parameter configuration is performed on each data cluster, and a time dimension relationship between the appearance time of early features and the appearance time of late features is established; pathogen type and environmental induction factor are used as data driving objects to capture and generate driving feature states; Step S3: Based on the time dimension relationship, the pathogenic microorganism target experimental samples at the early feature appearance time or the late feature appearance time are locked respectively to form the target feature matrix at the early feature appearance time and the late feature appearance time, and the feature correlation analysis between the pathogenic microorganism target experimental samples at the early and late stages is performed; Step S4: Iteratively evaluate the feature correlation to form a potential target matrix; based on the different row sequences in the potential target matrix corresponding to the pathogen type, mark the potential environmental inducing factors and output them to the experimenter port; The specific implementation process of step S4 includes: The feature correlation threshold is preset, and g=g+1 is set to perform iterative evaluation of feature correlation, and the iterative evaluation stops when r+g=G, where G is the time node number when the target experiment ends. When the iterative evaluation stops, the feature correlation C[(t r , t r+g )|S h ]The corresponding target feature matrix V N·I [t r+g (S h )],S h represents the hth pathogenic microorganism target experimental sample, t r represents the time when the rth early feature appears, g is the time delay scale, t r+g (S h ) represents the time of appearance of late characteristics, N and I represent the maximum coding number of pathogen type and environmental induction factor, respectively; For the selected target feature matrix V N·I [t r+g (S h )] perform matrix accumulation and summation operation, and record the matrix after accumulation and summation operation as the potential target matrix V(S h ), extract the potential target matrix V(S h ) to form the row sequence n|V(S h ) and mark the pathogen type ad n Potential environmental induction factors i is the coding number of the environmental induction factor.

2. A data-driven pathogenic microorganism target analysis method according to claim 1, characterized in that: The specific implementation process of step S1 includes: Establishing a pathogenic microorganism target experimental sample library, wherein the pathogenic microorganism target experimental sample library records different pathogenic microorganism target experimental samples, and each pathogenic microorganism target experimental sample is accompanied by different indicator category information, wherein the indicator category information includes pathogen type, environmental induction factor, early characteristic appearance time, and late characteristic appearance time; Based on the indicator category information, the pathogenic microorganism target experimental samples are separated to obtain data clusters with different indicator category information attributes, including pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters and late feature appearance time data clusters.

3. A data-driven pathogenic microorganism target analysis method according to claim 2, characterized in that: The specific implementation process of step S2 includes: Configure pathogen type data cluster AD = {ad n (S h )|n∈[1,N],h∈[1,H]}, environmental induction factor data cluster FD={fd i (S h )|i∈[1,I],h∈[1,H]}、early feature appearance time data cluster ZT={t r (S h )|r∈[1,R],h∈[1,H]} and late feature appearance time data cluster WT={t r+g (S h )|r∈[1, R], h∈[1, H]}, where H represents the total number of pathogenic microorganism target experimental samples, ad n (S h ), fd i (S h ) and t r (S h ) represent the pathogenic microorganism target experimental samples S h The pathogen type ad obtained during isolation n 、Environmental induction factor fd i and early feature appearance time t r , n represents the coding number of the pathogen type, R represents the maximum coding number of the early feature appearance time, and ad n (S h )=ad n 、fd i (S h )=fd i and t r (S h )=t r , early feature appearance time t r (S h ) and late feature appearance time t r+g (S h ) have a time dimension correspondence relationship with a time delay scale g; Capture data driven object: pathogen type ad n and environmental induction factor fd i , generate the driving characteristic state S h [ad n , fd i ], in pathogenic microorganism target experimental sample S h If pathogen type ad is captured at the same time n and environmental induction factor fd i , then let the driving characteristic state S h [ad n , fd i ]=1, if pathogen type ad is not captured at the same time n and environmental induction factor fd i , then let the driving characteristic state S h [ad n , fd i ]=0.

4. A data-driven pathogenic microorganism target analysis method according to claim 3, characterized in that: The specific implementation process of step S3 includes: The driving feature state is used as the matrix element with row number n and column number i to respectively constitute the early feature appearance time t r (S h ) and late feature appearance time t r+g (S h ) target feature matrix V under N·I [t r (S h )] and V N·I [t r+g (S h )]; Evaluate pathogenic microorganism target experimental samples S based on target feature matrix h The correlation between the early and late features In the formula, NUM{V N·I [t r (S h )]∩V N·I [t r+g (S h )]} represents the target feature matrix V N·I [t r (S h )] and V N·I [t r+g (S h )] after the Boolean intersection operation between the matrix elements with a value of 1, NUM{V N·I [t r (S h )]∪V N·I [t r+g (S h )]} represents the target feature matrix V N·I [t r (S h )] and V N·I [t r+g (S h )] contains the number of matrix elements with a value of 1 after the Boolean union operation between them.

5. A data-driven pathogenic microorganism target analysis system, which executes the pathogenic microorganism target analysis method according to any one of claims 1 to 4, characterized in that: The system includes: a target sample archiving module, a driving feature generation module, a feature matrix analysis module and a potential target classification processing module; The target sample archiving module is used to establish a pathogenic microorganism target experimental sample library, store target experimental sample data and generate data clusters; The driving feature generation module is used to configure data cluster parameters, establish time dimension relationships, and capture and generate driving feature states; The feature matrix analysis module is used to lock samples and construct a target feature matrix to analyze the correlation between early and late features; The potential target classification processing module is used to iteratively evaluate feature relevance, generate a potential target matrix and mark environmental induction factors.

6. The data-driven pathogenic microorganism target analysis system according to claim 5, characterized in that: The target sample archiving module includes a sample storage unit and a data cluster generation unit; The sample storage unit is used to store uniquely coded pathogenic microorganism target experimental samples and record pathogen types, environmental inducing factors, and the time of appearance of early and late stage characteristics; The data cluster generating unit separates data clusters based on indicator category information, including pathogen type data clusters, environmental induction factor data clusters, early feature appearance time data clusters, and late feature appearance time data clusters.

7. The data-driven pathogenic microorganism target analysis system according to claim 5, characterized in that: The driving feature generation module includes a parameter configuration unit and a driving capture unit; The parameter configuration unit is used to configure the data cluster parameters and establish a time delay scale correspondence between the early feature appearance time and the late feature appearance time; The driving capture unit is used to generate a driving characteristic state by taking the pathogen type and the environmental induction factor as data driving objects.

8. The data-driven pathogenic microorganism target analysis system according to claim 5, characterized in that: The feature matrix analysis module includes a matrix construction unit and a correlation analysis unit; The matrix construction unit is used to construct early and late target feature matrices using the driving feature states as matrix elements; The correlation analysis unit quantifies the early and late feature synergy and calculates feature correlation based on Boolean matrix intersection and union operations.

9. The data-driven pathogenic microorganism target analysis system according to claim 5, characterized in that: The potential target classification processing module includes an iterative evaluation unit and a potential target marking unit; The iterative evaluation unit iteratively adjusts the time delay scale by presetting the feature correlation threshold to screen the target feature matrix with high correlation; The potential target marking unit is used to accumulate the screened matrix to generate a potential target matrix, extract row sequences and mark potential environmental induction factors corresponding to pathogen types.

Citation Information

Patent Citations

  • Intelligent analysis and processing system and method for disease hidden danger based on epidemic disease big data

    CN119339973A