A knowledge graph-based microbial secondary metabolite retrieval method
By employing a knowledge graph-based method for retrieving microbial secondary metabolites, a correlation map between secondary metabolites and microorganisms is constructed. Combining feature transformation and similarity analysis, and introducing material temporal features, this approach addresses the scarcity of rare medicinal plant symbiotic bacteria resources. It enables efficient and accurate screening of microbial substitutes, reduces research costs, and alleviates the supply-demand imbalance of medicinal resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-03-27
AI Technical Summary
In the current technology, rare medicinal plant symbiotic bacteria resources are scarce. Traditional secondary metabolite research processes are complicated, inefficient, and costly, and lack targeted screening guidelines, making it difficult to quickly identify suitable microorganisms for target secondary metabolites. As a result, the contradiction between supply and demand of medicinal resources cannot be effectively alleviated.
Based on knowledge graphs, key features of secondary metabolites are linked to microorganisms. Through feature transformation and domain knowledge constraints, combined with a hierarchical retrieval mechanism of similarity analysis and attention weight, material time sequence features are introduced to verify metabolic rhythm adaptability, forming a multi-dimensional and precise retrieval system.
Rapidly identify suitable microorganisms for target secondary metabolites, shorten research cycles, reduce research costs, improve search efficiency, broaden the scope of strain screening, ensure the reliability of alternative bacteria in practical applications, and alleviate the supply and demand contradiction of medicinal resources.
Smart Images

Figure CN121478844B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of secondary metabolite retrieval technology, specifically a knowledge graph-based method for retrieving microbial secondary metabolites. Background Technology
[0002] Current research on secondary metabolites of medicinal actinomycetes still employs the conventional strategy of strain fermentation and crude product extraction. Subsequent chromatographic techniques such as thin-layer chromatography, column chromatography, and HPLC are needed for product separation and purification, followed by spectroscopic identification using MS, UV, IR, and NMR, and finally, biological methods for evaluating the activity of individual compounds. This research process is complex and lengthy, resulting in prolonged research time and low efficiency. Furthermore, the stringent requirements for experimental consumables and equipment in chromatographic separation and spectroscopic identification further increase research costs. Moreover, traditional strategies lack targeted strain screening guidance, relying solely on experimental methods to conduct research blindly, making it difficult to quickly identify suitable strains for the target secondary metabolites. This severely restricts the research efficiency and quality improvement of medicinal actinomycete secondary metabolites and fails to effectively alleviate the supply-demand imbalance of rare medicinal plant resources.
[0003] In summary, there is an urgent need for a new knowledge graph-based method for retrieving microbial secondary metabolites to achieve efficient and accurate retrieval of substitutes for rare medicinal plant symbiotic bacteria, quickly identify suitable microorganisms corresponding to target secondary metabolites, improve research efficiency, reduce research costs, and alleviate the supply and demand contradiction of medicinal resources. Summary of the Invention
[0004] The purpose of this application is to provide a knowledge graph-based method for retrieving microbial secondary metabolites, in order to solve the technical problems in the existing technology, such as the scarcity of rare medicinal plant symbiotic bacteria resources, the limitation of large-scale utilization, the complicated, inefficient, and costly traditional secondary metabolite research process, the lack of targeted screening guidance, the difficulty in quickly identifying suitable microbial substitutes for target products, and the inability to effectively alleviate the contradiction between supply and demand of medicinal resources.
[0005] To achieve the above objectives, this application provides a knowledge graph-based method for retrieving microbial secondary metabolites, applied to retrieving alternative terms for symbiotic bacteria of rare medicinal plants, including:
[0006] A knowledge graph is constructed based on the key features of secondary metabolites and the corresponding microorganisms to obtain a retrieval knowledge graph. The key features include at least the chemical formula and structural formula of the secondary metabolites. The retrieval knowledge graph takes the key features as input and outputs the microorganisms corresponding to the key features.
[0007] The key features of the target secondary metabolite are subjected to feature transformation processing, which is constrained by preset domain knowledge to obtain candidate transformation features, which are potential alternatives to the key features.
[0008] The search is performed on the knowledge graph based on the candidate transformation features. The search is based at least on similarity and attention weights, and the microorganisms corresponding to the candidate transformation features are output as alternatives to the microorganisms corresponding to the target secondary metabolites.
[0009] Preferably, the feature transformation process includes:
[0010] Obtain the domain knowledge constraints, which are the constraints on the chemical changes that occur between compounds corresponding to secondary metabolites.
[0011] Based on the aforementioned domain knowledge constraints, the key features of the target secondary metabolite are augmented with data. These key features correspond to the target chemical formula features and the target structural formula features. The data augmentation includes a first augmentation for the target chemical formula features and a second augmentation for the target structural formula features.
[0012] The results of the first expansion and the second expansion are the candidate transformation features obtained from the feature transformation process.
[0013] Preferably, the first expansion is as follows:
[0014] Based on the domain knowledge constraints, a first change condition corresponding to the target chemical formula feature is determined. The first change condition is used to characterize the constraint rules for obtaining the target chemical formula feature from other candidate chemical formula features.
[0015] Based on the first change condition, the candidate chemical formula features corresponding to the target chemical formula feature are determined, and the determined candidate chemical formula features are the result of the first expansion.
[0016] Preferably, the second extension is as follows:
[0017] Based on the domain knowledge constraints, a second change condition corresponding to the target structural feature is determined. The second change condition is used to characterize the constraint rules for obtaining the target structural feature from other candidate structural features.
[0018] Based on the second change condition, the candidate structural features corresponding to the target structural feature are determined, and the determined candidate structural features are the result of the second expansion.
[0019] Preferably, the first change condition is as follows:
[0020] Based on the aforementioned domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target chemical formula features and the corresponding candidate chemical formula features. This feature extraction includes at least the extraction of element conservation features, functional group derivation features, molecular skeleton splicing and cyclization features, small molecule removal features of condensation reactions, and structural unit homology features of secondary metabolism between precursors and products. The result of this feature extraction is defined as the first change condition.
[0021] Preferably, the second change condition is as follows:
[0022] Based on the aforementioned domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target structural features and the corresponding candidate structural features. This feature extraction includes at least the extraction of molecular skeleton splicing and rearrangement features, functional group transformation site features, cyclization and ring-opening configuration features, chemical bond breaking and bonding features, and side chain group spatial derivation features. The result of this feature extraction is defined as the second change condition.
[0023] Preferably, the retrieval based on the candidate transformation features in the retrieval knowledge graph includes:
[0024] The candidate transformation features are analyzed to obtain candidate chemical formula features and candidate structural formula features; wherein, the candidate chemical formula features are potential substitutes corresponding to the target chemical formula features, and the candidate structural formula features are potential substitutes corresponding to the target structural features;
[0025] Based on the candidate chemical formula features and / or candidate structural formula features, a similarity analysis is performed on the retrieval knowledge graph, and chemical formula features and / or structural formula features that meet the preset similarity threshold are output. The secondary metabolites and corresponding microorganisms corresponding to the output are used as the first retrieval results.
[0026] Based on the first search result, and using a preset attention weighting mechanism, the overall probability value of the first search result being converted into the target secondary metabolite is analyzed. The first search result that meets the preset overall probability value threshold is output as the second search result. The alternative of the microorganism corresponding to the target secondary metabolite corresponding to the second search result is output.
[0027] Preferably, the attention weighting mechanism includes:
[0028] Initialize the attention weight allocation dimension, which corresponds to the synthesis matching dimension of the candidate chemical formula feature and the synthesis matching dimension of the candidate structural formula feature, respectively;
[0029] Calculate the synthesis and transformation matching degree between the candidate chemical formula feature and the target chemical formula feature, and the synthesis and transformation matching degree between the candidate structural formula feature and the target structural formula feature corresponding to the first search result, and use these as the initial attention weight values for each dimension.
[0030] The initial attention weights for each dimension are normalized to obtain the normalized attention weights.
[0031] Based on the normalized attention weight, the overall probability value of synthesizing the target secondary metabolite corresponding to the compound in the first search result is calculated in a weighted manner. The overall probability value is compared with the overall probability value threshold, and the first search result that meets the threshold requirement is output as the second search result.
[0032] Preferably, the alternative to the microorganism corresponding to the target secondary metabolite corresponding to the second search result further includes:
[0033] A material temporal feature is constructed, which is used to characterize the time-series generation pattern of primary and secondary metabolites during the production of secondary metabolites by microorganisms.
[0034] Extract the material time sequence characteristics of the target secondary metabolites and the material time sequence characteristics of the microorganisms corresponding to the second search results;
[0035] The material time sequence characteristics of the target secondary metabolite are compared with the material time sequence characteristics of the microorganisms corresponding to the second search result. Microorganisms that meet the preset material time sequence characteristic similarity threshold are selected and output as alternatives to the microorganisms corresponding to the target secondary metabolite.
[0036] Preferably, the construction of the material temporal characteristics includes:
[0037] Divide the growth sequence stages in the process of microbial production of secondary metabolites;
[0038] Samples were collected at each time stage to detect and obtain the types and abundance data of primary and secondary metabolites at the corresponding growth stages.
[0039] Using growth time as a dimension, species and abundance data are correlated and integrated in chronological order to construct a time-series dataset;
[0040] The time-series dataset is standardized to obtain the time-series characteristics of the substance.
[0041] Beneficial Effects: The knowledge graph-based microbial secondary metabolite retrieval method of this application constructs a retrieval knowledge graph that associates key features of secondary metabolites with microorganisms. Combined with a bidirectional feature expansion strategy constrained by knowledge in the field of microbial secondary metabolism, and a hierarchical retrieval mechanism integrating similarity analysis and attention weights, and incorporating material temporal features for metabolic rhythm adaptability verification, this method forms a multi-dimensional and precise retrieval system. This method effectively overcomes the limitations of traditional techniques, such as lack of targeted screening guidance, single feature dimensions, and large retrieval errors. It eliminates the reliance on symbiotic bacteria of rare medicinal plants, enabling rapid identification of suitable microbial substitutes for target secondary metabolites. This significantly shortens the research cycle, reduces the proportion of invalid experiments, and lowers equipment and material costs. Simultaneously, it considers both chemical feature matching and metabolic mechanism compatibility, improving the practical reliability of substitute bacteria, broadening the range of strain screening, and reducing the omission of effective strains. This effectively alleviates the supply and demand contradiction of medicinal resources and provides sufficient strain resources to support the large-scale production of pharmaceutical active compounds. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A flowchart illustrating the knowledge graph-based microbial secondary metabolite retrieval method provided in this application embodiment;
[0044] Figure 2 This is a flowchart illustrating the feature transformation process provided in an embodiment of this application.
[0045] Figure 3 A flowchart illustrating the process of searching a knowledge graph based on candidate transformation features, as provided in this application embodiment.
[0046] Figure 4 A flowchart illustrating the attention weighting mechanism provided in an embodiment of this application;
[0047] Figure 5 An optimized flowchart of the second search result output provided in the embodiments of this application;
[0048] Figure 6 A flowchart illustrating the construction of material timing features provided in this application embodiment.
[0049] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0050] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0051] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0052] Actinomycetes are a typical type of microorganism among plant symbiotic fungi. Addressing the scarcity of rare medicinal plant symbiotic fungi and the difficulty of quickly identifying alternative microorganisms or actinomycetes using traditional research methods, this embodiment provides a knowledge graph-based method for retrieving microbial secondary metabolites. This method is used to retrieve symbiotic alternatives or other microbial or compound options corresponding to target secondary metabolites. It achieves efficient and accurate retrieval through the construction of a retrieval knowledge graph, feature transformation, and targeted searching.
[0053] Reference Figure 1 , Figure 1 A flowchart illustrating the knowledge graph-based microbial secondary metabolite retrieval method provided in this application embodiment.
[0054] like Figure 1 As shown, this embodiment discloses a knowledge graph-based method for retrieving microbial secondary metabolites, applied to retrieving alternative terms for symbiotic bacteria of rare medicinal plants, including:
[0055] S10: Construct a knowledge graph based on the key features of secondary metabolites and the corresponding microorganisms to obtain a retrieval knowledge graph. The key features include at least the chemical formula and structural features of the secondary metabolites. The retrieval knowledge graph takes the key features as input and outputs the microorganisms corresponding to the key features.
[0056] In this specific application, the data sources are first collected: sample data containing secondary metabolites (such as active compounds produced by symbiotic bacteria of rare medicinal plants) and corresponding microorganisms (symbiotic bacteria such as actinomycetes and fungi) are extracted from public biological databases and related literature. Then, key features are extracted: for each secondary metabolite, chemical formula features (such as elemental composition and atomic ratio) and structural features (such as molecular skeleton, functional group type and site) are extracted and associated with the strain number and classification information of the corresponding microorganism. Finally, a retrieval knowledge graph is constructed: using the Neo4j graph tool, two types of core nodes are defined: "secondary metabolite nodes" and "microorganism nodes," with "production" as the link between nodes; chemical formula and structural feature attributes are added to product nodes, and species information attributes are added to microorganism nodes to complete the graph construction. When using it, inputting chemical formula or structural feature will match and output the corresponding microorganism node information.
[0057] S20: Perform feature transformation processing on the key features of the target secondary metabolite. The feature transformation processing is constrained by the preset domain knowledge to obtain candidate transformation features, which are potential substitutes for the key features.
[0058] The feature transformation process of this embodiment will now be described in detail. This step is the core link to effectively expand the key features of the target and obtain potential alternative features. In this method, the feature transformation process uses the chemical change constraint of secondary metabolites as the domain knowledge constraint, and performs feature expansion on the target chemical formula feature and the target structural formula feature respectively, and finally obtains the candidate transformed features that meet the retrieval requirements.
[0059] Reference Figure 2 , Figure 2 This is a flowchart illustrating the feature transformation process provided in an embodiment of this application.
[0060] like Figure 2 As shown, the specific feature transformation process includes:
[0061] S21: Obtain domain knowledge constraints, which are the constraints on the chemical changes that occur between compounds corresponding to secondary metabolites.
[0062] In the specific application of this embodiment, domain knowledge constraints are obtained from well-known rules in the field of microbial secondary metabolism, metabolic pathway databases, and published literature on secondary metabolic synthesis. Specifically, these constraints can be: constraints on chemical changes such as condensation, rearrangement, and functional group transformation between compounds, covering the constraints of element conservation, molecular skeleton evolution, small molecule removal / introduction, and chemical bond formation / breaking, which are integrated to form a standardized set of chemical change constraint rules as the basis for feature expansion.
[0063] S22: Based on domain knowledge constraints, data expansion is performed on the key features of the target secondary metabolite. These key features correspond to the target chemical formula features and the target structural formula features. This data expansion includes a first expansion for the target chemical formula features and a second expansion for the target structural formula features.
[0064] Specifically, the first expansion is as follows:
[0065] Based on domain knowledge constraints, the first change condition corresponding to the target chemical formula feature is determined. The first change condition is used to characterize the constraint rules for obtaining the target chemical formula feature from other candidate chemical formula features.
[0066] Based on the first change condition, the candidate chemical formula features corresponding to the target chemical formula feature are determined, and the determined candidate chemical formula features are the result of the first expansion.
[0067] In the specific application of this embodiment, based on the aforementioned set of chemical change constraints for microbial secondary metabolism, a first change condition is determined. This condition specifically characterizes the constraints on element conservation, atomic ratio, and small molecule removal / introduction in the synthesis of the target chemical formula feature from the candidate chemical formula feature. Based on this first change condition, derivation and deduction are performed on the target chemical formula feature to generate multiple candidate chemical formulas that conform to the chemical change rules. After screening and eliminating invalid features, the resulting valid candidate chemical formulas are the first expansion result of the target chemical formula feature.
[0068] As a preferred embodiment of this invention, the first variation condition is specifically as follows:
[0069] Based on domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target chemical formula features and the corresponding candidate chemical formula features. This feature extraction includes at least the extraction of element conservation features, functional group derivation features, molecular skeleton splicing and cyclization features, small molecule removal features of condensation reactions, and structural unit homology features of secondary metabolism. The result of this feature extraction is defined as the first change condition.
[0070] In the specific application of this embodiment, based on the chemical change constraint rules in the field of microbial secondary metabolism, feature extraction is carried out on the potential chemical change rules between the target chemical formula features and the candidate chemical formula features. The element conservation features of precursors and products, functional group derivation features, molecular skeleton splicing and cyclization features, small molecule removal features of condensation reactions, and homology features of structural units specific to secondary metabolism are extracted in sequence. The various features extracted above are integrated and summarized to form a standardized feature rule set, and this feature rule set is directly defined as the first change condition.
[0071] Specifically, the second expansion is as follows:
[0072] Based on domain knowledge constraints, the second change condition corresponding to the target structural feature is determined. The second change condition is used to characterize the constraint rules of other candidate structural features to obtain the target structural feature.
[0073] Based on the second change condition, the candidate structural features corresponding to the target structural features are determined, and the determined candidate structural features are the result of the second expansion.
[0074] In the specific application of this embodiment, based on the aforementioned set of chemical change constraints for microbial secondary metabolism, a second change condition is determined. This condition specifically characterizes the constraints on the molecular skeleton evolution, functional group types and site changes, chemical bond formation / breaking, ring system configuration, and fusion mode of the synthesis of the target structural feature from the candidate structural feature. Based on this second change condition, derivation and extrapolation are performed on the target structural feature to generate multiple candidate structural features that conform to the chemical change rules. After screening and eliminating invalid features that do not conform to constraints such as skeleton splicing and functional group transformation, the resulting valid candidate structural features are the second expansion result of the target structural feature.
[0075] As a preferred embodiment of this invention, the second variation condition is specifically as follows:
[0076] Based on domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target structural features and the corresponding candidate structural features. This feature extraction includes at least the extraction of molecular skeleton splicing and rearrangement features, functional group transformation site features, cyclization and ring-opening configuration features, chemical bond breaking and bonding features, and side chain group spatial derivation features. The result of this feature extraction is defined as the second change condition.
[0077] In the specific application of this embodiment, based on the set of chemical change constraint rules in the field of microbial secondary metabolism, feature extraction is carried out on the potential chemical change rules between the target structural features and the candidate structural features. In sequence, molecular skeleton splicing and rearrangement features, functional group transformation site features, cyclization and ring-opening configuration features, chemical bond breaking and bonding features, and side chain group spatial derivation features are extracted. The various features extracted above are integrated and summarized to form a standardized set of feature rules, and this set of feature rules is directly defined as the second change condition.
[0078] It should be noted that the chemical formula and structural formula of a compound are two complementary key dimensions for characterizing its core chemical properties. They are related to each other and each has its own irreplaceable role: the chemical formula accurately reflects the quantitative essence of the compound, such as its elemental composition and atomic ratio, in a textual form, which is suitable for efficient textual feature analysis and matching with metabolic transformation rules; the structural formula, on the other hand, presents the qualitative differences such as the topology of the molecular skeleton, the spatial sites of functional groups, and the way chemical bonds are connected in a graphical form, which can capture structural-specific information that cannot be covered by the chemical formula. This application employs a differentiated processing path combining textual and image features for two types of features. It can not only uncover potential substitutes that structural features cannot cover but chemical features can effectively expand upon (such as cross-structure type substitutes based on element conservation and small molecule transformation) through textual feature processing, but also capture complementary substitutes that chemical features cannot cover but structural features can effectively expand upon (such as isomers and skeleton analogs) through image feature processing. The two form a two-way cross-validation and complementary coverage, which not only improves the reliability of potential substitutes, but also maximizes the discovery of non-overlapping multiple substitutes, significantly broadening the retrieval space and providing more comprehensive and three-dimensional feature support for subsequent precise retrieval based on knowledge graphs.
[0079] S23: The results of the first expansion and the second expansion are the candidate transformation features obtained by feature transformation processing.
[0080] S30: Based on the candidate transformation features, a search is performed in the retrieval knowledge graph. This search is based at least on similarity and attention weights, and the microorganisms corresponding to the candidate transformation features are output as alternatives to the microorganisms corresponding to the target secondary metabolites.
[0081] Reference Figure 3 , Figure 3 This is a flowchart illustrating the process of searching a knowledge graph based on candidate transformation features, as provided in an embodiment of this application.
[0082] Based on the aforementioned candidate transformation features, this embodiment will perform a precise search within the knowledge graph using these features to screen for microbial alternatives corresponding to the target secondary metabolites. The search first parses the candidate transformation features to obtain candidate chemical and structural formulas, and then gradually obtains the results through similarity analysis and attention weighting.
[0083] like Figure 3 As shown, specifically, the retrieval is performed in the knowledge graph based on the candidate transformation features, including:
[0084] S31: Analyze the candidate transformation features to obtain candidate chemical formula features and candidate structural formula features; wherein, the candidate chemical formula features are potential substitutes corresponding to the target chemical formula features, and the candidate structural formula features are potential substitutes corresponding to the target structural features.
[0085] S32: Based on the candidate chemical formula features and / or candidate structural formula features, perform similarity analysis in the retrieval knowledge graph, output the chemical formula features and / or structural features that meet the preset similarity threshold, and use the corresponding secondary metabolites and corresponding microorganisms as the first search results.
[0086] In the specific application of this embodiment, the candidate chemical formula features and candidate structural formula features obtained from the analysis are used as the retrieval basis. Similarity analysis is carried out separately or jointly in the retrieval knowledge graph. The analysis dimensions cover the matching degree of chemical formula elemental composition, the fit between structural skeleton and functional group, etc. The preset similarity threshold can be set according to the actual situation. Chemical formula features and / or structural formula features with a matching degree not lower than the threshold are selected and output. The secondary metabolites corresponding to these features and their associated microbial information are summarized as the first retrieval result.
[0087] S33: For the first search result, based on the preset attention weighting mechanism, analyze the overall probability value that the first search result can be converted into the target secondary metabolite, output the first search result that meets the preset overall probability value threshold as the second search result, and output the substitute for the microorganism corresponding to the target secondary metabolite corresponding to the second search result.
[0088] Reference Figure 4 , Figure 4 A flowchart illustrating the attention weighting mechanism provided in this application embodiment.
[0089] like Figure 4 As shown, in a preferred embodiment of this example, the attention weighting mechanism includes:
[0090] S3311: Initialize the attention weight allocation dimension, which corresponds to the synthesis matching dimension of the candidate chemical formula feature and the synthesis matching dimension of the candidate structural formula feature, respectively.
[0091] S3312: Calculate the synthesis and transformation matching degree between the candidate chemical formula feature and the target chemical formula feature, and the synthesis and transformation matching degree between the candidate structural formula feature and the target structural formula feature corresponding to the first search result, and use these as the initial attention weight values for each dimension.
[0092] S3313: Normalize the initial attention weight values for each dimension to obtain the normalized attention weights.
[0093] S3314: Based on the normalized attention weight, calculate the overall probability value of the synthesis of the target secondary metabolite corresponding to the first search result, compare the overall probability value with the overall probability value threshold, and output the first search result that meets the threshold requirement as the second search result.
[0094] In this specific application, the attention weight allocation dimensions are first initialized by setting weight dimensions corresponding to the synthesis matching dimensions of the candidate chemical formula features and the candidate structural formula features. Then, the synthesis conversion matching degree between the candidate chemical formula features and the target chemical formula features, and the synthesis conversion matching degree between the candidate structural formula features and the target structural formula features, are calculated in the first search results. These matching degree results are used as the initial attention weight values for the corresponding dimensions. Normalization is performed on the initial attention weight values for each dimension to obtain normalized attention weights. Based on these weights, a weighted calculation is performed on the two types of matching degrees to obtain the overall probability value of synthesizing the target secondary metabolite of the compound corresponding to the first search result. This probability value is compared with a preset threshold, and the first search result that meets the threshold requirement is selected and output as the second search result.
[0095] This embodiment obtains the second search result based on the aforementioned similarity analysis and attention weighting. However, traditional microbial alternative screening methods only focus on the chemical characteristic matching of secondary metabolites, without considering the time-dependent generation patterns of primary and secondary metabolites during the production of secondary metabolites by microorganisms. This can easily lead to selected strains having matching chemical characteristics but inconsistent metabolic rhythms, resulting in the inability to stably synthesize the target secondary metabolite and significantly reducing the practical applicability of the alternative bacteria. To further improve the accuracy of microbial alternative screening and ensure the metabolic compatibility between the alternative and target bacteria, it is necessary to introduce material temporal characteristics for secondary verification. This verification can accurately match strains with compatible metabolic mechanisms by characterizing the time-dependent generation patterns of secondary metabolites produced by microorganisms, ensuring that the alternative bacteria can stably produce the target secondary metabolite.
[0096] Reference Figure 5 , Figure 5 An optimized flowchart of the second search result output provided in the embodiments of this application.
[0097] like Figure 5 As shown, specifically, the output of the microbial alternatives corresponding to the target secondary metabolite of the second search result also includes:
[0098] S3321: Construct material temporal characteristics, which are used to characterize the time-series generation patterns of primary and secondary metabolites during the production of secondary metabolites by microorganisms.
[0099] S3322: Extract the material temporal characteristics of the target secondary metabolites and the material temporal characteristics of the microorganisms corresponding to the second search results;
[0100] S3323: Compare the material time sequence characteristics of the target secondary metabolite with the material time sequence characteristics of the microorganisms corresponding to the second search result, screen out the microorganisms that meet the preset material time sequence characteristic similarity threshold, and output them as alternatives to the microorganisms corresponding to the target secondary metabolite.
[0101] In this specific application, a material time-series feature is first constructed. Combined with data related to the microbial metabolic cycle, indicators such as the generation amount and rate of primary and secondary metabolites are integrated along the time dimension to form a material time-series feature system characterizing the temporal generation patterns of both. Subsequently, the material time-series features corresponding to the target secondary metabolite and the material time-series features corresponding to each microorganism in the second search results are extracted, and both types of features are converted into standardized feature vectors. Finally, a preset similarity algorithm is used to compare the similarity of the two types of standardized feature vectors. A preset material time-series feature similarity threshold is set, and microorganisms with a similarity not lower than this threshold are selected and output as replacements for the microorganisms corresponding to the target secondary metabolite.
[0102] Reference Figure 6 , Figure 6 A flowchart illustrating the construction of material timing features provided in this application embodiment.
[0103] like Figure 6 As shown, specifically, in a preferred embodiment of this example, the construction of material time sequence characteristics includes:
[0104] S33211: Defining the growth sequence stages in the process of microbial production of secondary metabolites;
[0105] S33212: Collect samples at each time stage, detect and obtain the types and abundance data of primary metabolites and secondary metabolites at the corresponding growth time stages;
[0106] S33213: Using growth time as the dimension, species and abundance data are correlated and integrated in chronological order to construct a time-series dataset;
[0107] S33214: Standardize the time series dataset to obtain the material time series characteristics.
[0108] In the specific application of this embodiment, the construction of material temporal characteristics is carried out as follows: First, based on the general laws of microbial growth and metabolism, the growth time sequence in the process of producing secondary metabolites is divided, for example, into four core time sequence stages: lag phase, logarithmic growth phase, stationary phase, and death phase. Samples are collected at fixed points in each time sequence stage, and liquid chromatography-mass spectrometry (LC-MS) is used for detection to obtain the specific types of primary metabolites (such as sugars and amino acids) and secondary metabolites in the corresponding stage, and the abundance data of each type of product are simultaneously statistically analyzed. Using growth time sequence as the dimension, the above-mentioned type and abundance data are correlated and integrated in chronological order to build a structured time sequence dataset. This time sequence dataset is then normalized to eliminate dimensional differences between different detection dimensions. After standardization, the material temporal characteristics are obtained.
[0109] It should be noted that the material time-series characteristics involved in this embodiment are constructed based on the inherent laws of microbial metabolism and conventional research techniques: the process of microorganisms producing secondary metabolites is inevitably accompanied by phased changes in the growth sequence. The types and abundances of primary and secondary metabolites will naturally exhibit regular fluctuations with stages such as the lag phase and logarithmic growth phase. Through mature detection techniques in this field, such as liquid chromatography-mass spectrometry, relevant data of each time-series stage can be accurately collected. The material time-series characteristics formed after standardization and integration have sufficient operability and reproducibility.
[0110] The core scientific basis for introducing the temporal characteristics of this substance for secondary verification is that the biosynthesis of the target secondary metabolite is not an isolated process. It depends on the basic elements such as carbon, hydrogen, and oxygen provided by the primary metabolite, as well as key precursor substances. The temporal characteristics of the substance can directly reflect whether microorganisms have the ability to continuously synthesize these basic metabolites. This is a prerequisite for ensuring that microorganisms can serve as effective substitutes. It can eliminate ineffective strains with matching chemical characteristics but lacking basic metabolic support from the metabolic mechanism level, and greatly improve the reliability of the actual application of the substitute.
[0111] It is worth noting that the technical solution of this embodiment is not limited to searching for microbial substitutes: the compounds corresponding to the candidate transformation features discovered during the search process can themselves serve as precursors, directly generating the target secondary metabolites through conventional methods such as chemical synthesis or biotransformation. This design breaks through the limitation of traditional technologies that only focus on microbial screening, providing a dual-pathway option for obtaining target secondary metabolites: microbial substitution and direct compound synthesis, further expanding the application scenarios and practical value of the technical solution.
[0112] In summary, the knowledge graph-based microbial secondary metabolite retrieval method of this embodiment has at least the following technological innovations:
[0113] We construct a retrieval knowledge graph that focuses on the association between key chemical features (chemical formulas and structural formulas) of secondary metabolites and microorganisms. With the direct mapping relationship between features and microorganisms as the core, we overcome the limitations of traditional methods that lack targeted screening guidance.
[0114] We propose a bidirectional feature expansion strategy based on knowledge constraints in the field of microbial secondary metabolism, which mines potential substitutes for chemical formula and structural features respectively, thus solving the problems of single feature dimensions and insufficient coverage of substitutes in traditional research.
[0115] A hierarchical retrieval mechanism combining similarity analysis and attention weighting is designed. By calculating and weighting the matching degree in two dimensions, the mechanism can accurately screen alternative bacteria and avoid the errors of traditional single-dimensional retrieval.
[0116] The study introduces material time-series characteristics to characterize the periodic generation patterns of primary and secondary metabolites, and adds a metabolic rhythm fit verification step to overcome the shortcomings of traditional methods that only focus on chemical characteristics and ignore the fit of metabolic mechanisms.
[0117] Based on the above-mentioned technological innovations, the knowledge graph-based microbial secondary metabolite retrieval method of this embodiment achieves at least the following technical effects:
[0118] By eliminating the reliance on symbiotic bacteria of rare medicinal plants, we can quickly identify suitable microbial substitutes for target secondary metabolites, significantly shorten the research cycle, and solve the problems of time-consuming and inefficient traditional methods.
[0119] By targeting and precisely screening, the proportion of invalid experiments is reduced, the use of expensive consumables and equipment such as chromatographic separation and spectroscopic identification is reduced, and research costs are significantly reduced.
[0120] By balancing chemical characteristic matching and metabolic rhythm adaptability, the reliability of alternative bacteria in practical applications is improved, ensuring that they can stably produce target secondary metabolites and effectively alleviate the contradiction between supply and demand of medicinal resources.
[0121] Expanding the screening scope of microbial alternatives reduces the omission of effective strains and provides sufficient strain resources to support the large-scale production of pharmaceutical active compounds.
[0122] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.
[0123] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1.A knowledge graph-based method for retrieving microbial secondary metabolites, characterized in that, The application is applied to the alternative of retrieving symbiotic bacteria of rare medicinal plants, and includes: Based on the key features of secondary metabolites and the corresponding microorganisms, a knowledge graph is constructed to obtain a retrieval knowledge graph, wherein the key features at least include chemical formula features and structural formula features of the secondary metabolites, and the retrieval knowledge graph takes the key features as input and outputs the microorganisms corresponding to the key features; The key features of the target secondary metabolites are processed by feature transformation, and the feature transformation is constrained by a preset domain knowledge to obtain a candidate transformed feature, which is a potential alternative of the key features; Based on the candidate transformed feature, a retrieval is performed in the retrieval knowledge graph, which is based at least on similarity and attention weight, and the microorganisms corresponding to the candidate transformed feature are output as the alternative of the microorganisms corresponding to the target secondary metabolites; The feature transformation processing includes: Obtaining the domain knowledge constraint, which is a constraint condition of chemical changes between secondary metabolite corresponding compounds; Based on the domain knowledge constraint, the key features of the target secondary metabolites are data augmented, which correspond to target chemical formula features and target structural formula features, and the data augmentation includes first augmentation for the target chemical formula features and second augmentation for the target structural formula features; The results of the first augmentation and the second augmentation are the candidate transformed features obtained by the feature transformation processing; The retrieval based on the candidate transformed feature in the retrieval knowledge graph includes: Analyzing the candidate transformed feature to obtain a candidate chemical formula feature and a candidate structural formula feature; wherein the candidate chemical formula feature is a potential alternative of the target chemical formula feature, and the candidate structural formula feature is a potential alternative of the target structural formula feature; Based on the candidate chemical formula feature and / or the candidate structural formula feature, a similarity analysis is performed in the retrieval knowledge graph, and chemical formula features and / or structural formula features satisfying a preset similarity threshold are output, and the corresponding secondary metabolites and corresponding microorganisms are taken as the first retrieval result; Based on the first retrieval result, a preset attention weight mechanism is used to analyze the overall probability value of the first retrieval result that can be converted into the target secondary metabolite, and the first retrieval result satisfying a preset overall probability value threshold is output as the second retrieval result, and the alternative of the microorganisms corresponding to the target secondary metabolite corresponding to the second retrieval result is output. 2.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 1, characterized in that, The first augmentation is specifically: Based on the domain knowledge constraint, a first change condition corresponding to the target chemical formula feature is determined, which is used to represent the constraint rule of other candidate chemical formula features to obtain the target chemical formula feature; Based on the first change condition, the candidate chemical formula features corresponding to the target chemical formula feature are determined, and the determined candidate chemical formula features are the results of the first augmentation. 3.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 1, characterized in that, The second augmentation is specifically: Based on the domain knowledge constraints, a second change condition corresponding to the target structural feature is determined. The second change condition is used to characterize the constraint rules for obtaining the target structural feature from other candidate structural features. Based on the second change condition, the candidate structural features corresponding to the target structural feature are determined, and the determined candidate structural features are the result of the second expansion. 4.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 2, characterized in that, The first change condition is specifically: Based on the aforementioned domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target chemical formula features and the corresponding candidate chemical formula features. This feature extraction includes at least the extraction of element conservation features, functional group derivation features, molecular skeleton splicing and cyclization features, small molecule removal features of condensation reactions, and structural unit homology features of secondary metabolism between precursors and products. The result of this feature extraction is defined as the first change condition. 5.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 3, characterized in that, The second change condition is as follows: Based on the aforementioned domain knowledge constraints, feature extraction is performed on the potential chemical change patterns between the target structural features and the corresponding candidate structural features. This feature extraction includes at least the extraction of molecular skeleton splicing and rearrangement features, functional group transformation site features, cyclization and ring-opening configuration features, chemical bond breaking and bonding features, and side chain group spatial derivation features. The result of this feature extraction is defined as the second change condition. 6.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 1, characterized in that, The attention weighting mechanism includes: Initialize the attention weight allocation dimension, which corresponds to the synthesis matching dimension of the candidate chemical formula feature and the synthesis matching dimension of the candidate structural formula feature, respectively; Calculate the synthesis and transformation matching degree between the candidate chemical formula feature and the target chemical formula feature, and the synthesis and transformation matching degree between the candidate structural formula feature and the target structural formula feature corresponding to the first search result, and use these as the initial attention weight values for each dimension. The initial attention weights for each dimension are normalized to obtain the normalized attention weights. Based on the normalized attention weight, the overall probability value of synthesizing the target secondary metabolite corresponding to the compound in the first search result is calculated in a weighted manner. The overall probability value is compared with the overall probability value threshold, and the first search result that meets the threshold requirement is output as the second search result. 7.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 1, characterized in that, The method of outputting the alternative terms for the microorganism corresponding to the target secondary metabolite corresponding to the second search result also includes: A material temporal feature is constructed, which is used to characterize the time-series generation pattern of primary and secondary metabolites during the production of secondary metabolites by microorganisms. Extract the material time sequence characteristics of the target secondary metabolites and the material time sequence characteristics of the microorganisms corresponding to the second search results; The material time sequence characteristics of the target secondary metabolite are compared with the material time sequence characteristics of the microorganisms corresponding to the second search result. Microorganisms that meet the preset material time sequence characteristic similarity threshold are selected and output as alternatives to the microorganisms corresponding to the target secondary metabolite. 8.The knowledge graph-based microbial secondary metabolite retrieval method according to claim 7, characterized in that, The construction of the material temporal characteristics includes: Divide the growth sequence stages in the process of microbial production of secondary metabolites; Collecting samples at each time sequence stage, detecting and obtaining the type and abundance data of primary metabolites and secondary metabolites corresponding to the growth time sequence stage; Integrating the type and abundance data in time sequence according to time sequence as a dimension to construct a time sequence dataset; Standardizing the time sequence dataset to obtain the material time sequence characteristics.
Citation Information
Patent Citations
Intestinal microorganism intelligent question-answering system based on knowledge graph and large model
CN119226487A
System and method for determining microbiome from host metabolome using a machine learning model
US20240404637A1