Analysis method and equipment for circulating free microbial community

By identifying species and extracting features from circulating microbial communities, a community feature matrix is ​​generated. Combined with host attribute information, a classifier is used to analyze whether tumor cells exist in the host, which solves the problem of insufficient sensitivity in early tumor detection in existing technologies and achieves comprehensiveness and accuracy in early cancer detection.

CN121905291APending Publication Date: 2026-04-21TIANJIN TUMOR HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN TUMOR HOSPITAL
Filing Date
2025-12-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, methods for early tumor diagnosis by detecting tumor-released ctDNA in the blood have insufficient sensitivity, especially when the ctDNA content is low in the early stages of tumors. This limits their application in the early detection of various types of tumors and also limits their accuracy in locating the tumor's origin tissue.

Method used

By identifying species and extracting features from circulating microbial communities, a community feature matrix is ​​generated. Combined with host attribute information, the presence of tumor cells in the host is analyzed. A classifier is used for analysis and tracing to obtain the presence of tumor cells and the location of the tissue of origin.

Benefits of technology

It improves the sensitivity and accuracy of early tumor detection, is applicable to a variety of tumor types, enables early cancer detection from the perspective of the microbial community, and enhances the comprehensiveness and accuracy of the analysis of the presence of tumor cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905291A_ABST
    Figure CN121905291A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for analyzing a circulating free microbial community, which can be applied to the technical field of artificial intelligence. The method comprises: based on community data of a to-be-detected microbial community, performing microbial species identification on the to-be-detected microbial community to obtain a species identification result of the to-be-detected microbial community, the to-be-detected microbial community being extracted from a host; based on a species identification result, performing species feature extraction on the community data to obtain respective species features of a plurality of target microorganisms; generating a community feature matrix of the to-be-detected microbial community based on the respective species features of the plurality of target microorganisms and the respective species category features of the plurality of target microorganisms; based on the community feature matrix and host attribute information of the host, whether tumor cells exist in the host or not is analyzed, and an analysis result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and specifically to a method and apparatus for analyzing circulating free microbial communities. Background Technology

[0002] Currently, early tumor diagnosis is crucial for cancer prevention and treatment. Related technologies typically generate diagnostic results about the tumor by detecting substances released by tumors in the blood, such as circulating tumor deoxyribonucleic acid (ctDNA).

[0003] In realizing the concept of this disclosure, it was found that the related technology has at least the following problems: the above methods rely on tumor burden, and the content of ctDNA is low in the early stage of tumors, resulting in insufficient detection sensitivity, which limits its application in the early detection of various types of tumors and its accuracy in locating the tumor origin tissue. Summary of the Invention

[0004] In view of the above problems, this disclosure provides a method, apparatus, equipment, media and program product for the analysis of circulating free microbial communities.

[0005] According to one aspect of this disclosure, a method for analyzing circulating free microbial communities is provided, comprising: identifying microbial species in the microbial community to be tested based on community data, obtaining species identification results for the microbial community to be tested, wherein the microbial community to be tested is extracted from a host; extracting species features from the community data based on the species identification results, obtaining species features of multiple target microorganisms; generating a community feature matrix of the microbial community to be tested based on the species features and species category features of the multiple target microorganisms; and analyzing the presence of tumor cells in the host based on the community feature matrix and host attribute information of the host, obtaining analysis results.

[0006] According to embodiments of this disclosure, the above analysis method further includes: when the analysis results indicate the presence of tumor cells in the host, performing source tracing analysis on the tumor cells based on the community feature matrix, the host attribute information of the host, and the analysis results, to obtain source tracing analysis results that locate the tissue of origin of the tumor cells.

[0007] According to embodiments of this disclosure, the above-mentioned extraction of species features from the community data based on the species identification results to obtain the species features of each of the multiple target microorganisms includes: determining the microbial data of each of the target microorganisms from the community data based on the species identification results and the species categories of each of the multiple target microorganisms; and extracting species features from the microbial data of each of the target microorganisms to obtain the species features of each of the target microorganisms.

[0008] According to embodiments of this disclosure, the aforementioned microbial data includes microbial species abundance; the aforementioned extraction of species features from the microbial data of each of the aforementioned target microorganisms to obtain the species features of each of the aforementioned target microorganisms includes: discretizing the aforementioned microbial species abundance based on a predetermined threshold to obtain an updated microbial species abundance; and extracting species features from the updated microbial species abundance of each of the aforementioned target microorganisms to obtain the species features of each of the aforementioned target microorganisms.

[0009] According to embodiments of this disclosure, a first classifier is used to analyze whether tumor cells exist in the host based on the aforementioned community feature matrix and the host attribute information of the host, thereby obtaining analysis results. The analysis method further includes: determining the test species characteristics and test species category characteristics of each microorganism in the test sample, generating a test community feature matrix of the test sample; inputting the test community feature matrix into the first classifier to obtain test analysis results; and using an interpretable model to determine the relationship between the test analysis results and the test species category characteristics in the test community feature matrix, thereby determining the species category characteristics of multiple target microorganisms from multiple test species category characteristics.

[0010] According to embodiments of this disclosure, the first classifier is trained through the following operations: based on a predetermined threshold, the abundance of sample microorganisms of each target microorganism in the training samples is discretized to obtain updated sample microorganism species abundance; features are extracted from the updated sample microorganism species abundance to obtain sample species features of each target microorganism; based on the sample species features and sample species category features of each microorganism, a sample community feature matrix of the training samples is generated; the sample community feature matrix and the sample host attribute information of the sample host are input into the initial first classifier to obtain sample analysis results; and based on the sample analysis results and the primary classification labels corresponding to the training samples, the parameters of the initial first classifier are tuned to obtain the first classifier.

[0011] According to embodiments of this disclosure, a second classifier is used to perform source tracing analysis on the tumor cells based on the aforementioned community feature matrix, the host attribute information of the aforementioned host, and the aforementioned analysis results, to obtain source tracing analysis results locating the tissue of origin of the aforementioned tumor cells; the second classifier is trained through the following operations: when the aforementioned sample analysis results indicate that the sample host of the aforementioned sample host has tumor cells, the aforementioned sample community feature matrix, the aforementioned sample host's sample host attribute information, and the aforementioned sample analysis results are input into an initial second classifier to obtain sample source tracing analysis results; and based on the sample source tracing analysis results and the secondary classification labels of the aforementioned training samples, the parameters of the initial second classifier are tuned to obtain the aforementioned second classifier.

[0012] According to embodiments of this disclosure, the above-mentioned identification of microbial species in the microbial community to be tested based on the community data of the microbial community to be tested, and obtaining the species identification result of the microbial community to be tested, includes: performing gene sequencing on the host cell to obtain gene sequencing data; and performing microbial species identification on the microbial community to be tested based on the gene sequencing data, and obtaining the species identification result of the microbial community to be tested.

[0013] According to embodiments of this disclosure, the above-mentioned gene sequencing of host cells to obtain gene sequencing data includes: extracting free deoxyribonucleic acid from the host cells to obtain gene data to be tested; and performing gene sequencing on the gene data to be tested based on a list of noisy genes to obtain the above-mentioned gene sequencing data.

[0014] The second aspect of this disclosure provides an analysis apparatus for circulating free microbial communities, comprising: an identification module for identifying microbial species in the microbial community to be tested based on community data, thereby obtaining a species identification result for the microbial community to be tested, wherein the microbial community to be tested is extracted from a host; an extraction module for extracting species features from the community data based on the species identification result, thereby obtaining species features of multiple target microorganisms; a generation module for generating a community feature matrix of the microbial community to be tested based on the species features and species category features of the multiple target microorganisms; and an analysis module for analyzing whether tumor cells exist in the host based on the community feature matrix and host attribute information of the host, thereby obtaining an analysis result.

[0015] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0016] A fourth aspect of this disclosure also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0017] The fifth aspect of this disclosure also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0018] According to embodiments of this disclosure, the microbial community to be tested is extracted from the host, and species identification, species feature extraction, and generation of a community feature matrix are performed. The presence of host tumor cells is then analyzed in conjunction with host attribute information. By employing species feature extraction from community data to obtain species characteristics related to species categories and forming a community feature matrix, and linking it to host attributes for tumor cell analysis, this approach at least partially addresses the sensitivity limitations of related technologies in early tumor detection. It achieves early tumor detection from a microbial community perspective, improves the comprehensiveness and accuracy of tumor cell presence analysis, and is applicable to various tumor types. Attached Figure Description

[0019] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0020] Figure 1 The illustration schematically depicts application scenarios of methods, apparatus, devices, media, and program products for analyzing circulating free microbial communities according to embodiments of the present disclosure;

[0021] Figure 2 A flowchart illustrating a method for analyzing circulating free microbial communities according to embodiments of the present disclosure is shown schematically.

[0022] Figure 3 This schematically illustrates a data flow diagram of training a first classifier according to an embodiment of the present disclosure;

[0023] Figure 4 A flowchart illustrating a method for analyzing circulating free microbial communities according to another embodiment of this disclosure is shown schematically;

[0024] Figure 5 A schematic diagram illustrating the structure of an analytical apparatus for circulating free microbial communities according to embodiments of the present disclosure is shown; and

[0025] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing an analytical method for circulating free microbial communities according to embodiments of the present disclosure. Detailed Implementation

[0026] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0030] In the technical solution disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.

[0031] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this disclosure all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0032] The research revealed that cancer is a significant public health issue, and early detection is crucial for improving patient survival and treatment outcomes. Traditional cancer screening methods, such as imaging and tissue biopsies, are often invasive, costly, or limited to specific cancer types. Liquid biopsy, a non-invasive technique, offers the potential for early detection by detecting substances released by tumors in the blood, such as ctDNA.

[0033] However, liquid biopsy methods in related technologies rely on tumor burden, and the ctDNA content is low in early-stage cancers, especially stage I, resulting in insufficient detection sensitivity and limiting their application in the early detection of multiple cancers. Furthermore, the accuracy of these methods in locating the tissue of origin of cancer is limited, making it difficult to meet clinical needs.

[0034] Therefore, embodiments of this disclosure provide a method for analyzing circulating free microbial communities, comprising: identifying microbial species in the microbial community to be tested based on community data, obtaining species identification results of the microbial community to be tested, wherein the microbial community to be tested is extracted from the host; extracting species features from the community data based on the species identification results, obtaining species features of multiple target microorganisms; generating a community feature matrix of the microbial community to be tested based on the species features and species category features of multiple target microorganisms; and analyzing the presence of tumor cells in the host based on the community feature matrix and host attribute information of the host, obtaining analysis results.

[0035] It should be noted that the methods for analyzing circulating free microbial communities involved in the embodiments of this disclosure are not intended for the diagnosis or treatment of cancer. Rather, they are used to obtain predictive analysis results on the presence of tumor cells in the host and source analysis results on the origin of tumor cells by analyzing the community data of the microbial community to be tested and the host attribute information, which are used to assist in the analysis of whether a patient has cancer and the origin of cancer tissue.

[0036] Figure 1The illustration schematically depicts application scenarios of methods, apparatus, devices, media, and program products for analyzing circulating free microbial communities according to embodiments of the present disclosure.

[0037] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0038] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0040] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0041] It should be noted that the method for analyzing circulating free microbial communities provided in this embodiment can generally be executed by server 105. Correspondingly, the device for analyzing circulating free microbial communities provided in this embodiment can generally be located in server 105. The method for analyzing circulating free microbial communities provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the device for analyzing circulating free microbial communities provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0042] It should be understood that Figure 1 The number of first terminal devices, second terminal devices, third terminal devices, networks, and servers shown in the diagram is merely illustrative. Depending on implementation needs, any number of first terminal devices, second terminal devices, third terminal devices, networks, and servers can be included.

[0043] The following will be based on Figure 1 The described scene, through Figures 2-4 The analytical method for circulating free microbial communities according to the disclosed embodiments is described in detail.

[0044] Figure 2 A flowchart illustrating a method for analyzing circulating free microbial communities according to an embodiment of the present disclosure is shown schematically.

[0045] like Figure 2 As shown, the method includes operations S210 to S240.

[0046] In operation S210, based on the community data of the microbial community to be tested, the microbial species of the microbial community to be tested are identified to obtain the species identification results of the microbial community to be tested, wherein the microbial community to be tested is extracted from the host.

[0047] In operation S220, based on the species identification results, species characteristics are extracted from the community data to obtain the species characteristics of multiple target microorganisms.

[0048] In operation S230, a community feature matrix of the microbial community to be tested is generated based on the species characteristics and species category characteristics of the multiple target microorganisms.

[0049] In operation S240, based on the community feature matrix and the host's host attribute information, the presence of tumor cells in the host is analyzed to obtain the analysis results.

[0050] The microbial community to be tested can be a microbial group extracted from the host. Community data can be a collection of various information about the microbial community, such as species composition information, sequence information, and taxonomic information.

[0051] The microbial community to be tested may include free microorganisms; the sequence information may include free microbiome DNA (cmDNA).

[0052] The host is not limited to any individual requiring analysis for the presence of tumor cells in their body. It is understood that in the medical field, tumor cells can be divided into benign and malignant tumor cells. Therefore, the analysis results can include the presence of tumor cells in the host, or the presence of malignant tumor cells, i.e., cancer cells.

[0053] Furthermore, it is understandable that if malignant tumor cells are confirmed in a host, the host can be considered to have a cancer risk. Conversely, if the absence of such cells, the host can be considered not to have a cancer risk.

[0054] Species identification results can include the species category and abundance of each microorganism or microbial data in the community data.

[0055] There are no restrictions on the method of identifying microbial species. Species can be identified by using sequence alignment or taxonomic classification tools to annotate the species categories of microorganisms. Statistical analysis can then be performed based on the identified species categories of each microorganism to obtain the species abundance, frequency of occurrence, etc.

[0056] There are no restrictions on the taxonomic classification tools, such as Kraken2 or Bracken.

[0057] There are no restrictions on the method of extracting species features from community data. Multi-dimensional species feature extraction can be performed on multiple microorganisms included in the community data, such as extracting read counts of the taxa to which each microorganism belongs, and extracting differential features of each microorganism.

[0058] The differential characteristics can be obtained by comparing the abundance differences between the current community data and the community data of different control groups.

[0059] The filling positions of each species category feature in the community feature matrix can be preset, and the corresponding species features can be filled into those positions according to the preset filling positions. The corresponding species feature can be a species feature that belongs to the same target microorganism as the species category feature at that filling position.

[0060] In some embodiments, the species category features of multiple target microorganisms can be output as species category feature representations of multiple target microorganisms through an embedding layer, and the species features of multiple target microorganisms can be projected onto the same latent dimension through a linear layer to create species feature representations. By combining the two feature representations element-wise, the final microbial feature embedding, i.e., the community feature matrix, is formed.

[0061] There are no restrictions on host attribute information; it can include basic host information such as age and gender.

[0062] The presence of tumor cells in a host can be analyzed using a community feature matrix and host attribute information. The analysis method is not limited; it can be obtained by classification using a first classifier or by joint analysis of microbial features and host attribute information through a multi-group data fusion algorithm.

[0063] In some embodiments, after obtaining the species identification results, the species characteristics of each of the multiple target microorganisms, the community characteristic matrix, and the analysis results, they can be stored in a database. The analysis results can then be displayed on a screen.

[0064] According to embodiments of this disclosure, by extracting the microbial community from the host, species identification and species feature extraction are sequentially performed to generate a community feature matrix, and the presence of host tumor cells is analyzed in conjunction with host attribute information. Because this method utilizes species feature extraction from community data to obtain species characteristics related to species categories, forming a community feature matrix, and then correlates it with host attributes for tumor cell analysis, it at least partially addresses the sensitivity limitations of related technologies in early cancer detection. It achieves early cancer detection from a microbial community perspective, improves the comprehensiveness and accuracy of tumor cell presence analysis, and is applicable to multiple cancer types.

[0065] According to embodiments of this disclosure, the above-described analysis method also includes operable methods.

[0066] When the analysis results indicate the presence of tumor cells in the host, the tumor cells are traced back to their origin based on the community feature matrix, the host's host attribute information, and the analysis results, thus obtaining the traced origin analysis results that locate the tissue of origin of the tumor cells.

[0067] Source tracing analysis results can include information such as the tissue of origin of the tumor cells. These results can help pinpoint the location of the lesion in the host.

[0068] If the analysis results indicate the presence of tumor cells in the host, source tracing analysis can be performed on these tumor cells. The method of source tracing analysis is not limited; it can be obtained by using a second classifier, or by using a joint analysis algorithm to analyze the community feature matrix, host attribute information, and analysis results.

[0069] According to embodiments of this disclosure, when tumor cells are found to be present in the host, a source tracing analysis is performed based on the community feature matrix, the host's host attribute information, and the analysis results to obtain the source tracing analysis results of the originating tissue. This can be combined with the characteristics of the originating tissue and the biological behavior of the tumor to more accurately determine the disease progression trend, so as to provide a basis for targeted diagnostic plans.

[0070] According to embodiments of this disclosure, based on species identification results, species characteristics are extracted from community data to obtain the species characteristics of multiple target microorganisms, which may include the following operations.

[0071] Based on the species identification results and the species categories of multiple target microorganisms, the microbial data of each target microorganism is determined from the community data; and the species characteristics of each target microorganism are extracted from the microbial data of each target microorganism to obtain the species characteristics of each target microorganism.

[0072] For hosts with different host attribute information, species categories that are more relevant or related to the analysis results can be pre-determined, thereby identifying more accurate microbial data for the target microorganism.

[0073] For example, for hosts of different ages and genders, existing research or clinical data can be used to screen for species of microorganisms that are more associated with tumor development and progression in those hosts, thereby reducing interference from irrelevant microbial data.

[0074] When extracting species characteristics from the microbial data of each target microorganism, multi-dimensional analysis can be performed on the selected target microorganism data. For example, extracting read counts of the taxa to which each microorganism belongs, and extracting differential features of each target microorganism.

[0075] According to embodiments of this disclosure, by selecting target microorganism data from community data based on their respective species categories, more effective data can be filtered out. Feature extraction is then performed on the filtered data, thereby improving data dimensionality reduction and reducing resource waste. Furthermore, by reducing noisy or invalid data, the accuracy of the obtained species characteristics of each target microorganism can be improved.

[0076] According to embodiments of this disclosure, the microbial data includes microbial species abundance; extracting species characteristics from the microbial data of each target microorganism to obtain the species characteristics of each target microorganism may include the following operations.

[0077] The microbial species abundance is discretized based on a predetermined threshold to obtain the updated microbial species abundance; and species characteristics are extracted from the updated microbial species abundance of each target microorganism to obtain the species characteristics of each target microorganism.

[0078] The method for discretizing microbial species abundance is not limited; histogram optimization can be used to discretize continuous features. By binning continuous microbial species abundance and generating discrete feature histograms, memory usage can be reduced and computation accelerated.

[0079] Multiple predetermined thresholds can be preset, and the abundance of each target microorganism can be compared with these thresholds. Target microorganisms that are less than a first predetermined threshold and greater than a second predetermined threshold can be classified into a microbiome corresponding to a predetermined range. The predetermined range is composed of the first and second predetermined thresholds.

[0080] Statistical features can be extracted from target microorganisms belonging to the same microbiome, and the extracted statistical features can be used as the updated microbial species abundance. Statistical features can include, for example, mean abundance, median abundance, or abundance standard deviation.

[0081] Based on this, species characteristics can be extracted from the updated microbial species abundance of each target microorganism, thereby obtaining species characteristics that can reflect the distribution pattern and biological significance of the abundance of each target microorganism.

[0082] According to embodiments of this disclosure, when the abundance of microbial species is discretized based on a predetermined threshold, continuous abundance data can be transformed into structured information, thereby reducing the dimensionality of the original data to reduce memory usage and accelerate subsequent calculations. At the same time, the species features extracted based on the discretization results can better reflect actual biological significance, such as the clinical correlation between abundance levels, thus providing more efficient input data for subsequent community analysis and disease association model construction.

[0083] According to embodiments of this disclosure, a first classifier is used to analyze whether tumor cells exist in the host based on the community feature matrix and the host attribute information of the host, and the analysis results are obtained; the above analysis method may also include the following operations.

[0084] The test species characteristics and test species category characteristics of each microorganism in the test sample are determined, and a test community feature matrix of the test sample is generated. The test community feature matrix is ​​input into the first classifier to obtain the test analysis results. The relationship between the test analysis results and the test species category characteristics in the test community feature matrix is ​​determined using an interpretable model, and the species category characteristics of multiple target microorganisms are determined from multiple test species category characteristics.

[0085] The first classifier is not limited and can be a Light Gradient Boosting Machine Classifier (LightGBM). By inputting the community feature matrix and host attribute information into the first classifier, an analysis result indicating the presence of tumor cells in the host can be output.

[0086] In some embodiments, the species category characteristics of multiple target microorganisms that are more relevant or correlated with the analysis results can be determined by performing interpretability analysis on the analysis results of the first classifier.

[0087] For example, the same microbial species identification and feature extraction methods used for community data can be employed to obtain the test species characteristics and test species category characteristics of each microorganism in the test sample. Then, based on the same method used to generate the community feature matrix of the microbial community to be tested, which is based on the species characteristics and category characteristics of multiple target microorganisms, a test community feature matrix is ​​generated.

[0088] The test analysis results are obtained by inputting the test community feature matrix and the host attribute information of the test host of the test samples into the first classifier. The classification results are then analyzed using interpretable models such as Shapley Additive Explanations (SHAP) value analysis or Local Interpretable Model-agnostic Explanations (LIME) interpreters. This allows for the quantification of the contribution of each test species category feature to the test classification results, thus determining the correlation strength between the test analysis results and the test category features in the test community feature matrix.

[0089] Furthermore, based on the contribution ranking, test species category features that have a high impact on the test analysis results can be screened out, and these test species category features can be used as the species category features of multiple target microorganisms that are more relevant to the analysis results.

[0090] According to embodiments of this disclosure, a test community feature matrix is ​​constructed by integrating the test species characteristics and test species category characteristics of each microorganism in the test sample. This matrix is ​​then input into a first classifier, which outputs the test analysis results. Simultaneously, an interpretable model is introduced to determine the correlation between the test analysis results and the test species category characteristics, more accurately identifying species category characteristics that have a high impact on the analysis results. This identifies key microbial characteristics driving cancer detection, providing more effective data support for subsequent analysis of whether the host contains tumor cells and for tracing the origin of tumor cells.

[0091] According to embodiments of this disclosure, the above steps can facilitate the efficient processing of large-scale microbial data by the first classifier and enable higher-performance early cancer detection.

[0092] In some embodiments, microbial diversity and community analysis can be performed on test samples to assess differences in microbial composition and structural characteristics among samples. For example, at the species level, alpha diversity (α-Diversity) is calculated using the vegan R package (v2.6.10), including the Shannon Diversity Index (SDI) and Observed Species Richness (OSR). Beta diversity (β-Diversity) is assessed by calculating the Euclidean Distance (ED) of the relative abundance table using the `vegdist` function. Community structure is visualized using Principal Coordinates Analysis (PCoA) and Non-metric Multidimensional Scaling (NMDS), implemented using the `cmdscale` and `metaMDS` functions, respectively. The composition of microbial gradients is projected onto the NMDS space using the `envfit` function to show their orientation. The statistical differences in α-diversity between groups were assessed using the Wilcoxon Rank-Sum Test (WRT) (for comparisons between two groups), while the differences in β-diversity were tested using the adonis2 function in the vegan package with 999 permutations in a permutational analysis of variance (PERMANOVA).

[0093] In some embodiments, the microbial load and community composition of test samples can be assessed to quantify the total number of microorganisms and their phylum-level composition characteristics. For example, bacterial load (BL) is calculated by comparing the horizontally allocated microbial reads with the total number of sequencing reads after quality control. The conversion ratio is used for estimation. Community composition is characterized based on relative abundance features derived from a normalized abundance table. Taxonomic groups are aggregated at the phylum level, with relative abundance calculated in R using the `decostand` function from the `vegan` R package, and the normalization method set to "total". The seven most abundant phyla are selected based on the cumulative abundance of all samples, and the remaining phyla are merged into the "Other" category by summing their relative abundance.

[0094] In some embodiments, differential abundance and association analyses of the test samples can also be performed to identify differentially abundant microorganisms between groups and analyze their relationship with host attribute information. For example, two methods are used to identify taxa with differential abundance between groups. For multi-group comparisons (e.g., 13 cancer types), the Linear Discriminant Analysis Effect Size (LEfSe) method in the microeco R package is used, with a Linear Discriminant Analysis (LDA) score threshold ≥2.0 and a Wilcoxon p-value (WP) <0.05. A Microbiome Multivariable Association with Linear Models (MaAsLin2) model is used to model the multivariate association between microbial species and host attribute information, such as age, sex, and disease state, using log-transformed relative abundance data in the analysis.

[0095] Figure 3 The diagram illustrates a data flow graph for training a first classifier according to an embodiment of the present disclosure.

[0096] like Figure 3 As shown, training the first classifier may include operations S310 to S350.

[0097] In operation S310, based on a predetermined threshold, the abundance of sample microorganisms of each target microorganism in the training samples is discretized to obtain the updated abundance of sample microorganisms.

[0098] In operation S320, feature extraction is performed on the abundance of microbial species in the updated sample to obtain the sample species characteristics of each target microorganism.

[0099] In operation S330, a sample community feature matrix of training samples is generated based on the sample species characteristics and sample species category characteristics of each microorganism.

[0100] In operation S340, the sample community feature matrix and the sample host attribute information of the sample host are input into the initial first classifier to obtain the sample analysis results.

[0101] In operation S350, based on the sample analysis results and the primary classification labels corresponding to the training samples, the parameters of the initial first classifier are tuned to obtain the first classifier.

[0102] The same method used to determine the updated microbial species abundance can be employed to determine the updated sample microbial species abundance, which will not be elaborated upon here.

[0103] The same method used to determine the community feature matrix as described above can be used to determine the sample community feature matrix, which will not be elaborated here.

[0104] The same methods used to determine the species characteristics mentioned above can be employed to determine the species characteristics of the sample, which will not be elaborated upon here.

[0105] In some embodiments, the abundance of microbial species in the sample of each target microorganism can be binned, a histogram can be constructed, and the gradient and second derivative of each bin can be calculated to determine the optimal split point, thereby reducing the amount of computation and improving the splitting efficiency.

[0106] When tuning the initial first classifier, the loss function can be determined based on the sample analysis results and the primary classification labels corresponding to the training samples, and the initial first classifier can be tuned based on the loss function.

[0107] Specifically, when tuning the initial first classifier, it can be initialized first. Taking a tree-based initial first classifier as an example, a decision tree can be initialized first, setting the initial predicted values ​​to the mean of the target variable or a certain constant, and presetting key hyperparameters, including a learning rate of 0.1, a number of leaf nodes of 31, a feature sampling ratio of 0.85, and a data sampling ratio of 0.8. Simultaneously, an improved ensemble learning algorithm based on gradient boosting decision trees can be selected to lay the foundation for subsequent parameter tuning.

[0108] Furthermore, when adjusting parameters, decision trees can be iteratively constructed. For example, a leaf-order growth strategy can be used to calculate the gradient of the loss function, which is the residual predicted by the current model. Based on the residual, the node that contributes the most to the loss function can be selected for splitting, and a histogram-optimized decision tree can be constructed, generating a new tree in each iteration.

[0109] In each iteration, a portion of the sample community feature matrix and sample host attribute information can be randomly sampled to reduce the risk of overfitting and accelerate training. For example, training can be performed using a feature score of 0.85 and a bagging score of 0.8.

[0110] There are no restrictions on the loss function used; logarithmic loss can be used. Furthermore, gradient boosting can be used to update the model, gradually reducing residuals and optimizing overall prediction performance.

[0111] In some embodiments, 5-fold cross-validation can be applied during training to monitor performance on the validation set. If there is no performance improvement after 100 consecutive rounds, early stopping is triggered to prevent overfitting.

[0112] In some embodiments, the first classifier may output the presence or absence of tumor cells in the host, or the presence or absence of cancer.

[0113] In some embodiments, a sensitivity of approximately 87% and its 95% confidence interval can be calculated based on a pre-set 99.5% specificity threshold.

[0114] According to embodiments of this disclosure, by discretizing and extracting features of the abundance of target microorganisms in training samples, and combining host attribute information to construct a sample community feature matrix for classifier training and parameter tuning, the learning ability of the first classifier to learn the association pattern between microbial features and host attributes is improved, thereby enhancing the model's classification accuracy and generalization performance.

[0115] According to embodiments of this disclosure, a second classifier is used to perform source tracing analysis on tumor cells based on the community feature matrix, host attribute information of the host, and analysis results, to obtain source tracing analysis results that locate the tissue of origin of the tumor cells; the second classifier is trained through the following operations.

[0116] When the sample analysis results indicate that the sample host contains tumor cells, the sample community feature matrix, the sample host attribute information, and the sample analysis results are input into the initial second classifier to obtain the sample source tracing analysis results; and based on the sample source tracing analysis results and the secondary classification labels of the training samples, the parameters of the initial second classifier are tuned to obtain the second classifier.

[0117] There are no restrictions on the second classifier; it can be a PyTorch-based Transformer architecture, a neural network model, etc.

[0118] The first classifier and the second classifier can be trained jointly. After the first classifier is trained, the second classifier can be trained based on the sample analysis results output by the first classifier, combined with the sample community feature matrix and the sample host attribute information of the sample host.

[0119] Sample host attribute information can include sample host age and sample host gender, which are transformed through an embedding layer and represented as follows: and The sample analysis results can be transformed into the latent space using a fully connected layer, and the resulting embedding representation can be... .

[0120] By generating multiple target microorganisms through an embedding layer The abundance of multiple target microorganisms is then projected onto the same latent dimension through a linear layer to create abundance embeddings. These two representations are combined element-wise to form the sample community feature matrix of the target microorganism. All processed feature embeddings are concatenated along the feature dimension to obtain the input representation of the second classifier: The embeddings of the connections can be processed by the Transformer encoder to generate the final multi-class predictions.

[0121] Performance was evaluated using overall accuracy, precision per class, recall, F1 score, and top k accuracy (k=1~5).

[0122] The Transformer model can be trained in PyTorch for 100 epochs with a batch size of 16. The optimizer can be an Adaptive Moment Estimation Weight Decay Optimizer (AdamW). This optimizer is configured with a learning rate and weight decay factor for the training process, and the beta value can be set to (0.9, 0.95). To dynamically control the learning rate during training, a Plateau decay scheduler can be introduced. This scheduler adjusts the learning rate by monitoring the validation loss; if there is no improvement after 10 consecutive epochs, the learning rate is reduced by a factor of 0.5, and the learning rate must not fall below 1e-7.

[0123] According to embodiments of this disclosure, when sample analysis results indicate the presence of tumor cells in the host, secondary classification can be performed by introducing a sample community feature matrix, host attribute information, and preliminary analysis results, thereby improving the accuracy and specificity of source tracing analysis. Furthermore, the initial second classifier is tuned using secondary classification labels, and through joint training, the model's ability to identify tumors from different sources or types can be further optimized.

[0124] In some embodiments, a prospective cohort analysis can be performed to validate the performance of the first classifier and analyze cancer risk associations. The locked first classifier is applied to the prospective cohort, and its performance is compared with models using common protein biomarkers such as alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), and carbohydrate antigen 199 (CA199) by area under the curve (AUC). Time-dependent receiver operating characteristic (timeROC) analysis is performed using the timeROC R package to assess the predictive accuracy at different lead times before clinical diagnosis.

[0125] In some embodiments, time-event analysis can be performed to assess the association between species characteristics of cmDNA and the risk of new-onset cancer. For each cmDNA characteristic, the sample hosts are divided into high-risk and low-risk groups based on the median abundance of the cmDNA characteristic in the study cohort. Univariate and multivariate Cox Proportional Hazards Regression Models are used to model the time to first cancer diagnosis. Multivariate analysis includes confounding variables such as age, sex, and smoking status. The cumulative hazard curve is estimated from the fitted Cox Proportional Hazards model using the `survival` package, visualized using the `survminer` package, and the difference in cumulative incidence between the high- and low-risk groups is assessed using the log-rank test.

[0126] In some embodiments, to study the performance of the first classifier in sample host groups, such as control groups, high-risk control groups, and cancer case groups, gene set enrichment analysis (GSEA) is used to sort the sample analysis results of the first classifier according to the inter-group difference values. Accuracy is then tested.

[0127] According to embodiments of this disclosure, based on community data of the microbial community to be tested, microbial species identification is performed on the microbial community to be tested to obtain species identification results of the microbial community to be tested, which may include the following operations.

[0128] Gene sequencing is performed on the host cells to obtain gene sequencing data; based on the gene sequencing data, the microbial species of the microbial community to be tested are identified to obtain the species identification results of the microbial community to be tested.

[0129] When performing gene sequencing on host cells, suitable host samples such as tissues, blood, or secretions can be selected, and total nucleic acids can be extracted. Sequencing technologies, such as high-throughput sequencing or low-coverage whole-genome sequencing, are then used to generate raw gene sequencing data containing both the host genome and the microbial genome.

[0130] When identifying species in a microbial community based on gene sequencing data, the raw sequencing data can be preprocessed, such as by removing sequencing adapters, filtering low-quality sequences, and removing host genome sequences, to obtain a microbial sequence dataset.

[0131] Microbial sequences can be compared with known microbial reference databases, such as databases containing characteristic genes or whole genome sequences of each species. Sequence similarity search algorithms can be used to determine the taxonomic affiliation of each sequence by calculating the similarity between sequences. This yields species identification results containing the names and relative proportions of each species in the microbial community being tested.

[0132] According to embodiments of this disclosure, species identification of the microbial community to be tested based on gene sequencing data and obtaining species identification results can more accurately analyze the types of microorganisms contained in the community and their taxonomic information.

[0133] According to embodiments of this disclosure, gene sequencing of host cells to obtain gene sequencing data may include the following operations.

[0134] Free deoxyribonucleic acid (DNA) was extracted from host cells to obtain the gene data to be tested; and gene sequencing was performed on the gene data to be tested based on a list of noisy genes to obtain gene sequencing data.

[0135] When performing gene sequencing on gene data to be tested based on a list of noisy genes, the range of target genes to be detected can be determined based on the list of noisy genes, and specific sequencing probes or primers can be determined based on this, so as to ensure that the sequencing process can specifically capture the sequence information of target genes other than those in the list of noisy genes.

[0136] Furthermore, during the sequencing process, sequencing parameters that match the target gene can be set to prioritize the acquisition of the target gene's sequence signal. After sequencing is completed, the raw sequencing data is preliminarily filtered to remove low-quality reads and adapter sequences, thereby obtaining gene sequencing data from a noise-removed gene list.

[0137] In some embodiments, noise control can be performed using the decontamR package to remove potentially contaminating taxa based on the concentration of each microorganism and a predefined list of noise genes.

[0138] The noise gene list can include host gene data. Therefore, the gene sequence data of microorganisms can be obtained through the noise gene list.

[0139] According to embodiments of this disclosure, gene sequencing is performed on the gene data to be tested based on a list of noisy genes to obtain gene sequencing data. This can reduce interference from non-noisy gene sequences by selectively filtering out noisy gene regions. Simultaneously, it can reduce the amount of irrelevant data generated, thus reducing the burden of subsequent data storage and analysis.

[0140] Figure 4 A flowchart illustrating a method for analyzing circulating free microbial communities according to another embodiment of this disclosure is shown schematically.

[0141] like Figure 4 As shown, the method includes operations S410 to S470.

[0142] In operation S410, based on the community data of the microbial community to be tested, the microbial species of the microbial community to be tested are identified to obtain the species identification results of the microbial community to be tested, wherein the microbial community to be tested is extracted from the host.

[0143] In operation S420, based on the species identification results and the species categories of multiple target microorganisms, the microbial data of each target microorganism is determined from the community data.

[0144] In operation S430, species characteristics are extracted from the microbial data of each target microorganism to obtain the species characteristics of each target microorganism.

[0145] In operation S440, based on the community feature matrix and the host's host attribute information, the presence of tumor cells in the host is analyzed, and the analysis results are obtained.

[0146] In operation S450, determine whether the analysis results indicate the presence of tumor cells.

[0147] When operating the S460, if the analysis results indicate that there are no tumor cells in the host, the analysis results are output to the display interface.

[0148] When operating S470, if the analysis results indicate the presence of tumor cells in the host, the tumor cells are traced back to their origin based on the community feature matrix, the host's host attribute information, and the analysis results. The traced-back analysis results are then used to locate the tissue of origin of the tumor cells, and the analysis results and traced-back analysis results are output to the display interface.

[0149] Based on the above-described method for analyzing circulating free microbial communities, this disclosure also provides an analytical apparatus for circulating free microbial communities. The following will be combined with... Figure 5 The device is described in detail.

[0150] Figure 5 A schematic block diagram of an analytical apparatus for circulating free microbial communities according to an embodiment of the present disclosure is shown.

[0151] like Figure 5 As shown, the analysis device 500 for circulating free microbial communities in this embodiment includes an identification module 510, an extraction module 520, a generation module 530, and an analysis module 540.

[0152] The identification module 510 is used to identify microbial species in the microbial community to be tested based on the community data, and to obtain the species identification result of the microbial community to be tested, wherein the microbial community to be tested is extracted from the host. In one embodiment, the identification module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0153] The extraction module 520 is used to extract species characteristics from community data based on species identification results, thereby obtaining the species characteristics of multiple target microorganisms. In one embodiment, the extraction module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0154] The generation module 530 is used to generate a community feature matrix of the microbial community to be tested based on the species characteristics and species category characteristics of the multiple target microorganisms. In one embodiment, the generation module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0155] Analysis module 540 is used to analyze whether tumor cells exist in the host based on the community feature matrix and the host's host attribute information, and obtain analysis results. In one embodiment, analysis module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0156] According to embodiments of this disclosure, any plurality of modules among the identification module 510, extraction module 520, generation module 530, and analysis module 540 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the identification module 510, extraction module 520, generation module 530, and analysis module 540 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the identification module 510, extraction module 520, generation module 530, and analysis module 540 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0157] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing an analytical method for circulating free microbial communities according to embodiments of the present disclosure.

[0158] like Figure 6 As shown, an electronic device 600 according to an embodiment of this disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this disclosure.

[0159] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 602 and / or RAM 603. It should be noted that programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.

[0160] According to embodiments of this disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0161] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0162] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0163] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the method for analyzing circulating free microbial communities provided in embodiments of this disclosure.

[0164] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0165] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0166] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0167] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0169] Those skilled in the art will understand that the features described in the various embodiments of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0170] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A method for analyzing circulating free microbial communities, characterized in that, The analytical method includes: Based on the community data of the microbial community to be tested, the microbial species of the microbial community to be tested are identified to obtain the species identification results of the microbial community to be tested, wherein the microbial community to be tested is extracted from the host. Based on the species identification results, species characteristics are extracted from the community data to obtain the species characteristics of each of the multiple target microorganisms; Based on the species characteristics and species category characteristics of each of the target microorganisms, a community feature matrix of the microbial community to be tested is generated; Based on the community feature matrix and the host attribute information of the host, the presence of tumor cells in the host is analyzed to obtain the analysis results.

2. The analytical method according to claim 1, characterized in that, The analytical method further includes: If the analysis results indicate the presence of tumor cells in the host, the tumor cells are traced back to their origin based on the community feature matrix, the host attribute information of the host, and the analysis results, thereby obtaining the traced origin analysis results that locate the tissue of origin of the tumor cells.

3. The analytical method according to claim 1, characterized in that, Based on the species identification results, species characteristics are extracted from the community data to obtain the species characteristics of multiple target microorganisms, including: Based on the species identification results and the respective species categories of the multiple target microorganisms, microbial data for each target microorganism is determined from the community data; and Species characteristics are extracted from the microbial data of each target microorganism to obtain the species characteristics of each target microorganism.

4. The analytical method according to claim 3, characterized in that, The microbial data includes microbial species abundance; The process of extracting species characteristics from the microbial data of each target microorganism to obtain the species characteristics of each target microorganism includes: The microbial species abundance is discretized based on a predetermined threshold to obtain the updated microbial species abundance; and Species characteristics are extracted from the updated microbial species abundance of each target microorganism to obtain the species characteristics of each target microorganism.

5. The analytical method according to claim 1, characterized in that, The first classifier is used to analyze whether tumor cells exist in the host based on the community feature matrix and the host attribute information of the host, and the analysis results are obtained. The analytical method further includes: The test species characteristics and test species category characteristics of each microorganism in the test sample are determined, and the test community characteristic matrix of the test sample is generated. The test community feature matrix is ​​input into the first classifier to obtain the test analysis results; and An interpretable model is used to determine the relationship between the test analysis results and the test species category features in the test community feature matrix, and to determine the species category features of multiple target microorganisms from multiple test species category features.

6. The analytical method according to claim 5, characterized in that, The first classifier is trained through the following operations: Based on a predetermined threshold, the abundance of microbial species in each target microorganism in the training samples is discretized to obtain the updated abundance of microbial species in the samples. Feature extraction is performed on the updated sample microbial species abundance to obtain the sample species characteristics of each target microorganism; Based on the sample species characteristics and sample species category characteristics of each microorganism, a sample community feature matrix of the training samples is generated; The sample community feature matrix and the sample host attribute information of the sample host are input into the initial first classifier to obtain the sample analysis results; as well as Based on the sample analysis results and the primary classification labels corresponding to the training samples, the parameters of the initial first classifier are tuned to obtain the first classifier.

7. The analytical method according to claim 6, characterized in that, Using a second classifier based on the community feature matrix, the host attribute information of the host, and the analysis results, the tumor cells are subjected to source tracing analysis to obtain source tracing analysis results that locate the tissue of origin of the tumor cells; The second classifier is trained through the following operations: If the sample analysis results indicate that the sample host contains tumor cells, the sample community feature matrix, the sample host attribute information, and the sample analysis results are input into the initial second classifier to obtain the sample source tracing analysis results. as well as Based on the sample source analysis results and the secondary classification labels of the training samples, the parameters of the initial second classifier are tuned to obtain the second classifier.

8. The analytical method according to claim 1, characterized in that, The method of identifying microbial species in the microbial community based on the community data of the microbial community to be tested, and obtaining the species identification results of the microbial community to be tested, includes: Gene sequencing is performed on host cells to obtain gene sequencing data; Based on the gene testing data, microbial species are identified in the microbial community to be tested, and the species identification results of the microbial community to be tested are obtained.

9. The analytical method according to claim 8, characterized in that, The gene sequencing of the host cell to obtain gene sequencing data includes: Free deoxyribonucleic acid was extracted from the host cells to obtain the gene data to be tested; and Based on the list of noisy genes, gene sequencing is performed on the gene data to be tested to obtain the gene sequencing data.

10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.