Intestinal flora data analysis system and method, electronic equipment and storage medium

Through the blocked parallel upload and automated processing process of the data analysis system, the slow upload and complex analysis of intestinal microbial sequencing data are solved, efficient and automated intestinal microbial sequencing data analysis is achieved, and visual analysis reports are generated, which improves the efficiency and value of intestinal microbial research.

CN120387149APending Publication Date: 2025-07-29BEIJING FUMART BIOTECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510751292.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the upload speed of intestinal microbial sequencing data is slow, the analysis process is complex and requires a lot of manual intervention, resulting in low processing efficiency and lack of in-depth analysis, which limits the efficiency and value of intestinal microbial research.

Method used

It provides a data analysis system for intestinal flora, including a data upload module, a process setting module, a data analysis module and a data mining module. Through block parallel upload, automated processing process and batch processing, combined with the data mining module, in-depth analysis is carried out to generate a visual analysis report.

Benefits of technology

It improves data upload efficiency, realizes the full process of automated processing of intestinal microbial sequencing data, improves analysis efficiency and depth, generates visual analysis reports, and enhances the value of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387149A_ABST
    Figure CN120387149A_ABST
Patent Text Reader

Abstract

The invention provides an intestinal flora data analysis system and method, electronic equipment and a storage medium, and the method comprises a data uploading module which is used for efficiently uploading to-be-processed intestinal flora sequencing data, and improving the uploading rate; the flow setting provides an automatic whole flow for the processing process of the intestinal flora sequencing data, and manual intervention is not needed; the data analysis module automatically analyzes the intestinal flora sequencing data uploaded by the data uploading module based on the processing flow formed by the flow setting module; the data mining module is used for performing further data mining on an analysis result obtained by analysis of the data analysis module so as to realize deep analysis on the intestinal flora sequencing data; and the analysis report generation module generates an analysis report based on an analysis result obtained by analysis of the data analysis module and / or mining data obtained by the data mining module, and performs visual display. According to the embodiment of the invention, efficient processing and deep analysis of the intestinal flora sequencing data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of bioinformatics data analysis, and particularly to a data analysis system, method, electronic device and storage medium for gut microbiota. Background Art

[0002] Bioinformatics analysis of gut microbiota is a research field that uses the theories, methods and technologies of bioinformatics to process, analyze and interpret gut microbiota-related data, so as to reveal the structure, function of gut microbiota and its relationship with host health and diseases.

[0003] With the in-depth research on gut microbiota, a large amount of gut microbiota sequencing data needs to be analyzed and processed efficiently and accurately. However, in the process of implementing the present disclosure, it is found that there are at least the following technical problems in the prior art: slow data upload speed, complex analysis process and requiring a large amount of manual intervention, low processing efficiency and lack of in-depth analysis, etc., which limit the efficiency and value of gut microbiota research. Summary of the Invention

[0004] The present disclosure provides a data analysis system, method, electronic device and storage medium for gut microbiota, so as to improve the data upload efficiency, automatically analyze and process the gut microbiota sequencing data in a full process, and realize the improvement of processing efficiency and the in-depth value mining of the analysis data of gut microbiota sequencing data.

[0005] According to one aspect of the present disclosure, there is provided a data analysis system for gut microbiota, including: a data upload module, a process setting module, a data analysis module, a data mining module and an analysis report generation module; wherein,

[0006] The data upload module obtains a first data file to be uploaded, and the first data file includes first gut microbiota sequencing data of multiple samples; performs block processing and parallel upload on the first gut microbiota sequencing data of the multiple samples;

[0007] The process setting module obtains at least one analysis index of each first gut microbiota sequencing data in the first data file, and generates a processing process corresponding to the first gut microbiota sequencing data based on the dependency relationship between the at least one analysis index and the default processing items corresponding to the at least one analysis index. The processing process includes processing nodes corresponding to the default processing items and processing nodes corresponding to each analysis index. Among them, the processing nodes with dependency relationships are connected in series, and the processing nodes without dependency relationships are set in parallel; the default processing items include at least one of data preprocessing, data quality control and feature extraction;

[0008] The data analysis module obtains the available resource amount, determines the batch processing quantity based on the available resource amount and the resource consumption of each first intestinal flora sequencing data for executing the processing flow, creates threads with the batch processing quantity, and performs batch processing on the first intestinal flora sequencing data in the first data file through the threads with the batch processing quantity until the processing of multiple first intestinal flora sequencing data in the first data file is completed, obtaining the index values of at least one analysis index corresponding to each first intestinal flora sequencing data, wherein each thread executes the processing flow on one first intestinal flora sequencing data;

[0009] The data mining module performs data mining based on the sample tags of the first intestinal flora sequencing data and the index values of at least one analysis index corresponding to the first intestinal flora sequencing data, obtaining mined data;

[0010] The analysis report generation module generates an analysis report based on the index values of at least one analysis index corresponding to each first intestinal flora sequencing data and / or the mined data, and displays the analysis report; the analysis report includes a visualization chart formed by the at least one analysis index and / or the mined data.

[0011] Optionally, the analysis indexes of the first intestinal flora sequencing data include at least one of the following: microbial health index, microbial colonization resistance, diversity, microbial species, enterotype, FB ratio, intestinal immunity evaluation index, nutrient metabolism ability evaluation index, short-chain fatty acid synthesis ability, nutrient synthesis ability, toxin degradation ability evaluation index, antibiotic resistance evaluation index, allergy risk evaluation index, and disease risk evaluation index, wherein the nutrient synthesis ability includes at least one of vitamin synthesis ability, natural pigment synthesis ability, amino acid synthesis ability, glutathione synthesis ability, and bile acid synthesis ability.

[0012] Optionally, the data uploading module is specifically configured to: identify the repeated sequencing fragments in the first intestinal flora sequencing data of the multiple samples, replace the repeated sequencing fragments with a set identifier to obtain the second intestinal flora sequencing data, forming a second data file, where the second data file includes the first intestinal flora sequencing data and / or the second intestinal flora sequencing data; perform block processing on the data in the second data file to obtain multiple data blocks; perform parallel uploading on the multiple data blocks, and after the uploaded multiple data blocks form a second data file, restore the set identifier in the data of the second data file to the repeated sequencing fragment to form the first data file.

[0013] Optionally, the process setting module is further configured to: obtain a pre-configured processing process; and / or, store the generated processing process; determine a processing process adapted to the at least one analysis metric from the pre-configured processing process or the stored processing process.

[0014] Optionally, the data analysis module stores an analysis algorithm for each of the analysis metrics and a processing algorithm for each of the default processing items;

[0015] The data analysis module is further configured to sequentially call the processing algorithms of the default processing items and the analysis algorithms of the analysis metrics according to the processing nodes in the processing process to obtain the metric values of at least one analysis metric;

[0016] and / or, the data analysis module is further configured to determine the status evaluation data of the intestinal flora of the sample corresponding to the first intestinal flora sequencing data based on the metric values of at least one analysis metric corresponding to the first intestinal flora sequencing data, wherein the status evaluation data is obtained based on a fusion algorithm or a determination rule of the metric values of at least one analysis metric corresponding to the first intestinal flora sequencing data.

[0017] Optionally, the data mining module is configured to: perform at least one level of data grouping on the first intestinal flora sequencing data of the multiple samples based on the sample labels corresponding to the first intestinal flora sequencing data, and perform data mining based on the metric values of at least one analysis metric in the first intestinal flora sequencing data in at least one level of grouping and the grouping labels of each level to obtain mined data; wherein the mined data includes the contribution degree data of each level of sample labels to the at least one analysis metric, the differences and correlations between at least one analysis metric corresponding to different groupings in the same level and / or different levels of groupings.

[0018] Optionally, the data mining module executes at least one of the following data mining methods:

[0019] Input the metric values of at least one analysis metric in the first intestinal flora sequencing data in at least one level of grouping and the grouping labels of each level into a data mining model to obtain the mined data;

[0020] For any level of grouping, determine the difference data between at least one analysis metric in the multiple groupings of the level, and determine the contribution degree data of the grouping label of the level to the at least one analysis metric based on the difference data;

[0021] For different levels of grouping, determine the difference data between at least one analysis metric in each level of grouping, compare the difference data between at least one analysis metric in different levels of grouping, and determine the contribution degree data of the grouping labels of different levels to the at least one analysis metric;

[0022] Based on the metric values of at least one analysis metric corresponding to different groups at the same level and / or different-level groups, generate a comparison data table and / or chart for the at least one analysis metric, and display the differences and correlations between the at least one analysis metric corresponding to different groups at the same level and / or different-level groups through the comparison data table and / or chart.

[0023] According to another aspect of the present disclosure, a data analysis method for gut microbiota is provided, including:

[0024] Display an interactive page of a gut microbiota data analysis system, where the interactive page includes an upload control, a process setting control, a data analysis control, a data grouping control, and a data mining control;

[0025] In response to a trigger operation on the upload control, obtain a first data file to be uploaded, where the first data file includes first gut microbiota sequencing data of multiple samples; perform chunking processing and parallel upload on the first gut microbiota sequencing data of the multiple samples, and display the completed-upload first gut microbiota sequencing data on the interactive page;

[0026] In response to a trigger operation on the process setting control, determine at least one analysis metric corresponding to the first gut microbiota sequencing data, and generate a processing flow corresponding to the first gut microbiota sequencing data based on the dependency relationship between the at least one analysis metric and the default processing items corresponding to the at least one analysis metric. The processing flow includes processing nodes corresponding to the default processing items and processing nodes corresponding to each analysis metric, where the processing nodes with a dependency relationship are serially connected, and the processing nodes without a dependency relationship are set in parallel; the default processing items include at least one of data preprocessing, data quality control, and feature extraction;

[0027] In response to a trigger operation on the data analysis control, obtain the available resource amount, determine the batch processing quantity based on the available resource amount and the resource consumption of each first gut microbiota sequencing data for executing the processing flow, create threads with the batch processing quantity, and perform batch processing on the first gut microbiota sequencing data in the first data file through the threads with the batch processing quantity until the processing of the multiple first gut microbiota sequencing data in the first data file is completed, obtaining metric values of at least one analysis metric corresponding to each first gut microbiota sequencing data, where each thread executes the processing flow on one first gut microbiota sequencing data;

[0028] In response to a trigger operation on the data grouping control, determine at least one-level data grouping for the first gut microbiota sequencing data of the multiple samples;

[0029] In response to a triggering operation on the data mining control, data mining is performed based on the index values of at least one analysis index of the first gut microbiota sequencing data in each hierarchical grouping and the grouping labels of each level to obtain mined data;

[0030] An analysis report is presented, the analysis report is generated based on the index values of at least one analysis index corresponding to each of the first gut microbiota sequencing data and / or the mined data, and the analysis report includes a visualization chart formed by the at least one analysis index and / or the mined data.

[0031] Optionally, the process setting control includes an index selection control and a process generation control;

[0032] The method further includes: in response to a triggering operation on the index selection control, presenting optional index items; in response to a selection operation on the optional index items, determining at least one analysis index corresponding to the first gut microbiota sequencing data; in response to a triggering operation on the process generation control, generating a processing process corresponding to the first gut microbiota sequencing data;

[0033] Alternatively, the process setting control further includes a process selection control;

[0034] The method further includes:

[0035] In response to a triggering operation on the process selection control, presenting a pre-configured processing process and / or the generated processing process; in response to a selection operation on the pre-configured processing process or the stored processing process, determining a processing process adapted to the at least one analysis index.

[0036] Optionally, the method further includes: presenting the processing process and each processing node in the processing process on the interaction page; and, during the execution of the processing process, rendering the executed processing nodes through a first rendering attribute and rendering the unexecuted processing nodes through a second rendering attribute, the first rendering attribute being different from the second rendering attribute.

[0037] Optionally, the method further includes: during the batch processing of the first gut microbiota sequencing data of multiple samples, presenting a status identifier and / or a processing progress of each of the first gut microbiota sequencing data, the status identifier including at least one of a completed status identifier, a processing status identifier, and an unprocessed status identifier.

[0038] Optionally, the first gut microbiota sequencing data includes a plurality of sample labels;

[0039] The method further includes: in response to a triggering operation on the data grouping control, displaying a grouping page, where the grouping page includes a grouping editing area and a label display area, and the grouping editing area includes a hierarchical division symbol; the label display area displays the multiple sample labels; in response to a triggering operation on the hierarchical division symbol, setting a new level; in response to a dragging operation on the sample label, setting the sample label corresponding to the new level; in response to an attribute setting operation on the sample label, setting the grouping attribute of the sample label in the new level to form the data grouping of at least one level.

[0040] According to another aspect of the present disclosure, there is provided an electronic device, which includes:

[0041] at least one processor; and

[0042] a memory communicatively connected to the at least one processor; wherein,

[0043] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data analysis method of the intestinal flora according to any embodiment of the present disclosure.

[0044] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the data analysis method of the intestinal flora according to any embodiment of the present disclosure when executed.

[0045] The technical solution of the embodiment of the present disclosure improves the upload efficiency of data files through the data upload module, automatically generates an automated processing flow for intestinal flora sequencing data through the process setting module, and reduces labor consumption. Through the data analysis module, batch processing of multiple intestinal flora sequencing data is realized, and the analysis efficiency of the intestinal flora sequencing data is improved. Through the data mining module, deep learning is performed on the analysis results respectively corresponding to multiple intestinal flora sequencing data to obtain mined data, which improves the learning depth of the intestinal flora sequencing data and the value of data analysis. Through the analysis report generation module, an analysis report including visual charts is generated and visually displayed, improving the display effect of the analysis results.

[0046] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0048] Figure 1 This is a schematic diagram of the structure of a data analysis system for intestinal flora provided by an embodiment of the present disclosure;

[0049] Figure 2 This is a flow chart of a method for analyzing intestinal flora data provided by an embodiment of the present disclosure;

[0050] Figure 3 Schematic diagram of an interactive page of the intestinal flora data analysis system provided by an embodiment of the present disclosure;

[0051] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0054] Figure 1 The following is a schematic diagram of the structure of a data analysis system for intestinal flora provided by an embodiment of the present disclosure. The data analysis system for intestinal flora can be integrated into electronic devices such as computers, servers, or terminal devices, where terminal devices include but are not limited to mobile phones, tablet computers, and PCs.

[0055] The intestinal flora data analysis system includes a data upload module 110, a process setting module 120, a data analysis module 130, a data mining module 140, and an analysis report generation module 150. The data upload module 110, the process setting module 120, the data analysis module 130, the data mining module 140, and the analysis report generation module 150 are connected in sequence.

[0056] Among them, the data upload module 110 is used to efficiently upload the intestinal flora sequencing data to be processed and improve the upload rate; the process setting module 120 provides an automated full process for the processing of intestinal flora sequencing data without manual intervention; the data analysis module 130 performs automated analysis on the intestinal flora sequencing data uploaded by the data upload module 110 based on the processing process formed by the process setting module 120; the data mining module 140 is used to further perform data mining on the analysis results obtained by the data analysis module 130 to achieve in-depth analysis of the intestinal flora sequencing data; the analysis report generation module 150 generates an analysis report based on the analysis results obtained by the data analysis module 130 and / or the mining data obtained by the data mining module 140, and performs a visual display.

[0057] The following is an introduction to the various modules in the intestinal flora data analysis system:

[0058] The data uploading module 110 obtains a first data file to be uploaded, wherein the first data file includes first intestinal flora sequencing data of multiple samples; and performs block processing and parallel uploading on the first intestinal flora sequencing data of the multiple samples.

[0059] The first data file includes first intestinal flora sequencing data to be processed. The first data file may be imported from an external storage device, read from a database, or transmitted from a device that generates intestinal flora sequencing data, including but not limited to a sequencer, or may be transmitted from other electronic devices with data storage capabilities. The method for obtaining the first data file is not limited herein. The first data file may be in the form of a file package or a data set, and the data format of the first data file is not limited herein.

[0060] The first gut microbiota sequencing data of multiple samples in the first data file. Optionally, the samples here can be subject samples, that is, by collecting samples from different subjects, different subject samples are obtained, and each subject sample is sequenced to obtain the first gut microbiota sequencing data of different subject samples. The subjects here can be human subjects or animal subjects, etc. Optionally, the samples here can be time samples, that is, by collecting samples from the same subject at different times, different time samples are obtained, and each time sample is sequenced to obtain the first gut microbiota sequencing data of different time samples.

[0061] In some embodiments of the present disclosure, the first data file is directly block-processed to obtain multiple data blocks. Specifically, it can be block-processing the first data file based on the block size; or, block-processing the first data file based on the number n of data blocks to obtain n data blocks. By uploading the multiple data blocks in parallel, the upload duration of the first data file is shortened and the upload efficiency is improved.

[0062] In some embodiments of the present disclosure, the data upload module 110 is specifically configured to: identify the repeated sequencing fragments in the first gut microbiota sequencing data of the multiple samples, replace the repeated sequencing fragments with a set identifier to obtain the second gut microbiota sequencing data, and form a second data file. The second data file includes the first gut microbiota sequencing data and / or the second gut microbiota sequencing data; block-process the data in the second data file to obtain multiple data blocks; upload the multiple data blocks in parallel, and after the uploaded multiple data blocks form the second data file, restore the set identifier in the data of the second data file to the repeated sequencing fragment to form the first data file.

[0063] Replace the repeated sequencing fragments in the first gut microbiota sequencing data of multiple samples in the first data file with a set identifier to obtain a second data file. The data volume of the second data file is smaller than that of the first data file, and the upload efficiency is improved by reducing the upload data volume.

[0064] Optionally, traverse each sequencing fragment in the multiple first gut microbiota sequencing data, compare the sequencing fragment to be identified with the sequenced fragments that have been traversed in the multiple first gut microbiota sequencing data. If the sequencing fragment to be identified is the same as the sequenced fragment that has been traversed, determine that the sequencing fragment to be identified is a repeated sequencing fragment. Replace the repeated sequencing fragment with a set identifier, and carry the set identifier at the position of the sequenced fragment that is the same as the repeated sequencing fragment to establish the corresponding relationship between the set identifier and the repeated sequencing fragment for subsequent restoration processing.

[0065] Optionally, obtain a pre-set sequencing fragment set, which includes historical sequencing fragments in the historical data file and whose occurrence frequency meets the set frequency threshold. Traverse each to-be-identified sequencing fragment in multiple first gut microbiota sequencing data and compare it with each historical sequencing fragment in the sequencing fragment set. If a to-be-identified sequencing fragment is identical to any historical sequencing fragment in the sequencing fragment set, determine that the to-be-identified sequencing fragment is a duplicate sequencing fragment. It can be understood that when the first data file is uploaded, the sequencing fragment set is updated based on the first data file. Each historical sequencing fragment in the sequencing fragment set corresponds to a set identifier respectively.

[0066] Replace the duplicate sequencing fragment with a set identifier, where the data volume of the set identifier is smaller than that of the duplicate sequencing fragment, achieving the effect of reducing the uploaded data volume. Different set identifiers correspond to different duplicate sequencing fragments, and each duplicate sequencing fragment corresponds to a unique set identifier to avoid data confusion.

[0067] It can be understood that if the first gut microbiota sequencing data includes duplicate sequencing fragments, replace the duplicate sequencing fragments with set identifiers to obtain the second gut microbiota sequencing data; if the first gut microbiota sequencing data does not include duplicate sequencing fragments, keep the first gut microbiota sequencing data unchanged. The first gut microbiota sequencing data without set identifiers and the second gut microbiota sequencing data with set identifiers form the second data file. Among them, the first gut microbiota sequencing data and / or the second gut microbiota sequencing data may include sequencing fragments carrying set identifiers.

[0068] Optionally, determine whether the data volume of the second data file is greater than the data volume threshold. If so, perform block processing on the second data file and upload the obtained multiple data blocks in parallel. If not, upload the second data file as a whole.

[0069] After the second data file is uploaded, restore the set identifiers in the data of the second data file to the duplicate sequencing fragments to form the first data file, ensuring the accuracy of the first data file. Optionally, it can be based on the set identifier corresponding to each historical sequencing fragment in the sequencing fragment set to restore the set identifier in the second data file. Or, for the set identifier carried in the sequencing fragment in the second data file, replace the set identifier in the second gut microbiota sequencing data with the corresponding sequencing fragment and delete the set identifier carried in the second data file.

[0070] Optionally, the data upload module 110 supports the first data file in sequencing data formats such as FASTQ, FASTA, and BAM, and can identify the file format of the first data file through file header information and data characteristics. When the upload of the first data file is completed, integrity verification is performed on the uploaded first data file to ensure the integrity of the uploaded data. Among them, the integrity verification can be implemented through the MD5 algorithm.

[0071] For the uploaded first intestinal flora sequencing data, it can be automatically stored in the dynamic database, which supports operations such as addition, deletion, modification, and query of the uploaded first intestinal flora sequencing data, and can achieve efficient storage and rapid retrieval of multiple first intestinal flora sequencing data. For example, the dynamic database can be a MySQL relational database. Regularly back up and optimize the dynamic database to ensure data security and system performance.

[0072] The process setting module 120 obtains at least one analysis index of each first intestinal flora sequencing data in the first data file, and generates a processing process corresponding to the first intestinal flora sequencing data based on the dependency relationship between the at least one analysis index and the default processing items corresponding to the at least one analysis index. The processing process includes processing nodes corresponding to the default processing items and processing nodes corresponding to each analysis index. Among them, the processing nodes with a dependency relationship are connected in series, and the processing nodes without a dependency relationship are set in parallel.

[0073] The analysis index of the first intestinal flora sequencing data is used to characterize the state of the first intestinal flora sequencing data in different analysis dimensions. Optionally, the analysis index of the first intestinal flora sequencing data includes at least one of the following: microbial health index, microbial colonization resistance, diversity, microbial species, enterotype, FB ratio, intestinal immunity assessment index, nutrient metabolism ability assessment index, short-chain fatty acid synthesis ability, nutrient synthesis ability, toxin degradation ability assessment index, antibiotic resistance assessment index, allergy risk assessment index, and disease risk assessment index.

[0074] Among them, the microbial health index can characterize whether the intestinal flora is in a balanced state and whether there are potential risks. A high microbial health index usually indicates a rich and diverse intestinal flora with normal functions, which has a positive effect on maintaining intestinal and overall health; a low index may imply dysbiosis, such as a decrease in beneficial bacteria and an increase in harmful bacteria.

[0075] Microbial colonization resistance reflects the ability of the intestinal flora to resist the invasion of foreign harmful microorganisms and colonize in the intestine. Strong colonization resistance means that the intestinal flora can prevent pathogenic bacteria from taking root in the intestine by competing for nutrients, producing antibacterial substances, etc., thereby protecting the intestine from infection and maintaining the stability of the intestinal microecosystem.

[0076] Diversity includes species diversity, gene diversity, etc. High species diversity indicates the presence of various different types of microorganisms in the gut. Different types of microorganisms can cooperate with each other to jointly perform various physiological functions, such as helping with digestion, synthesizing vitamins, etc., which helps to maintain the stability and balance of the gut ecosystem. Abundant gene diversity means that the microbial community has a wider range of functional potential and can better cope with environmental changes and the different needs of the host.

[0077] The microbial species identify the specific microbial species present in the gut, such as the types and relative proportions of beneficial bacteria like Bifidobacterium and Lactobacillus, as well as conditional pathogenic or harmful bacteria like Escherichia coli and Clostridium. Understanding the composition of microbial species helps to determine whether the structure of the gut microbiota is reasonable, and whether there is abnormal proliferation or absence of certain specific microorganisms, which is closely related to gut health and disease status.

[0078] According to the composition and structural characteristics of the gut microbiota, the gut microbiota is divided into different types. Different enterotypes may be related to the host's diet, genetic factors, and health status, and can reflect the unique characteristics of an individual's gut microbiota, which helps to understand the interaction between the gut microbiota and the host in a personalized way, as well as the susceptibility to different diseases.

[0079] The FB ratio usually refers to the ratio of Bacteroides to Firmicutes. These two types of bacteria are the main phyla in the gut microbiota, and changes in their ratio can reflect the overall structural changes of the gut microbiota. For example, in obese people, the FB ratio may change, with a relative increase in Firmicutes and a relative decrease in Bacteroides, suggesting that this ratio may be associated with physiological processes such as energy metabolism and body weight regulation.

[0080] Intestinal immunity assessment indicators are used to measure the impact of the gut microbiota on the intestinal immune system. The gut microbiota can interact with immune cells to promote the development and maturation of immune cells and regulate the intensity and direction of the immune response. Normal intestinal immunity assessment indicators indicate that the gut microbiota can maintain the balance of the intestinal immune system, enhance the gut's resistance to pathogens, and prevent the occurrence of infections and inflammatory diseases; abnormal indicators may suggest defects or overactivation of intestinal immune function, increasing the risk of disease. Nutrient metabolism ability assessment indicators reflect the metabolic ability of the gut microbiota to various nutrients, including carbohydrates, proteins, fats, vitamins, etc.

[0081] Gut microbiota participates in the digestion and absorption of nutrients, for example, by breaking down dietary fiber to produce short-chain fatty acids and synthesizing B vitamins and vitamin K. Nutrient metabolism capacity assessment indicators can indicate whether the gut microbiota is functioning properly in nutrient metabolism and whether it can provide the host with necessary nutritional support, and are closely related to the host's nutritional status and health.

[0082] Short-chain fatty acid synthesis capacity reflects the ability of the intestinal flora to metabolize dietary fiber. Short-chain fatty acids, such as acetic acid, propionic acid, and butyric acid, are important products of the fermentation of carbohydrates such as dietary fiber by the intestinal flora. They play an important role in maintaining intestinal mucosal barrier function, regulating intestinal immunity, and providing energy. Strong short-chain fatty acid synthesis capacity indicates that the intestinal flora can effectively utilize dietary fiber to produce beneficial metabolites, which are beneficial to intestinal health. Decreased synthesis capacity may lead to impaired intestinal mucosal barrier function and increase the risk of intestinal inflammation.

[0083] Nutrient synthesis capacity includes evaluation indicators of nutrient metabolism capacity, absorption, utilization and conversion of specific nutrients. For example, the intestinal flora's ability to absorb and transport trace elements such as iron and zinc, as well as the synthesis and utilization of certain specific amino acids, is evaluated. Nutrient synthesis capacity can further understand the specific role of the intestinal flora in maintaining the host's nutritional balance, and whether there is a potential risk of nutrient deficiency or excess. Among them, the nutrient synthesis capacity includes at least one of the following: vitamin synthesis capacity, natural pigment synthesis capacity, amino acid synthesis capacity, glutathione synthesis capacity, and bile acid synthesis capacity.

[0084] The toxin degradation capacity assessment index characterizes the ability of the intestinal flora to degrade toxins and harmful substances within the intestine. Intestinal flora can participate in the degradation or transformation of toxins and harmful substances within the intestine, such as endotoxins and drug metabolites. Intestinal flora with strong toxin degradation capacity can reduce the damage these harmful substances cause to the intestine and overall health, protect intestinal mucosal cells, and reduce the risk of inflammation and disease. Conversely, insufficient toxin degradation capacity may lead to the accumulation of toxins in the body, triggering a series of health issues.

[0085] Antibiotic resistance assessment indicators characterize the resistance of intestinal flora to antibiotics, help predict possible resistance problems when using antibiotics to treat diseases, and evaluate the ecological changes in intestinal flora caused by antibiotic use, providing a basis for the rational use of antibiotics and prevention of the spread of resistant bacteria.

[0086] Allergy risk assessment indicators are used to characterize the potential risk of allergic reactions in sampled subjects. The interaction between the intestinal flora and the immune system plays a key role in the development and progression of allergic reactions. Certain specific microbial community structures and metabolites may influence the immune system's recognition and response to allergens.

[0087] Disease risk assessment indicators characterize the risk levels of sampling subjects for different diseases, such as, including but not limited to, inflammatory bowel disease, obesity, diabetes, and cardiovascular diseases. The types and abundance changes of microorganisms in the gut microbiota of different disease domains, as well as the levels of metabolites, are related. Through the disease risk assessment indicators, the disease risk information of the sampling subject can be prompted.

[0088] In this embodiment, the above-mentioned indicators can be analyzed for the first gut microbiota sequencing data, improving the diversity and comprehensiveness of the analysis indicators to meet the analysis requirements of different data files.

[0089] There may be different analysis requirements for different first data files. That is to say, different first data files may correspond to different analysis indicators. At least one analysis indicator corresponding to the first data file can be carried in the configuration file of the first data file, or input by the operator, which is not limited here.

[0090] The analysis indicators corresponding to the first data file are different, and the processing processes performed on each first gut microbiota sequencing data in the first data file are different. It can be stated that the processing processes of different first gut microbiota sequencing data in the first data file are independent of each other and do not interfere with each other. The processing process is automatically generated according to at least one analysis indicator corresponding to the first data file, reducing manual intervention in the analysis process of the first data file, reducing labor consumption, and realizing the full-process automatic processing of the first gut microbiota sequencing data.

[0091] The default processing items corresponding to the analysis indicators can be understood as the data processing items that need to be performed during the analysis of the analysis indicators for the first gut microbiota sequencing data. The default processing items corresponding to different analysis indicators can be the same or different. The duplicate removal process is performed on the default processing items corresponding to different analysis indicators respectively to obtain the union of the default processing items. For example, when there is an overlap in the default processing items corresponding to two analysis indicators, the default processing item can be executed once for the first gut microbiota sequencing data, improving the reusability of the processed data and avoiding the repeated processing process of the first gut microbiota sequencing data.

[0092] Optionally, the default processing items include at least one of data preprocessing, data quality control, and feature extraction; among them, data preprocessing includes, but is not limited to, removing adapter sequences and contaminant sequences, sequence splicing, and data standardization. Data quality control includes quality assessment of the first gut microbiota sequencing data and removal of low-quality data. Among them, quality assessment can include base quality assessment and overall data quality assessment. Feature extraction is used to extract the feature information of the first gut microbiota sequencing data.

[0093] According to the dependency relationships between at least one analysis metric and at least one default processing item, a processing flow is established. For example, if analysis metric a depends on analysis metric b, then the processing process of analysis metric b is located before the processing process of analysis metric a. If both analysis metric a and analysis metric b depend on default processing item c, then the processing process of default processing item c is located before the processing process of analysis metric b. Specifically, processing nodes corresponding to each analysis metric and each default processing item are established, and based on the dependency relationships between at least one analysis metric and at least one default processing item, the connection relationships between the above-mentioned processing nodes are established. The connection relationships include series and parallel. For example, processing nodes with dependency relationships are set in series, and processing nodes without dependency relationships are set in parallel. The two processing nodes set in parallel can be processed in parallel to improve the processing efficiency of the first gut microbiota sequencing data.

[0094] In some embodiments of the present disclosure, the process setting module may store a preconfigured processing flow, or store the generated processing flow. The generated processing flow can be understood as a processing flow constructed according to the dependency relationships between at least one analysis metric and the default processing item corresponding to the analysis metric. Storing the newly constructed processing flow facilitates direct invocation when there are the same analysis requirements in the future, without the need to reconstruct the processing flow. Among them, the preconfigured processing flow can be set by the operator according to the high-frequency analysis requirements for the first gut microbiota sequencing data, which can reduce the construction process of the processing flow.

[0095] Correspondingly, the process setting module is further configured to: determine a processing flow adapted to the at least one analysis metric from the preconfigured processing flow and / or the stored processing flow. Selecting a processing flow from the preconfigured processing flow and / or the stored processing flow improves the reusability of the processing flow and the processing efficiency.

[0096] The data analysis module 130 obtains the available resource amount, determines the batch processing quantity based on the available resource amount and the resource consumption of each first gut microbiota sequencing data for executing the processing flow, creates threads with the batch processing quantity, and performs batch processing on the first gut microbiota sequencing data in the first data file through the threads with the batch processing quantity until the processing of multiple first gut microbiota sequencing data in the first data file is completed, and obtains the index values of at least one analysis metric corresponding to each first gut microbiota sequencing data, where each thread executes the processing flow on one first gut microbiota sequencing data.

[0097] Since there is a large amount of first gut microbiota sequencing data and the processing process is complex and time-consuming, the overall processing of the first data file takes a long time. In the embodiments of the present disclosure, the first gut microbiota sequencing data is processed in batches to improve the overall processing efficiency of the first data file and reduce the processing time. The batch processing quantity is determined according to the available resource amount of the data analysis system of the gut microbiota to ensure the normal execution of the batch processing. Specifically, the ratio of the available resource amount to the resource consumption of each first gut microbiota sequencing data for executing the processing flow is rounded down to obtain the batch processing quantity.

[0098] Optionally, the resource consumption of each first gut microbiota sequencing data for executing the processing flow can be determined according to the historical resource consumption of the processing flow, and the average value of the historical resource consumption of the processing flow is determined as the resource consumption of the processing flow. Optionally, the resource consumption of each first gut microbiota sequencing data for executing the processing flow can be determined according to the sum of the node resource consumptions of each processing node in the processing flow, and the node resource consumption of each processing node is determined according to the historical resource consumption of the node.

[0099] If the batch processing quantity is greater than or equal to the quantity of the first gut microbiota sequencing data in the first data file, then all the first gut microbiota sequencing data in the first data file are processed in batches. If the batch processing quantity is less than the quantity of the first gut microbiota sequencing data in the first data file, then a first batch of processing is performed on the local first gut microbiota sequencing data in the first data file based on the batch processing quantity. After the first batch of processing is completed, the next batch of processing is performed based on the batch processing quantity until all the first gut microbiota sequencing data in the first data file are processed.

[0100] During the batch processing, M threads are constructed according to the batch processing quantity M, and each thread performs the analysis and processing of a first gut microbiota sequencing data. That is to say, each thread executes the processing flow on a first gut microbiota sequencing data to obtain the index values of the first gut microbiota sequencing data for at least one analysis index.

[0101] Optionally, the data analysis module 130 stores the analysis algorithms of each analysis index and the processing algorithms of each default processing item; the analysis algorithms and processing algorithms here can respectively include at least one of machine learning algorithms and mathematical model algorithms. For example, the processing algorithm of the default processing item of feature extraction can be a machine learning algorithm. For example, the analysis algorithms corresponding to the allergy risk assessment index and the disease risk assessment index are respectively machine learning algorithms, and the analysis algorithms corresponding to different analysis indexes are different. The analysis algorithms of each analysis index and the processing algorithms of the default processing items are preset to provide an algorithm basis for the processing of the first gut microbiota sequencing data.

[0102] The data analysis module 130 is also used to call the processing algorithm of each default processing item and the analysis algorithm of the analysis index in sequence according to the processing node in the processing flow to obtain the index value of at least one analysis index. Each processing node in the processing flow can be set with a node identifier to characterize the processing type of the processing node. The data analysis module 130 calls the processing algorithm / analysis algorithm corresponding to the processing node according to the node identifier of the processing node, wherein the data analysis module 130 calls the processing algorithm / analysis algorithm corresponding to the first processing node in the processing flow to obtain the processing result of the processing node, and uses the processing result of the first processing node as the trigger event of the next processing node. Based on the node identifier of the next processing node, the processing algorithm / analysis algorithm corresponding to the next node is called, and so on, that is, the processing result of the previous processing node is received as the trigger event of the next processing node until the processing flow is completed, thereby realizing batch full-process automatic processing of multiple first intestinal flora sequencing data, especially for the complex processing process of multiple analysis indicators, and improving processing efficiency.

[0103] Optionally, in the case where the processing result of the previous node can trigger multiple parallel processing nodes, multiple sub-threads are constructed to execute the multiple triggered processing nodes in parallel to further improve the processing efficiency.

[0104] In some embodiments of the present disclosure, the data analysis module 130 is further configured to determine, based on the value of at least one analytical indicator corresponding to the first intestinal flora sequencing data, intestinal flora status assessment data for the sample corresponding to the first intestinal flora sequencing data. This status assessment data may represent the overall status of the intestinal flora, which may include a healthy state and an unhealthy state, wherein the unhealthy state may include at least one abnormal level.

[0105] Optionally, the status assessment data is obtained based on a fusion algorithm of the index value of at least one analysis index corresponding to the first intestinal flora sequencing data. Specifically, the data is obtained by weightedly fusing the index value of at least one analysis index using a fusion weight corresponding to each analysis index. The fusion weight corresponding to the analysis index may be pre-set or obtained using a particle swarm algorithm.

[0106] Exemplarily, the method for obtaining the fusion weights corresponding to the analysis indicators includes: obtaining the weight set corresponding to each analysis indicator, initializing each weight in the weight set as a particle in the solution space according to the predefined particle swarm optimization algorithm, which is used to perform population iterative calculations on multiple particles and track the optimal particle in the solution space; for the weight value corresponding to each particle, calculating the individual fitness value of each particle, and determining the global optimal fitness value according to the individual fitness values of multiple particles; updating the weight value of each particle to obtain the updated particle, checking whether the iteration meets the end condition, if not, continuing to execute the step of checking whether the iteration meets the end condition according to the weight value corresponding to each updated particle, if not, continuing to execute according to the target weight value corresponding to each updated particle.

[0107] For example, assume there are m samples, and each sample has n analysis indicators x ij , where i represents the sample number and j represents the indicator number. For the indicator value y of each analysis indicator i , let the weight vector of the kth particle in the particle swarm optimization algorithm be The above fitness function can be where is the comprehensive evaluation value,

[0108] The above fitness function is only an example. In other embodiments, the fitness function can also be a fitness function based on information entropy, etc., which is not limited herein. By using the particle swarm optimization algorithm to determine the fusion weights corresponding to different analysis indicators, the fusion accuracy is improved.

[0109] Store the fusion weights corresponding to the analysis indicator combinations formed by at least one analysis indicator for subsequent calling. Correspondingly, before fusing at least one analysis indicator, it can be determined whether there are determined fusion weights. If so, directly read them; if not, re-determine them based on the particle swarm optimization algorithm and store them.

[0110] Optionally, the state evaluation data is obtained based on the determination rules of the indicator values of at least one analysis indicator corresponding to the first gut microbiota sequencing data. The determination rules include the indicator ranges corresponding to different state evaluation data for the analysis indicators. Match the indicator values of at least one analysis indicator with the indicator ranges corresponding to different state evaluation data to determine the state evaluation data of the gut microbiota.

[0111] On this basis, the indicator range of the analysis indicator can include an abnormal indicator range. In the case where the indicator range corresponding to the state evaluation data includes an abnormal indicator range, the state evaluation data is an abnormal state, and it can be determined that the analysis indicator corresponding to the state evaluation data is an abnormal indicator, realizing the abnormal detection of the indicator.

[0112] The data mining module 140 performs data mining based on the sample tags of the first gut microbiota sequencing data and the index values of at least one analysis index corresponding to the first gut microbiota sequencing data to obtain mined data.

[0113] The data analysis module 130 processes each first gut microbiota sequencing data independently and cannot understand the associations and differences between different first gut microbiota sequencing data. In this embodiment, by performing data mining on the index values of the analysis indexes corresponding to different first gut microbiota sequencing data respectively, high-value mined data is obtained, the analysis depth of the gut microbiota sequencing data is realized, and the data analysis value of the gut microbiota sequencing data is improved.

[0114] The first gut microbiota sequencing data corresponds to multiple sample tags, and the sample tags represent the attribute information of the sampling object of the first gut microbiota sequencing data. For example, the sample tags include but are not limited to disease type, region, sampling time, sample type, etc. By grouping multiple first gut microbiota sequencing data through the above data tags, at least one level of data grouping is obtained. In this embodiment, according to the grouping operation of the operator on the first gut microbiota sequencing data, free grouping of multiple first gut microbiota sequencing data can be realized, the diversity and flexibility of data grouping are improved, and different grouping requirements of the operator for the gut microbiota sequencing data are met.

[0115] In some embodiments of the present disclosure, the data mining module 140 is configured to: perform at least one level of data grouping on the first gut microbiota sequencing data of the multiple samples based on the sample tags corresponding to each of the first gut microbiota sequencing data, and perform data mining based on the index values of at least one analysis index of the first gut microbiota sequencing data in at least one level of grouping and the grouping tags of each level to obtain mined data.

[0116] Group multiple first gut microbiota sequencing data according to at least one sample tag. For example, select the sample tags corresponding to different levels and the division attribute values of the sample tags, match each first gut microbiota sequencing data with each sample tag and division data value, and form at least one level of data grouping.

[0117] For the divided data grouping, each data grouping may include at least one first gut microbiota sequencing data. Data mining is performed based on the index values of at least one analysis index corresponding to the first gut microbiota sequencing data in the same level of grouping, and / or data mining is performed based on the index values of at least one analysis index corresponding to the first gut microbiota sequencing data in different levels of grouping, so as to realize the flexibility and diversity of data mining.

[0118] It is understandable that the data grouping operation can be performed before the data analysis module 130 analyzes and processes each first intestinal flora sequencing data, or after the data analysis module 130 analyzes and processes each first intestinal flora sequencing data, which is not limited here.

[0119] Optionally, the mining data includes the contribution data of sample labels at each level to the at least one analysis indicator, the differences and correlations between different groups at the same level and / or at least one analysis indicator corresponding to groups at different levels. Among them, the contribution data of sample labels to at least one analysis indicator can be understood as the degree of influence of different attribute values of sample labels on the change in the indicator value of the analysis indicator. The greater the contribution data of sample labels to the analysis indicator, the greater the degree of influence of different attribute values of sample labels on the change in the indicator value of the analysis indicator. Here, the contribution data of a sample label to the analysis indicator can be obtained by mining, or the contribution data of a combination of two or more sample labels to the analysis indicator.

[0120] The difference and correlation between at least one analysis indicator corresponding to different groups at the same level can characterize the difference data and correlation data between the analysis indicators corresponding to different attribute values of sample labels corresponding to the same level, and further determine the relationship between the changing trend of the attribute values of sample labels at this level and the changing trend of the analysis indicators. It can characterize the difference data and correlation data between the analysis indicators corresponding to the combination of sample labels corresponding to different levels, and further determine the relationship between the changing trend of the attribute values of the combination of sample labels corresponding to different levels and the changing trend of the analysis indicators.

[0121] Optionally, the data mining method performed by the data mining module 140 includes: inputting the index value of at least one analytical indicator of the first intestinal flora sequencing data in the at least one hierarchical grouping and the grouping labels of each hierarchical level into a data mining model to obtain the mined data. In this embodiment, by setting a data mining model, which can be a machine learning model, such as a neural network model, etc., it is possible to learn the relationship between the index value of at least one analytical indicator of the first intestinal flora sequencing data in the at least one hierarchical grouping and the grouping labels of each hierarchical level to obtain the mined data. The specific structure of the data mining model is not limited here.

[0122] Optionally, the data mining method performed by the data mining module 140 includes: for groups at any level, determining the difference data between at least one analysis indicator in multiple groups at the level, and determining the contribution data of the grouping label at the level to the at least one analysis indicator based on the difference data. Among them, the indicator value of at least one analysis indicator corresponding to the first intestinal flora sequencing data in each group of at least two groups at the same level is statistically calculated to obtain statistical data values, which include but are not limited to mean, variance, and standard deviation. The difference data of the analysis indicators in different groups are determined by the statistical data values of the analysis indicators corresponding to different groups at the same level. The contribution data is positively correlated with the above-mentioned difference data, wherein the larger the difference data is, the greater the contribution data of the grouping label at the level to the analysis indicator is.

[0123] Optionally, the data mining method performed by the data mining module 140 includes: for groups at different levels, determining the difference data between at least one analytical indicator in each level of grouping, comparing the difference data between at least one analytical indicator in the groups at different levels, and determining the contribution data of the sub-lease labels at different levels to the at least one analytical indicator. Similarly, for groups at different levels, performing statistical calculations on the indicator values of at least one analytical indicator corresponding to the first intestinal flora sequencing data in each group to obtain statistical data values, determining the difference data of the combination of analytical indicators in different groups based on the statistical data values at different levels, and determining the contribution data of the sample label combination to each analytical indicator based on the difference data of the combination of analytical indicators in different groups.

[0124] Optionally, the data mining method performed by the data mining module 140 includes: generating a comparison data table and / or chart of the at least one analysis indicator based on the indicator value of the at least one analysis indicator corresponding to different groups at the same level and / or groups at different levels, and displaying the differences and correlations between the at least one analysis indicator corresponding to different groups at the same level and / or groups at different levels through the comparison data table and / or chart.

[0125] For the indicator value of at least one analysis indicator of different groups at the same level, a comparison data table and / or chart is constructed to show the difference and correlation between at least one analysis indicator corresponding to different groups at the same level; different groups at the same level correspond to the same sample label and different partitioning attribute values, and the difference and correlation between at least one analysis indicator corresponding to different partitioning attribute values of the sample label are shown through the above-mentioned comparison data table and / or chart.

[0126] Comparison data tables and / or charts are constructed for the indicator values of at least one analysis indicator corresponding to different hierarchical groupings, demonstrating the differences and correlations between the at least one analysis indicator corresponding to the different hierarchical groupings. Different groups at different hierarchical levels correspond to different sample labels and different partitioning attribute values, forming sample label combinations and partitioning attribute value combinations. The comparison data tables and / or charts are used to demonstrate the differences and correlations between the at least one analysis indicator corresponding to different sample label combinations and partitioning attribute value combinations.

[0127] In some embodiments of the present disclosure, the data mining module 140 is also used to cluster the index values of at least one analysis index of multiple first intestinal flora sequencing data to obtain multiple clusters, and perform data mining based on the sample labels corresponding to the first intestinal flora sequencing data in each cluster to obtain the correlation between the analysis index and the sample label.

[0128] An analysis report generation module 150 generates an analysis report based on the indicator value of at least one analysis indicator corresponding to each of the first intestinal flora sequencing data and / or the mined data, and displays the analysis report; the analysis report includes a visualization chart formed by the at least one analysis indicator and / or the mined data.

[0129] For the plurality of first intestinal flora sequencing data in the first data file, a data table and a chart are generated based on the index value of at least one analytical indicator corresponding to each first intestinal flora sequencing data. The data table may include, but is not limited to, the index value of the at least one analytical indicator corresponding to each first intestinal flora sequencing data, and statistical data values corresponding to the at least one analytical indicator. The statistical data values may include, but are not limited to, the mean, variance, and standard deviation. The chart may include, but is not limited to, a distribution graph of the index value of each analytical indicator for the plurality of first intestinal flora sequencing data within the overall index range and a distribution graph within different index sub-ranges.

[0130] The mined data can be displayed in the form of data tables and charts to achieve an intuitive display of the mined data.

[0131] In some embodiments of the present disclosure, the analysis report may include detailed results of the above-mentioned analysis indicators and support export in multiple formats. The analysis report generation module 150 supports a visual display function to facilitate intuitive display of the analysis report.

[0132] The technical solution provided in this embodiment improves the upload efficiency of data files through the data upload module, automatically generates an automated processing flow for gut microbiota sequencing data through the process setting module, and reduces labor consumption. Through the data analysis module, batch processing of multiple gut microbiota sequencing data is achieved, improving the analysis efficiency of gut microbiota sequencing data. Through the data mining module, deep learning is performed on the analysis results corresponding to multiple gut microbiota sequencing data respectively to obtain mined data, improving the learning depth of gut microbiota sequencing data and the value of data analysis. Through the analysis report generation module, an analysis report including visual charts is generated and visually displayed, improving the display effect of the analysis results.

[0133] Figure 2 FIG. is a flowchart of a method for analyzing gut microbiota data provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of efficiently processing multiple gut microbiota sequencing data and mining high-value data from the analysis results of gut microbiota sequencing data. This method can be executed by a gut microbiota data analysis system, and the gut microbiota data analysis system can be implemented in the form of hardware and / or software. As Figure 2 shown, the method includes:

[0134] S210. Display an interaction page of the gut microbiota data analysis system, where the interaction page includes an upload control, a process setting control, a data analysis control, a data grouping control, and a data mining control.

[0135] Exemplarily, refer to Figure 3 , Figure 3 FIG. is a schematic diagram of an interaction page of the gut microbiota data analysis system provided by an embodiment of the present disclosure. It can be understood that Figure 3 is only an example, and the specific form of the interaction page and the layout of each control in the interaction page are not limited here.

[0136] In this embodiment, by providing an interaction page, the interaction between the operator and the gut microbiota data analysis system can be realized, and the interaction operation in the process of processing gut microbiota sequencing data is simplified.

[0137] S220. In response to a trigger operation on the upload control, obtain a first data file to be uploaded, perform block parallel upload on the first data file, and display the first gut microbiota sequencing data that has completed the upload in the interaction page.

[0138] Among them, the first data file includes the first intestinal flora sequencing data of multiple samples; the block-parallel uploading of the first data file may include: identifying the repeated sequencing fragments in the first intestinal flora sequencing data of the multiple samples, replacing the repeated sequencing fragments with a set identifier to obtain the second intestinal flora sequencing data, forming a second data file, where the second data file includes the first intestinal flora sequencing data and / or the second intestinal flora sequencing data; performing block processing on the data in the second data file to obtain multiple data blocks, performing parallel uploading on the multiple data blocks, and after the uploaded multiple data blocks form a second data file, restoring the set identifier in the data of the second data file to the repeated sequencing fragment to complete the uploading of the first data file.

[0139] In this embodiment, by setting an upload control in the interactive page, the data source of the first data file can be selected through the upload control, and one-click uploading of the first data file can be realized.

[0140] See Figure 3 , in Figure 3 there is a data display area, which is used to display the uploaded first intestinal flora sequencing data of each sample. Among them, the data display area includes multiple data rows, and each data row displays a first intestinal flora sequencing data. It can be understood that the first intestinal flora sequencing data can correspond to multiple sample tags, and the data display area includes multiple data columns. The first intestinal flora sequencing data and sample tags are displayed through different data columns. For example, the first data column displays the first intestinal flora sequencing data, the second data column displays sample tag 1, the third data column displays sample tag 2, and so on.

[0141] In response to operations of adding, deleting, modifying, and querying the first intestinal flora sequencing data already displayed in the data display area, perform operations of adding, deleting, modifying, and querying the first intestinal flora sequencing data already displayed.

[0142] S230. In response to a trigger operation on the process setting control, determine at least one analysis index corresponding to the first intestinal flora sequencing data, and generate a processing process corresponding to the first intestinal flora sequencing data based on the dependency relationship between the at least one analysis index and the default processing items corresponding to the at least one analysis index. The processing process includes processing nodes corresponding to the default processing items and processing nodes corresponding to each analysis index. Among them, the processing nodes with a dependency relationship are connected in series, and the processing nodes without a dependency relationship are set in parallel; the default processing items include at least one of data preprocessing, data quality control, and feature extraction.

[0143] In some embodiments of the present disclosure, the process setting control includes an index selection control and a process generation control; wherein, the index selection control is used to select analysis indexes corresponding to the first data file, and the process generation control is used to trigger the generation of a processing process.

[0144] Optionally, in response to a trigger operation of the index selection control, display optional index items; in response to a selection operation on the optional index items, determine at least one analysis index corresponding to the first gut microbiota sequencing data; in response to a trigger operation of the process generation control, generate a processing process corresponding to the first gut microbiota sequencing data.

[0145] The trigger operation on the index selection control can be a click operation on the index selection control. The optional index items can be displayed in the form of a drop-down menu, a newly added display area, or a newly added display page, which is not limited herein. By displaying the optional index items for the operator, it is convenient to select at least one analysis index through an interactive operation, simplifying the process of determining the analysis index.

[0146] The trigger operation on the process generation control can be a click operation on the process generation control. By the trigger operation of the index selection control, trigger the execution of the construction process of the processing process, that is, obtain the default processing items corresponding to the selected analysis indexes, or the dependency relationships between the default processing items and the analysis indexes, establish the processing nodes and the connection relationships between the processing nodes, and form a processing process.

[0147] In this embodiment, by generating a processing process through an interactive operation on the interactive page, there is no need to write a process script, simplifying the generation method of the processing process and reducing the difficulty of analyzing gut microbiota sequencing data.

[0148] In some embodiments of the present disclosure, the process setting control further includes a process selection control; the process selection control is used to select a processing process adapted to the at least one analysis index from the stored processing processes.

[0149] Optionally, in response to a trigger operation on the process selection control, display the pre-configured processing process and / or the generated processing process; in response to a selection operation on the pre-configured processing process or the stored processing process, determine a processing process adapted to the at least one analysis index.

[0150] The pre-configured processing process and / or the generated processing process can be displayed in the form of a drop-down menu, a newly added display area, or a newly added display page, which is not limited herein. In response to a selection operation among the displayed multiple processing processes, determine a processing process adapted to the at least one analysis index for analyzing and processing each gut microbiota sequencing data in the first data file.

[0151] S240. In response to a trigger operation on the data analysis control, obtain the available resource amount, determine the batch processing quantity based on the available resource amount and the resource consumption of each first intestinal flora sequencing data for executing the processing flow, create threads of the batch processing quantity, and perform batch processing on the first intestinal flora sequencing data in the first data file through the threads of the batch processing quantity until the processing of multiple first intestinal flora sequencing data in the first data file is completed. Each thread executes the processing flow on one first intestinal flora sequencing data to obtain the index values of at least one analysis index corresponding to each first intestinal flora sequencing data.

[0152] Optionally, the batch processing quantity is displayed on the interaction page, and batch processing is performed on multiple first intestinal flora sequencing data displayed in the data display area according to the batch processing quantity. Exemplarily, if the batch processing quantity is M, the first M first intestinal flora sequencing data can be determined for analysis and processing according to the display order of the first intestinal flora sequencing data in the data display area, and after the first batch is completed, the second M first intestinal flora sequencing data are determined for analysis and processing until the analysis and processing of all first intestinal flora sequencing data are completed.

[0153] In some embodiments of the present disclosure, in order to display the data processing progress, during the batch processing of the first intestinal flora sequencing data of multiple samples, the status identifier and / or processing progress of each first intestinal flora sequencing data are displayed. The status identifier includes at least one of a completed status identifier, a processing status identifier, and an unprocessed status identifier. When the first intestinal flora sequencing data is not batch processed, the status identifier of the first intestinal flora sequencing data is the unprocessed status identifier; when the first intestinal flora sequencing data is batch processed, the status identifier of the first intestinal flora sequencing data is the processing status identifier; when the batch processing of the first intestinal flora sequencing data is completed, the status identifier of the first intestinal flora sequencing data is the completed status identifier. The above status identifiers can be set in the data row where each first intestinal flora sequencing data is located.

[0154] The processing progress can include the processing progress of each first intestinal flora sequencing data, and can also include the total processing progress of multiple first intestinal flora sequencing data in the first data file. The above processing progress can be a value between 0 and 1.

[0155] In some embodiments of the present disclosure, the processing flow and each processing node in the processing flow are displayed on the interaction page; and during the execution of the processing flow, the executed processing nodes are rendered through a first rendering attribute, and the unexecuted processing nodes are rendered through a second rendering attribute, and the first rendering attribute is different from the second rendering attribute.

[0156] Optionally, a process query control can be set on the interactive page. In response to the triggering operation of the process query control, the processing process and each processing node in the processing process are displayed on the interactive page, providing an intuitive display of the processing process.

[0157] Optionally, the above-mentioned processing process is respectively executed on the first gut microbiota sequencing data. During the processing of the first gut microbiota sequencing data, a process query is performed. On the basis of displaying each processing node in the processing process, unexecuted processing nodes and executed processing nodes are displayed through different rendering attributes, and the processing progress is displayed in the processing process. Among them, the rendering attribute can be a color attribute, that is, the colors of unexecuted processing nodes and executed processing nodes are different. For example, the rendering color of unexecuted processing nodes is gray, and the rendering color of executed processing nodes is green, improving the display effect of the processing progress.

[0158] In some embodiments of the present disclosure, an analysis result query control corresponding to each first gut microbiota sequencing data is set on the interactive page. The query control can be set in the display row where the first gut microbiota sequencing data is located. In response to the triggering operation on the analysis result query control, the index values of the analysis indexes of the first gut microbiota sequencing data corresponding to the data row where it is located are displayed.

[0159] S250. In response to the triggering operation on the data grouping control, determine at least one level of data grouping for the first gut microbiota sequencing data of the multiple samples.

[0160] Among them, the data grouping is determined based on the sample labels corresponding to each of the first gut microbiota sequencing data.

[0161] The first gut microbiota sequencing data includes multiple sample labels. In this embodiment, through the triggering operation on the data grouping control, free grouping of multiple first gut microbiota sequencing data is achieved through multiple sample labels.

[0162] In some embodiments of the present disclosure, in response to the triggering operation on the data grouping control, a grouping page is displayed. The grouping page includes a grouping editing area and a label display area. The grouping editing area includes a hierarchical division symbol; the label display area displays the multiple sample labels; in response to the triggering operation on the hierarchical division symbol, a new level is set; in response to the dragging operation on the sample label, the sample labels corresponding to the new level are set; in response to the attribute setting operation on the sample label, the grouping attributes of the sample labels in the new level are set to form the at least one level of data grouping.

[0163] Multiple sample labels are displayed through the label display area, providing optional labels for grouping. By triggering the hierarchical division symbol in the grouping editing area, a new level is set. At least one sample label is selected from the label display area as the grouping label for the new level. The grouping quantity of the new level and the grouping attributes corresponding to each group are determined through attribute setting operations. If only one new level is set, the data grouping of one level is determined. If two or more new levels are set, the data grouping of multiple levels is determined. The setting process of the data grouping for each level is the same and will not be elaborated here.

[0164] Match the label values of the sample labels of multiple first gut microbiota sequencing data with the sample labels corresponding to at least one data grouping respectively to determine the first gut microbiota sequencing data included in the data grouping of each level.

[0165] In some embodiments of the present disclosure, it can also be achieved through the tick operation on multiple first gut microbiota sequencing data in the data display area. This interactive page supports tick operations such as multi-selection and full selection, facilitating the quick determination of the first gut microbiota sequencing data corresponding to each group. It supports grouping according to multiple condition combinations, provides flexible condition setting options, and can customize combined conditions to meet different analysis requirements.

[0166] The grouped data can be saved as a new data set for convenient subsequent reuse. Each grouped data set can be named and described to facilitate management and identification.

[0167] S260. In response to the triggering operation on the data mining control, data mining is performed based on the index values of at least one analysis index of the first gut microbiota sequencing data in each level grouping and the grouping labels of each level to obtain mining data.

[0168] Among them, the mining data includes the contribution degree data of each level sample label to the at least one analysis index, and the differences and correlations between at least one analysis index corresponding to different groups in the same level and / or different level groups.

[0169] The triggering operation on the data mining control can be understood as a click operation on the data mining control. By setting the data mining control to trigger the automatic mining of the index values of at least one analysis index of each first gut microbiota sequencing data, one-key processing is realized, reducing the difficulty of data mining.

[0170] Optionally, an extension interface can be set in the interactive page. Through this extension interface, new data mining rules / algorithms can be written to realize the extension of data mining methods, improve the diversity of data mining, and avoid limitations.

[0171] S270. Display an analysis report, where the analysis report is generated based on the indicator value of at least one analysis indicator corresponding to each of the first intestinal flora sequencing data and / or the mined data, and the analysis report includes a visualization chart formed by the at least one analysis indicator and / or the mined data.

[0172] Optionally, the interactive page may include an analysis report query control, and in response to a query request to the analysis report query control, a report page is displayed, and the analysis report is displayed on the report page to achieve visual display of the analysis report.

[0173] The technical solution of this embodiment supports interaction between users and the intestinal flora data analysis system through the interactive page of the intestinal flora data analysis system, and reduces the difficulty of the data analysis process through interactive operations. Specifically, the intestinal flora data analysis system improves the efficiency of data file upload, automatically generates an automated processing flow for intestinal flora sequencing data, and reduces manpower consumption; realizes batch processing of multiple intestinal flora sequencing data, and improves the analysis efficiency of intestinal flora sequencing data; conducts deep learning on the analysis results corresponding to multiple intestinal flora sequencing data, obtains mining data, improves the learning depth of intestinal flora sequencing data, and improves the value of data analysis; generates analysis reports including visual charts, and performs visual display to improve the display effect of analysis results.

[0174] Figure 4 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0175] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the random access memory (RAM) 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the read-only memory (ROM) 12, and the random access memory (RAM) 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0176] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0177] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data analysis method of the gut microbiota.

[0178] In some embodiments, the data analysis method of the gut microbiota can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the read-only memory (ROM) 12 and / or the communication unit 19. When the computer program is loaded into the random access memory (RAM) 13 and executed by the processor 11, one or more steps of the data analysis method of the gut microbiota described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data analysis method of the gut microbiota by any other appropriate means (e.g., by means of firmware).

[0179] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0180] The computer programs for implementing the data analysis method of the gut microbiota of the present disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0181] The embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions for causing a processor to execute a data analysis method of the gut microbiota, the method including:

[0182] Displaying an interaction page of the data analysis system of the gut microbiota, the interaction page including an upload control, a process setting control, a data analysis control, a data grouping control, and a data mining control;

[0183] In response to a trigger operation on the upload control, obtain a first data file to be uploaded, where the first data file includes first gut microbiota sequencing data of multiple samples; identify repeated sequencing fragments in the first gut microbiota sequencing data of the multiple samples, replace the repeated sequencing fragments with a set identifier to obtain second gut microbiota sequencing data, and form a second data file, where the second data file includes the first gut microbiota sequencing data and / or the second gut microbiota sequencing data; perform chunking processing on the data in the second data file to obtain multiple data chunks, perform parallel upload on the multiple data chunks, and after the uploaded multiple data chunks form a second data file, restore the set identifier in the data of the second data file to the repeated sequencing fragment, and display the first gut microbiota sequencing data that has completed the upload on the interactive page;

[0184] In response to a trigger operation on the process setting control, determine at least one analysis index corresponding to the first gut microbiota sequencing data, and generate a processing process corresponding to the first gut microbiota sequencing data based on the dependency relationship between the at least one analysis index and the default processing items corresponding to the at least one analysis index, where the processing process includes processing nodes corresponding to the default processing items and processing nodes corresponding to each of the analysis indexes, and among them, the processing nodes with a dependency relationship are connected in series, and the processing nodes without a dependency relationship are set in parallel; the default processing items include at least one of data preprocessing, data quality control, and feature extraction;

[0185] In response to a trigger operation on the data analysis control, obtain the available resource amount, determine the batch processing quantity based on the available resource amount and the resource consumption of each first gut microbiota sequencing data for executing the processing process, create threads with the batch processing quantity, and perform batch processing on the first gut microbiota sequencing data in the first data file through the threads with the batch processing quantity until the processing of the multiple first gut microbiota sequencing data in the first data file is completed, and obtain the index values of at least one analysis index corresponding to each first gut microbiota sequencing data, where each thread executes the processing process on one first gut microbiota sequencing data;

[0186] In response to a trigger operation on the data grouping control, determine at least one - level data grouping for the first gut microbiota sequencing data of the multiple samples, and the data grouping is determined based on the sample labels corresponding to each first gut microbiota sequencing data;

[0187] In response to a triggering operation on the data mining control, data mining is performed based on the index values of at least one analysis index of the first gut microbiota sequencing data in each hierarchical grouping and the grouping labels of each level, and mining data is obtained. The mining data includes the contribution degree data of each hierarchical sample label to the at least one analysis index, and the differences and correlations between at least one analysis index corresponding to different groupings at the same level and / or different hierarchical groupings.

[0188] Display an analysis report, which is generated based on the index values of at least one analysis index corresponding to the first gut microbiota sequencing data and / or the mining data. The analysis report includes a visualization chart formed by the at least one analysis index and / or the mined data.

[0189] In the context of the present disclosure, a computer-readable storage medium may be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0190] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the electronic device. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0191] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0192] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0193] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitation is imposed herein.

[0194] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data analysis system for intestinal flora, characterized in that: Including: A data upload module, a process setting module, a data analysis module, a data mining module, and an analysis report generation module; among them, The data upload module obtains a first data file to be uploaded, and the first data file includes first intestinal flora sequencing data of multiple samples; performs chunking processing and parallel upload on the first intestinal flora sequencing data of the multiple samples; The process setting module obtains at least one analysis index of each first intestinal flora sequencing data in the first data file, and generates a processing process corresponding to the first intestinal flora sequencing data based on the dependency relationship between the at least one analysis index and the default processing items corresponding to the at least one analysis index. The processing process includes processing nodes corresponding to the default processing items and processing nodes corresponding to each analysis index. Among them, the processing nodes with a dependency relationship are connected in series, and the processing nodes without a dependency relationship are set in parallel; the default processing items include at least one of data preprocessing, data quality control, and feature extraction; The data analysis module obtains the available resource amount, determines the batch processing quantity based on the available resource amount and the resource consumption of each first intestinal flora sequencing data for executing the processing process, creates threads with the batch processing quantity, and performs batch processing on the first intestinal flora sequencing data in the first data file through the threads with the batch processing quantity until the processing of the multiple first intestinal flora sequencing data in the first data file is completed, and obtains the index values of at least one analysis index corresponding to each first intestinal flora sequencing data. Among them, each thread executes the processing process on one first intestinal flora sequencing data; The data mining module performs data mining based on the sample labels of the first intestinal flora sequencing data and the index values of at least one analysis index corresponding to the first intestinal flora sequencing data, and obtains mining data; The analysis report generation module generates an analysis report based on the index values of at least one analysis index corresponding to each first intestinal flora sequencing data and / or the mining data, and displays the analysis report; the analysis report includes a visualization chart formed by the at least one analysis index and / or the mined data.

2. The data analysis system for gut microbiota according to claim 1, wherein The analysis indexes for the first intestinal flora sequencing data include at least one of the following: microbial health index, microbial colonization resistance, diversity, microbial species, enterotype, FB ratio, intestinal immunity evaluation index, nutrient metabolism ability evaluation index, short-chain fatty acid synthesis ability, nutrient synthesis ability, toxin degradation ability evaluation index, antibiotic resistance evaluation index, allergy risk evaluation index, and disease risk evaluation index; among them, the nutrient synthesis ability includes at least one of vitamin synthesis ability, natural pigment synthesis ability, amino acid synthesis ability, glutathione synthesis ability, and bile acid synthesis ability.

3. The data analysis system for gut microbiota according to claim 1, wherein The data upload module is specifically used for: Identifying repeated sequencing fragments in the first intestinal flora sequencing data of the multiple samples, replacing the repeated sequencing fragments with set identifiers to obtain second intestinal flora sequencing data, and forming a second data file, wherein the second data file includes the first intestinal flora sequencing data and / or the second intestinal flora sequencing data; Performing block processing on the data in the second data file to obtain a plurality of data blocks; Multiple data blocks are uploaded in parallel, and a second data file is formed from the uploaded data blocks. The identifiers set in the data of the second data file are restored to the repeated sequencing fragments to form the first data file.

4. The data analysis system for gut microbiota according to any one of claims 1-3, characterized in that, The data analysis module stores an analysis algorithm for each analysis indicator and a processing algorithm for each default processing item; The data analysis module is further configured to sequentially call the processing algorithm of each of the default processing items and the analysis algorithm of the analysis indicator according to the processing nodes in the processing flow, so as to obtain an indicator value of at least one analysis indicator; And / or, the data analysis module is also used to determine the status evaluation data of the intestinal flora of the sample corresponding to the first intestinal flora sequencing data based on the indicator value of at least one analysis indicator corresponding to the first intestinal flora sequencing data, wherein the status evaluation data is obtained based on a fusion algorithm or judgment rule of the indicator value of at least one analysis indicator corresponding to the first intestinal flora sequencing data.

5. The data analysis system for gut microbiota according to claim 1, characterized in that The data mining module is used to: performing at least one level of data grouping on the first intestinal flora sequencing data of the plurality of samples based on the sample labels corresponding to the respective first intestinal flora sequencing data, and performing data mining based on the indicator value of at least one analysis indicator of the first intestinal flora sequencing data in the at least one level of grouping and the grouping labels of each level to obtain mined data; The mining data includes contribution data of sample labels at each level to the at least one analysis indicator, and differences and correlations between at least one analysis indicator corresponding to different groups at the same level and / or groups at different levels.

6. The data analysis system for gut microbiota according to claim 5, wherein The data mining module executes at least one of the following data mining methods: Inputting the index value of at least one analysis index of the first intestinal flora sequencing data in the at least one hierarchical grouping and the grouping labels of each level into a data mining model to obtain the mining data; For groups at any level, determining difference data between at least one analysis indicator in multiple groups at the level, and determining contribution data of group labels at the level to the at least one analysis indicator based on the difference data; For groups at different levels, determining difference data between at least one analysis indicator in each level of grouping, comparing the difference data between at least one analysis indicator in groups at different levels, and determining contribution data of sub-leasing tags at different levels to the at least one analysis indicator; Based on the indicator values of at least one analysis indicator corresponding to different groups at the same level and / or groups at different levels, a comparison data table and / or chart of the at least one analysis indicator is generated, and the differences and correlations between the different groups at the same level and / or at least one analysis indicator corresponding to groups at different levels are displayed through the comparison data table and / or chart.

7. A data analysis method for gut microbiota, characterized in that, include: An interactive page showing the intestinal flora data analysis system, including an upload control, a process setting control, a data analysis control, a data grouping control, and a data mining control; In response to a triggering operation on the upload control, a first data file to be uploaded is obtained, wherein the first data file includes first intestinal flora sequencing data of a plurality of samples; Processing the first intestinal flora sequencing data of the multiple samples in blocks and uploading them in parallel, and displaying the uploaded first intestinal flora sequencing data on an interactive page; In response to a triggering operation on the process setting control, determining at least one analysis indicator corresponding to the first intestinal flora sequencing data, and generating a processing process corresponding to the first intestinal flora sequencing data based on a dependency relationship between the at least one analysis indicator and a default processing item corresponding to the at least one analysis indicator, wherein the processing process includes a processing node corresponding to the default processing item and a processing node corresponding to each of the analysis indicators; In response to a triggering operation on the data analysis control, an available resource amount is obtained, a number of batch processing steps is determined based on the available resource amount and a resource consumption amount of executing the processing flow for each first intestinal flora sequencing data, the number of threads for the batch processing steps is created, and the first intestinal flora sequencing data in the first data file are batch processed by the number of threads for the batch processing steps until the processing of the plurality of first intestinal flora sequencing data in the first data file is completed, thereby obtaining an indicator value of at least one analysis indicator corresponding to each first intestinal flora sequencing data, wherein each thread executes the processing flow for one first intestinal flora sequencing data; In response to a triggering operation on the data grouping control, determining to perform at least one level of data grouping on the first intestinal flora sequencing data of the plurality of samples; In response to a triggering operation on the data mining control, data mining is performed based on an indicator value of at least one analysis indicator of the first intestinal flora sequencing data in each hierarchical group and a grouping label of each hierarchical level to obtain mining data; An analysis report is displayed, where the analysis report is generated based on the indicator value of at least one analysis indicator corresponding to each of the first intestinal flora sequencing data and / or the mining data, and the analysis report includes a visualization chart formed by the at least one analysis indicator and / or the mining data.

8. The method according to claim 7, characterized in that The process setting control includes an indicator selection control and a process generation control; The method further includes: in response to a trigger operation on the metric selection control, presenting selectable metric items; in response to a selection operation on the selectable metric items, determining at least one analysis metric corresponding to the first gut microbiota sequencing data; in response to a trigger operation on the process generation control, generating a processing process corresponding to the first gut microbiota sequencing data; Alternatively, the process setting control further includes a process selection control; The method further includes: in response to a trigger operation on the process selection control, presenting a pre-configured processing process and / or the generated processing process; in response to a selection operation on the pre-configured processing process or the stored processing process, determining a processing process adapted to the at least one analysis metric.

9. The method according to claim 7 or 8, characterized in that, The method further includes: presenting the processing process and each processing node in the processing process on the interaction page; and, during the execution of the processing process, rendering the executed processing nodes through a first rendering attribute and rendering the unexecuted processing nodes through a second rendering attribute, where the first rendering attribute is different from the second rendering attribute.

10. The method according to claim 7, characterized in that, The method further includes: during the batch processing of the first gut microbiota sequencing data of multiple samples, presenting a status identifier and / or a processing progress of each of the first gut microbiota sequencing data, where the status identifier includes at least one of a completion status identifier, a processing status identifier, and an unprocessed status identifier.

11. The method according to claim 7, characterized in that The first gut microbiota sequencing data includes a plurality of sample labels; The method further includes: in response to a trigger operation on the data grouping control, presenting a grouping page, where the grouping page includes a grouping editing area and a label display area, and the grouping editing area includes a hierarchical division symbol; the label display area presents the plurality of sample labels; in response to a trigger operation on the hierarchical division symbol, setting a new hierarchy; in response to a drag operation on the sample label, setting the sample label corresponding to the new hierarchy; in response to an attribute setting operation on the sample label, setting the grouping attribute of the sample label in the new hierarchy to form the data grouping of at least one hierarchy.

12. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data analysis method of gut microbiota according to any one of claims 7-11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions for causing a processor to implement the data analysis method of gut microbiota according to any one of claims 7-11 when executed.

Citation Information

Cited By

  • Targeted nutrient source mining method based on beef cattle intestine type-host gene interaction

    CN121884961A