Algorithm evaluation system taking metagenome sequencing data as training data

By establishing an algorithm evaluation system for metagenomic sequencing data, the problems of poor adaptability and insufficient stability of sample characteristics in the prior art are solved, and the diversity processing and accurate error detection of sequencing data are realized, which improves the adaptability and accuracy of data analysis.

CN120388616APending Publication Date: 2025-07-29MACAU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510388898.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the sequencing data characteristics of different types of samples in the evaluation of metagenomic sequencing data, resulting in species abundance deviations and lack of dynamic monitoring based on acquisition timestamps, which affects the stability of sequencing data.

Method used

By establishing an algorithm evaluation system for metagenomic sequencing data as training data, including data acquisition module, threshold adjustment module, morphological mapping module, abnormal filtering module, two-way verification module and result evaluation module, sequence length value comparison, bacterial abundance parameter screening, pathogen source parameter recording and base mass score analysis, adaptively adjust the abundance threshold, and perform multi-dimensional information correlation and dynamic threshold correction.

Benefits of technology

It improves the stability and accuracy of sequencing data, enhances the adaptability of samples from different sources in data analysis, and achieves more accurate error detection and diversity processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388616A_ABST
    Figure CN120388616A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of metagenome sequencing evaluation, in particular to an algorithm evaluation system taking metagenome sequencing data as training data, which comprises a data acquisition module for reading sequences based on clinical sample characteristics and sequencing fragments, comparing sequence length values and screening flora abundance parameters and pathogen source parameters. According to the method, an acquisition element set is established through sequence length value comparison, flora abundance parameter screening, pathogen source parameter recording and base mass fraction analysis, and data integrity and multi-dimensional information association are ensured in the initial stage. The method comprises the following steps: extracting a median value of a flora abundance parameter and a pathogen source parameter, combining fluctuation monitoring of a sequence length value under an acquisition timestamp, comparing base mass fraction distribution, and adaptively adjusting an abundance threshold value, so that data has targeted adaptation capability under different sample conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of metagenomic sequencing evaluation, and particularly to an algorithm evaluation system for using metagenomic sequencing data as training data. Background Art

[0002] The technical field of metagenomic sequencing evaluation involves high-throughput gene sequencing of microbial communities in the environment or in the host, and quality control, sequence assembly, species classification, and functional annotation of the sequencing data to analyze the composition and functional characteristics of the microbial community.

[0003] In the evaluation process of existing sequencing data, it mainly relies on fixed thresholds for quality screening, which is difficult to adapt to the sequencing data characteristics of different types of samples and is prone to species abundance deviation in complex sample backgrounds. Due to the lack of dynamic monitoring based on the collection timestamp, it is difficult to accurately identify the change trends of sequence length values and base quality scores in different batches, affecting the stability of the sequencing data. Summary of the Invention

[0004] The object of the present invention is to solve the deficiencies in the prior art, and to propose an algorithm evaluation system for using metagenomic sequencing data as training data.

[0005] To achieve the above object, the present invention adopts the following technical solution: An algorithm evaluation system for using metagenomic sequencing data as training data includes: A data acquisition module, which reads sequences based on clinical sample characteristics and sequencing fragments, compares sequence length values, screens out flora abundance parameters and pathogen source parameters, records base quality scores, compares the collection timestamp and the sample batch number, and generates a collection element set; A threshold adjustment module, which extracts the median values of the flora abundance parameters and pathogen source parameters based on the collection element set, monitors the fluctuations of the sequence length values at each collection timestamp, compares the base quality score distributions, records the adaptive abundance threshold, and generates a threshold regulation parameter; A morphological mapping module, which screens out the differences in sequence length values and base quality scores based on the threshold regulation parameter, compares the influence of the collection timestamp, and correlates the distributions of the flora abundance parameters and pathogen source parameters to generate a sequence morphological correlation after correcting the adaptive abundance threshold; An anomaly filtering module, which compares the distributions of the pathogen source parameters and flora abundance parameters based on the sequence morphological correlation, records the sequence error distribution and the potential error pattern library, and checks the threshold range of the base quality score, marks the non-conforming entries, and generates a filtering mark list.

[0006] Preferably, it further includes: The two-way verification module, based on the filtering tag list, conducts numerical statistics on the variance and skewness of base quality scores, generates potential error distribution measurement parameters, performs reverse alignment, and checks for differences in sample batch numbers and pathogen source parameters to generate sequence verification results; The result evaluation module, based on the sequence verification results, conducts reviews of potential error distribution measurement parameters and adaptive abundance thresholds, compares pathogen source parameters with flora abundance parameters, and checks for fluctuations in sequence length values and base quality scores, marks entries outside the deviation range, and generates comprehensive evaluation parameters.

[0007] Preferably, the data acquisition module includes: The sequence acquisition sub-module, based on clinical sample characteristics and sequencing fragments, performs sequence analysis, records sequence numbers, proofreads quality scores, confirms base positions, filters length values and flora abundance parameters, compares pathogen source parameters with sequence number tags, and generates sequence retrieval information; The parameter screening sub-module, based on the sequence retrieval information, conducts checks on length values and base quality scores, calibrates pathogen source parameters, records the correspondence between flora abundance parameters and collection timestamps, compares sample batch numbers with base position sequences, and generates quality screening records; The timing recording sub-module, based on the quality screening records, conducts association tests between collection timestamps and sample batch numbers, summarizes flora abundance parameters, proofreads the distribution of sequence length values and pathogen source parameters, compares base quality scores with pathogen source parameter tags, and generates a collection element set.

[0008] Preferably, the threshold adjustment module includes: The abundance monitoring sub-module, based on the collection element set, conducts statistics on flora abundance parameters and pathogen source parameters, confirms the median interval, proofreads the fluctuations of sequence length values and timestamps, marks flora abundance segmentation information, and generates abundance median elements; The fluctuation comparison sub-module, based on the abundance median elements, records changes in sequence length values at each collection timestamp, compares pathogen source parameters with flora abundance segmentation information, monitors the association of base quality score intervals, and generates a fluctuation distribution comparison; The threshold recording sub-module, based on the fluctuation distribution comparison, conducts comparisons of base quality score distributions, extracts threshold intervals for flora abundance parameters and pathogen source parameters, calibrates the change curve of sequence length values at timestamps, and generates threshold regulation parameters.

[0009] Preferably, the morphology mapping module includes: The morphology screening sub-module, based on the threshold regulation parameters, conducts difference comparisons between sequence length values and base quality scores, marks suspicious sections, extracts comparison results of collection timestamps and pathogen source parameter ranges, and generates difference comparison records; A parameter distribution sub-module, based on the difference comparison record, conducts a comparison of the distribution of the flora abundance parameter and the pathogen source parameter, sorts out the acquisition timestamp marks, checks the associated section of the sequence length value and the base quality score, and generates a distribution mapping record; A morphological correction sub-module, based on the distribution mapping record, conducts an adaptive abundance threshold calibration, monitors the mutual influence between the flora abundance parameter and the pathogen source parameter, verifies the difference section of the base quality score within the acquisition timestamp range, and generates a sequence morphology association.

[0010] Preferably, the anomaly filtering module includes: An error detection sub-module, based on the sequence morphology association, conducts a cross-comparison of the distribution of the pathogen source parameter and the flora abundance parameter to find the abnormal range, locates the abnormal points of the sequence length value and the base quality score, and generates an error section record; A pattern recording sub-module, based on the error section record, conducts an induction of the sequence error distribution, sorts out the potential error pattern library, cross-checks the abnormal points of the base quality score and the pathogen source parameter, and compares the distribution of the flora abundance parameter and the sequence length value, and generates a potential distribution record; A marking output sub-module, based on the potential distribution record, conducts a test of the base quality score threshold range, calibrates the abnormal position, compares the sequence error distribution with the difference of the flora abundance parameter, and identifies the items with the pathogen source parameter exceeding the range, and generates a filtering mark list.

[0011] Preferably, the two-way verification module includes: A statistics sub-module, based on the filtering mark list, conducts a cross-calculation of the variance and skewness values of the base quality score, sorts out the abnormal information of the flora abundance parameter, records the abnormal section of the pathogen source parameter, and generates an error measurement factor; A reverse verification sub-module, based on the error measurement factor, conducts a difference positioning of the pathogen source parameter, re-checks the abnormal information of the flora abundance parameter, calibrates the deviation point of the base quality score and the difference distribution of the sample batch number, and generates a batch verification record; A difference comparison sub-module, based on the batch verification record, conducts a segmented detection of the sample batch number and the pathogen source parameter, summarizes the abnormal position of the base quality score, and compares the flora abundance parameter with the sequence length value information, and generates a sequence verification result.

[0012] Preferably, the result evaluation module includes: An error review sub-module, based on the sequence verification result, conducts a count of the parameters for calculating the potential error distribution, cross-checks the pathogen source parameter and the flora abundance parameter, and compares the basic information of the sequence length value and the base quality score, and generates an error calculation record; A parameter comparison sub-module, based on the error measurement record, compares the pathogen source parameter with the flora abundance parameter, extracts deviation information, verifies the corresponding relationship between the base quality score and the sequence length value, correlates the potential error distribution measurement parameter, and generates a parameter comparison record; A fluctuation marking sub-module, based on the parameter comparison record, conducts a fluctuation check on the sequence length value and the base quality score, confirms the deviation section, verifies whether the pathogen source parameter and the flora abundance parameter meet the measurement reference range, and generates a comprehensive evaluation parameter.

[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by comparing the sequence length values, screening the flora abundance parameters, recording the pathogen source parameters, and analyzing the base quality scores, a set of acquisition elements is established to ensure the integrity of the data and the association of multi-dimensional information at the initial stage. By extracting the median values of the flora abundance parameter and the pathogen source parameter, combining with the fluctuation monitoring of the sequence length value under the acquisition timestamp, comparing the base quality score distribution, and adaptively adjusting the abundance threshold, the data has targeted adaptation ability under different sample conditions. Based on the threshold control parameter, screening the differences between the sequence length value and the base quality score, correlating the influence of the acquisition timestamp on the data stability, combining the distribution of the flora abundance parameter and the pathogen source parameter, and correcting the adaptive abundance threshold to make the identification standard more dynamic. By comparing the distributions of the pathogen source parameter and the flora abundance parameter, combining the sequence error distribution with the records in the potential error pattern library, screening the abnormal entries of the base quality score, making the error detection of the sequencing data more accurate. Based on the filtering mark list, conducting the variance and skewness statistics of the base quality score, combining with the reverse alignment result, and verifying the differences between the sample batch numbers and the pathogen source parameters, making the error screening have two-way verification ability. The overall processing logic constructs a hierarchical metagenomic data evaluation method through precise data acquisition, dynamic threshold adjustment, multi-dimensional form comparison, abnormal data filtering, and two-way verification, improving the stability and accuracy of the sequencing data, enhancing the adaptability of different source samples in data analysis, and making the diversity processing ability of the sequencing data stronger. Description of the Drawings

[0014] Figure 1 It is the system flowchart of the present invention. Specific Embodiments

[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0016] Please refer to Figure 1, the present invention provides a technical solution: an algorithm evaluation system using metagenomic sequencing data as training data includes: A data acquisition module, which reads sequences based on clinical sample characteristics and sequencing fragments, compares sequence length values, screens for microbiota abundance parameters and pathogen source parameters, records base quality scores, compares the acquisition timestamp and sample batch number, and generates an acquisition element set; A threshold adjustment module, which extracts the median values of microbiota abundance parameters and pathogen source parameters based on the acquisition element set, monitors the fluctuations of sequence length values at each acquisition timestamp, compares the base quality score distributions, records the adaptive abundance threshold, and generates a threshold regulation parameter; A morphological mapping module, which screens for differences in sequence length values and base quality scores based on the threshold regulation parameter, compares the influence of the acquisition timestamp, correlates the distributions of microbiota abundance parameters and pathogen source parameters, and generates a sequence morphological correlation after correcting the adaptive abundance threshold; An anomaly filtering module, which compares the distributions of pathogen source parameters and microbiota abundance parameters based on the sequence morphological correlation, records the sequence error distribution and the potential error pattern library, checks the range of base quality scores, marks the non-conforming entries, and generates a filtering mark list; A two-way verification module, which performs numerical statistics on the variance and skewness of base quality scores based on the filtering mark list, generates potential error distribution measurement parameters, performs reverse alignment, and checks the differences between the sample batch number and pathogen source parameters to generate a sequence verification result; A result evaluation module, which reviews the potential error distribution measurement parameters and the adaptive abundance threshold based on the sequence verification result, compares the pathogen source parameters with the microbiota abundance parameters, checks the fluctuations of sequence length values and base quality scores, marks the entries outside the range, and generates a comprehensive evaluation parameter.

[0017] The data acquisition module includes: A sequence acquisition sub-module, which parses sequences and records sequence numbers based on clinical sample characteristics and sequencing fragments, proofreads quality scores and confirms base positions, screens length values and microbiota abundance parameters, compares pathogen source parameters with sequence number markers, and generates sequence retrieval information; A parameter screening sub-module, which checks length values and base quality scores and calibrates pathogen source parameters based on the sequence retrieval information, records the correspondence between microbiota abundance parameters and acquisition timestamps, compares the sample batch number with the base position sequence, and generates a quality screening record; A time series recording sub-module, which performs an association test between the acquisition timestamp and the sample batch number and summarizes the microbiota abundance parameters based on the quality screening record, proofreads the distributions of sequence length values and pathogen source parameters, compares the base quality scores with pathogen source parameter markers, and generates an acquisition element set.

[0018] Specifically, for the received clinical sample data and sequencing fragment description file, record the base quality score of each sequencing fragment and the corresponding clinical sample identifier. If there are significant deviations, reconfirm according to the empirical set values. For example, the base quality score can be compared with the range of 0 to 40, and the sequence length can be compared with the range of 75bp to 300bp. When parsing all fragments, read their numbers line by line, verify the base position information and check whether it is consistent with the clinical sample characteristic information, and then associate and compare the pathogen source parameters with each sequence number. While confirming that the pathogen information matches the existing reference list information, use the set threshold range calculated based on historical experience to detect those base sites that may conflict with or not fully match the pathogen source. During the process, register the microbial community abundance parameter and the sequence length one by one. If it is found that the microbial community abundance parameter exceeds the standard threshold set in advance based on literature statistics (for example, measure the microbial community abundance according to the ratio of 0 to 1 and mark the part exceeding 0.5 as the high abundance interval), then retain the additional annotation information and conduct repeated verification. Whenever the quality of a sequencing fragment is confirmed, record all the obtained quality information and the sample number in a temporary database, and compare them one by one with the index positions given in the sequencing fragment description file. If there are differences in the base positions and their position information in the description file, the fragment needs to be reread and the integrity of the base sequence needs to be confirmed again. After all the control checks are completed, organize the finally successfully matched and confirmed sequence data into retrieval data and generate sequence retrieval information.

[0019] Based on the previously obtained sequence retrieval information, first screen the sequences according to the pre-given length range, such as 75bp to 300bp. If the length value of a certain sequence deviates from this range, it is considered potentially abnormal and reconfirmed. At the same time, compare the base quality score of each sequence with the range of 0 to 40. For items outside this range, further distinguish whether they need to be discarded or manually reviewed based on empirical thresholds, such as the high-confidence score range extracted from historical sequencing results, and calibrate the pathogen source parameters according to the results of these treatments. For example, refer to a pre-compiled pathogen reference list and list the main sequencing characteristics of known pathogens and their corresponding determination parameters in a clear text form, and then compare the specific sequence retrieval results item by item with this list. If it is found that an item has a large difference from the pathogen source parameters in this list, it needs to be key-checked. Then record the mapping relationship between the microbial community abundance parameter and the collection timestamp, generate the corresponding batch identifier with reference to information such as the experimental date and sampling equipment number, and then organize and verify the base positions in the sequence to check whether they match the corresponding timestamp and batch information. If there are inconsistent associations, record them in the abnormal mark list in a clear manner. After all the comparisons are completed, summarize the qualified sequence corresponding status and finally obtain the quality screening record.

[0020] Based on the quality screening records obtained in the previous step, first read the timestamps therein and check them one by one with the corresponding batch numbers. Select the sequence entries that are still considered valid after screening and statistically summarize the microbial community abundance parameters. For example, divide the microbial community abundance into a low abundance interval of 0 to 0.3, a medium abundance interval of 0.3 to 0.7, and a high abundance interval of above 0.7. And attach pathogen source parameters to each sequence in the temporary storage location. Compare the distribution of these source parameters with the sequence length values. If some sequence length values are concentrated at the extreme upper or lower limits, then combine with the previously set empirical interval such as 75bp to 300bp to judge whether additional repeated sequencing operations need to be performed. Then compare the corresponding base quality scores with the pathogen source parameter marks. If it is found that they are not within the pre-set score interval or conflict with the classification information in the source parameter list, then mark this sequence as a target of concern. After completing the above inspection operations item by item, correspond the associated information with the collection timestamp and batch number, and finally summarize the results to form a collection element set, generating a collection element set.

[0021] The threshold adjustment module includes: An abundance monitoring sub-module, based on the collection element set, statistically calculates the microbial community abundance parameters and pathogen source parameters and confirms the median interval, proofreads the fluctuation of the sequence length value and the timestamp, marks the segmented information of the microbial community abundance, and generates abundance median elements; A fluctuation comparison sub-module, based on the abundance median elements, records the changes in the sequence length values at each collection timestamp, compares the pathogen source parameters with the segmented information of the microbial community abundance, monitors the association of the base quality score intervals, and generates a fluctuation distribution comparison; A threshold recording sub-module, based on the fluctuation distribution comparison, compares the base quality score distributions and extracts the threshold intervals of the microbial community abundance parameters and pathogen source parameters, calibrates the change curve of the sequence length value at the timestamp, and generates threshold regulation parameters.

[0022] Specifically, according to the collected element set sorted out before, first batch - count the microbial abundance parameters among them and conduct a corresponding analysis with the pathogen source parameters. If the number of sequences with low microbial abundance is much smaller than the number of medium - abundance and high - abundance sequences, then combine the pre - established abundance segmentation ranges, for example, 0 to 0.3 is the low - abundance interval, 0.3 to 0.7 is the medium - abundance interval, and above 0.7 is the high - abundance interval, to confirm the median interval. Subsequently, compare these intervals with the timestamp information and check whether the sequence length values have fluctuated within a specific time period. For example, when the sequence length value is lower than 75bp or higher than 300bp in a certain time period, check whether the pathogen source parameters during this period match the reference list determined previously, and record the suspicious entries separately. After the check, mark the corresponding abundance segmentation results in the internal data structure, extract the median value segment from the statistical information and include it in the abundance element list used for subsequent data analysis, generating the abundance median element.

[0023] After obtaining the abundance median element, check the microbial abundance parameters and pathogen source parameters one by one, and record the change trend of the sequence length value in chronological order of the timestamp. If it is observed that the sequence length value repeatedly exceeds the empirical interval in some time periods, such as being higher than 300bp or lower than 75bp, then conduct a cause analysis in combination with the previously determined pathogen source information. For example, compare with the reference list of source parameters collected before to check whether there is a special pathogen type causing a significant increase or decrease in the sequencing fragments. At the same time, conduct an interval correlation detection on the base quality score, for example, compare the quality score with the interval of 0 to 40. If some sequences frequently fall into the lower score segment, mark them in a separate list. After completing these monitors, summarize all the information into a fluctuation distribution comparison and distinguish different abundance segments and the corresponding pathogen source parameters in it, generating the fluctuation distribution comparison.

[0024] Based on the obtained fluctuation distribution comparison, first correspond the overall distribution of the base quality score with the corresponding timestamp and conduct a cross - comparison with the microbial abundance parameters and pathogen source parameters. If it is found that the microbial abundance is high in some time periods while the base quality score is generally low, then refer to the score critical range set through statistical analysis beforehand, for example, consider the base quality score below 10 as a low value, to confirm whether it is necessary to adjust the subsequent test steps. Then, after confirming whether there are frequent jumps in the sequence length values of each time period, extract the corresponding threshold intervals based on the median characteristics of the historical sequencing data, which include the special value ranges of the pathogen source parameters. Recalculate the change curve of the sequence length value under the timestamp for these intervals and check them step by step at a fixed step size. If there is an obvious out - of - bounds at multiple consecutive collection points, reset the threshold interval to a stricter or wider dynamic range. After the steps are completed, record the latest threshold correction results in the database, and finally summarize the threshold regulation parameters that can be used for subsequent analysis and testing, generating the threshold regulation parameters. [[ID=,8]]

[0025] The morphology mapping module includes: A morphology screening sub-module, which, based on the threshold regulation parameter, compares the sequence length value with the base quality score difference and marks the suspicious sections, extracts the comparison result of the acquisition timestamp and the pathogen source parameter range, and generates a difference comparison record; A parameter distribution sub-module, which, based on the difference comparison record, compares the distribution of the flora abundance parameter with the pathogen source parameter and sorts out the acquisition timestamp marks, checks the associated sections of the sequence length value and the base quality score, and generates a distribution mapping record; A morphology correction sub-module, which, based on the distribution mapping record, performs adaptive abundance threshold calibration and monitors the mutual influence of the flora abundance parameter and the pathogen source parameter, verifies the difference sections of the base quality score within the acquisition timestamp range, and generates a sequence morphology association.

[0026] Specifically, based on the threshold regulation parameter obtained above, first read the respective interval ranges of the sequence length value and the base quality score listed therein and perform a difference comparison. For the sequence length value, a reference interval of 75bp to 300bp statistically obtained from historical sequencing experience can be selected. For the base quality score, a range from 0 to 40 can be selected as a preliminary limit and it is determined whether several additional values need to be increased or decreased through the existing sequencing statistics summary to form a more accurate difference boundary. During the comparison process, check each item of the length value and the base quality score of each sequence one by one. If it is found that the sequence length value deviates greatly from the above intervals, it is determined as a suspicious section. For these suspicious sections, mark them in combination with the acquisition timestamp and check whether there is a corresponding record in the previously obtained pathogen source parameter range. If this section deviates from both the allowable ranges of the sequence length value and the base quality score and there is no corresponding pathogen source parameter match, it is recorded as an abnormal difference point. At the same time, include the timestamp of the deviation and the corresponding flora abundance parameter in the list for subsequent processing. After all the screenings are completed, review the acquisition timestamp and the pathogen source parameter interval that may be associated. If some entries deviate from the sequence length value but the base quality score meets the interval, temporarily mark them as to be observed and continue to compare and analyze them in the subsequent steps. Finally, generate a difference comparison record.

[0027] According to the identification information of each suspicious section in the differential comparison record, classify and summarize these identifications with the ranges of the microbial flora abundance parameters and pathogen source parameters. If the microbial flora abundance parameter is less than 0.3 jointly deduced from historical literature and measured data, it is regarded as low abundance. When correlating the corresponding entries with the pathogen source parameters, if it is found that the parameter deviates from the reference range previously statistically obtained from clinical samples, it is recorded in a separate list. Conduct the same comparison for other abundance intervals in turn, and combine the collection timestamp under specific entries to examine the distribution within different batch numbers and different date ranges. Then, check the corresponding positions of the sequence length value and the base quality score during these periods. If both are in the suspicious interval, it accumulates further reference items for subsequent processing. If only one party deviates, it is judged whether additional manual verification is required according to the deviation degree recorded in the differential comparison record. For high-abundance entries, determine whether there are abnormalities according to the classification standard greater than 0.7. At the same time, conduct a re-comparison of the pathogen source parameters. After all entries are classified, form a multi-dimensional comparison of the obtained distribution information. After all the results are sorted out, finally generate a distribution mapping record.

[0028] Based on the data of each interval presented in the distribution mapping record, select the entries with microbial flora abundance parameters between 0.3 and 0.7 and normal pathogen source parameters as the main reference. Check the distribution of the base quality scores of these entries item by item and compare them with the statistical data of the known pathogen source parameters. If some entries continuously exceed the previously defined threshold in terms of the base quality score, it is necessary to re-judge whether the entry was misclassified in the previous microbial flora abundance classification. Subsequently, calibrate these newly identified situations together with the sequence length value to determine the adaptive abundance threshold. Position the threshold on the basis of not exceeding an empirical range from the historical sequencing data. For example, select 0.05 or 0.1 as the step size to adjust downward or upward according to the recent error assessment of the sequencing results. If the pathogen source parameters remain consistent within the adjustment range and the sequence length values corresponding to the collection timestamps are mostly distributed between 75bp and 300bp, it is considered that the calibration is effective. After the abundance parameters of all entries are re-allocated, check the differential sections of the base quality scores within each timestamp range, and finally generate a sequence form association.

[0029] The abnormal filtering module includes: An error detection sub-module, based on the sequence form association, conducts a cross-comparison of the distributions of the pathogen source parameters and the microbial flora abundance parameters to find the abnormal range, locates the abnormal points of the sequence length value and the base quality score, and generates an error section record; A pattern recording sub-module, based on the error section record, conducts an induction of the sequence error distribution and organizes a potential error pattern library, cross-checks the abnormal points of the base quality score and the pathogen source parameters, and compares the distributions of the microbial flora abundance parameters and the sequence length values to generate a potential distribution record; The marker output sub-module conducts a base quality score threshold range test and calibrates abnormal positions based on the potential distribution record, compares the sequence error distribution with the differences in microbial community abundance parameters, identifies entries with pathogen source parameters outside the range, and generates a filtered marker list.

[0030] Specifically, based on the induction results of pathogen source parameters and microbial community abundance parameters in the sequence morphology association, the suspicious distributions are compared one by one, and records with sequence length values outside the set range of 75bp to 300bp or base quality scores lower than the empirical lower limit obtained from historical sequencing data (such as 10 or 15) are found. These records are corresponded to the previously statistically obtained pathogen source parameters. If there are many such entries simultaneously within a collection time stamp, it indicates that there may be an abnormality during this time period. After locating the suspicious base quality score positions, they can be cross-checked horizontally with the typical distribution corresponding to the pathogen source parameters. If the gap is large, these entries are included in the abnormal range. After extracting all abnormal entries one by one and cross-checking them with the microbial community abundance parameters, if the abundance is lower than 0.3 or higher than 0.7 and does not match the pathogen source parameters, it is uniformly regarded as a serious abnormality. After completing the above cross-comparisons, all sequence length values and base quality score indexes with obvious deviations are incorporated into the summary list, and finally, an error section record is generated.

[0031] Based on the relevant indexes of each entry in the error section record, the entire sequencing information of the entry within the collection time stamp is compared and cross-checked according to the previously integrated pathogen source parameters and microbial community abundance parameters. First, the base quality scores of all abnormal positions are extracted and compared with the preset numerical range (such as 0 to 40). If it is lower than 10, it is considered that there may be a systematic deviation or local sequencing failure at this abnormal position and is incorporated into the potential error pattern library. Subsequently, the microbial community abundance parameters corresponding to these positions are checked. If the abundance is between 0.3 and 0.7 and the pathogen source parameters are inconsistent with the previous records, it is marked as a potential conflict type. Then, the distribution corresponding to the sequence length value is compared in the same way. If the phenomenon of being higher than 300bp or lower than 75bp continuously appears, it is listed as a potential extreme pattern. After all indexes are cross-checked, similar abnormal patterns are summarized and incorporated into the potential error pattern library, and finally, a potential distribution record is generated.

[0032] Based on the potential distribution record, first sort out the indices of all positions with abnormal base quality scores and check them according to the threshold range set previously based on historical sequencing results (such as 0 to 40). If records lower than 10 or higher than 35 are found, they are marked as suspicious points and compared with the differences in the microbial abundance parameters registered in the sequence error distribution. If some entries are abnormal in terms of base quality scores but the microbial abundance parameters do not exceed the previously set range of 0.3 to 0.7, they are only retained in the general abnormal list. If it is found that the pathogen source parameters are significantly higher or lower than the value range summarized from clinical samples in advance and are accompanied by deviations in base quality scores or microbial abundance parameters, they are marked as out-of-range entries. After completing all the comparisons, collect each abnormal piece of information and finally generate a filtering marker list.

[0033] The two-way verification module includes: A statistical sub-module that, based on the filtering marker list, performs cross-calculation of the variance and skewness values of base quality scores and sorts out abnormal information of microbial abundance parameters, records abnormal sections of pathogen source parameters, and generates error measurement elements; A reverse verification sub-module that, based on the error measurement elements, locates the differences in pathogen source parameters and re-checks the abnormal information of microbial abundance parameters, marks the deviation points of base quality scores and the differential distribution of sample batch numbers, and generates batch verification records; A difference comparison sub-module that, based on the batch verification records, performs segmented detection of sample batch numbers and pathogen source parameters and summarizes the positions of abnormal base quality scores, compares the microbial abundance parameters with the sequence length value information, and generates sequence verification results.

[0034] Specifically, based on the previously obtained list of filtering tags, first retrieve the base quality score deviation records listed therein and extract the corresponding microbial abundance parameters and pathogen source parameters item by item. Then, read the variance and skewness calculation requirements for each entry in the internal data structure. After arranging the base quality scores of each entry in sequence, use the method based on the cumulative difference square to obtain the variance value, and at the same time use the method of splitting data symmetry to obtain the skewness value. By comparing these variance and skewness results, determine whether the distribution of base quality scores shows an obvious tilt. For example, if a large number of values are concentrated in a set lower interval such as between 0 and 10, it indicates that the skewness may be greater than the average skewness benchmark calculated from historical sequencing statistics (for example, setting the average skewness of multiple historical batches to about 1.0, and if the current skewness exceeds 1.5, it can be regarded as an obvious tilt). For abnormal information of microbial abundance parameters, continue to read the abundance values recorded in the list and compare them with the interval between 0.3 and 0.7. If some entries are lower than 0.3 or higher than 0.7 and do not match the pathogen source parameters, classify them as severely abnormal and attach corresponding marks. During this period, also classify the entries with pathogen source parameters significantly deviating from the previously defined reference range (such as the corresponding interval of known pathogen types determined from clinical statistics) into a separate list. Finally, after summarizing the variance and skewness calculations of all entries and the corresponding abundance and source abnormal information, output these sorted results to form error measurement elements.

[0035] Based on the error measurement elements, first extract all pathogen source parameter deviation information from them and compare it with the existing abnormal data of microbial abundance parameters. Locate the difference points by retrieving the specific values of pathogen source parameters row by row in the original summary data. For those entries that show as exceeding the previously established range based on experience in statistical comparison (for example, the proportion of some rare pathogens is higher than 5% or lower than 0.1%, and this proportion is obtained by comprehensively statistical analysis of a large number of clinical sequencing results), define them as high-priority difference targets. Then, retrieve the abnormal information of microbial abundance parameters again to check whether these high-priority difference targets also show a situation inconsistent with the original preset in the abundance interval (such as less than 0.3 or greater than 0.7). If there are deviations in both aspects, list them as key review targets. Then, further compare the base quality score deviation points of these targets with the corresponding sample batch numbers. For the batch numbers, first define a retrieval range from 1 to N and compare their difference distributions with the pathogen source parameters one by one. If such abnormalities occur frequently in the same batch, mark them additionally and record them in the subsequent correction list. After completing these operations, summarize all batch verification information uniformly to generate a batch verification record.

[0036] Based on the batch proofreading records, first read the specific segmented definitions of the sample batch number and pathogen source parameters. The source parameters can be segmented according to the previously obtained range in clinical statistics (for example, the occurrence rate of some pathogens in the range of 0.5% to 2% is regarded as medium, less than 0.5% is regarded as rare, and more than 2% is regarded as common). In this way, classify the pathogen source parameters under each batch, and retrieve the flora abundance parameters within the same batch for vertical comparison. For the abnormal positions of the base quality scores, map them to the corresponding batch numbers item by item according to the deviation point indexes obtained from the previous statistics. If the base quality scores of a large number of entries are lower than 10 or higher than 35 under a certain batch and there are significant differences from the pathogen source parameters, it is regarded as a key concern item. After concentrating these deviation items, compare the difference between the flora abundance parameters and the known safe range (0.3 to 0.7) again. If there are obvious deviations at the same time, it indicates that there is a high probability of sequencing abnormality in this batch. Finally, integrate all the segmented detection results and correspond them to the respective sequence length value information (usually check whether there are extreme over-limit according to the standard range of 75bp to 300bp). After completion, summarize all the comparison information into the sequence verification results.

[0037] The result evaluation module includes: The error review sub-module, based on the sequence verification results, calculates the potential error distribution measurement parameters, cross-checks the pathogen source parameters and the flora abundance parameters, compares the sequence length values with the basic information of the base quality scores, and generates an error measurement record; The parameter comparison sub-module, based on the error measurement record, compares the pathogen source parameters with the flora abundance parameters, extracts the deviation information, verifies the corresponding relationship between the base quality scores and the sequence length values, and associates the potential error distribution measurement parameters to generate a parameter comparison record; The fluctuation marking sub-module, based on the parameter comparison record, checks the fluctuations of the sequence length values and the base quality scores, confirms the deviation sections, and verifies whether the pathogen source parameters and the flora abundance parameters meet the measurement reference range to generate a comprehensive evaluation parameter.

[0038] Specifically, based on the sequence verification results, all potential error distribution information is first listed and the corresponding measurement parameters, such as variance and skewness, are retrieved accordingly. The pathogen source parameters and bacterial abundance parameters are checked in these potential error entries. If the source parameters of some entries are obviously in conflict with the previously defined statistical boundaries of known pathogens (for example, the proportion of common pathogens is between 2% and 10%, and the proportion of rare pathogens is less than 0.5%), the basic information of the sequence length values of such entries (whether it is less than 75bp or greater than 300bp) and base quality scores (whether it falls between 0 and 40) are further compared. After completing the comparison for each entry, it is recorded, and the part that is confirmed to have no reference interval to match or has exceeded the previously listed range is recorded as high-risk content. After all potential error distribution measurement parameters are counted, the corresponding entry information is merged, and the error measurement record is finally output.

[0039] Based on the error measurement records, we first summarize the pathogen source parameters and bacterial abundance parameters marked in the previous article, and compare the two one by one, paying special attention to those entries with abundance greater than 0.7 or less than 0.3 and pathogen source parameters in the rare range (such as less than 0.5%). Then, we record the specific correlation between the base quality score and sequence length value of these entries. The base quality score value of each entry is compared with the range of 0 to 40. If most of them fall between 5 and 15 and the sequence length value does not fit the range of 75 bp to 300 bp, they are marked as key concerns. Then, they are cross-referenced with the previously identified potential error distribution measurement parameters (including variance and skewness). If the skewness of some entries is determined to be significantly higher than the pre-calculated reference benchmark (such as more than 1.5) and the variance is also large (such as more than 200), they are additionally searched and the corresponding information is summarized in the same record. Finally, after all entries are sorted and formed into an overall comparison, the parameter comparison record is generated.

[0040] Based on the parameter comparison record, first retrieve the corresponding list between the sequence length value and the base quality score and expand it internally in chronological order or batch order. Vertically check the length value of each sequence against the base quality score range from 0 to 40. If some entries repeatedly appear in the rows and columns of extremely low base quality scores (below 10) or extremely high length values (above 300bp), they are marked as having obvious fluctuations. Further, check these fluctuation points against the reference range corresponding to the pathogen source parameters (for example, if the pathogens corresponding to the source parameters usually do not have extremely long or extremely short sequencing fragments but anomalies occur), and also perform the same check on the microbiota abundance parameters to see if they are below 0.3 or above 0.7. If both deviate, include this entry in the final list of deviation segments. After that, traverse the pathogen source parameters and microbiota abundance parameters of all entries and compare them with the measurement reference ranges derived from historical sequencing data or clinical statistical data. If the overall distribution continues to be in the abnormal segment, mark it in the same way. After completing these verification operations, record the integrated deviation segment as the comprehensive evaluation parameter.

Claims

1. An algorithm evaluation system using metagenomic sequencing data as training data, characterized in that The system includes: A data acquisition module, which reads sequences based on clinical sample characteristics and sequencing fragments, compares sequence length values, screens for microbiota abundance parameters and pathogen source parameters, records base quality scores, compares the acquisition timestamp and the sample batch number, and generates an acquisition element set; A threshold adjustment module, which extracts the median values of the microbiota abundance parameters and pathogen source parameters based on the acquisition element set, monitors the fluctuations of sequence length values at each acquisition timestamp, compares the base quality score distributions, records the adaptive abundance threshold, and generates threshold regulation parameters; A morphological mapping module, which screens for differences in sequence length values and base quality scores based on the threshold regulation parameters, compares the influence of the acquisition timestamp, correlates the distributions of the microbiota abundance parameters and pathogen source parameters, and generates a sequence morphological correlation after correcting the adaptive abundance threshold; An anomaly filtering module, which compares the distributions of the pathogen source parameters and the microbiota abundance parameters based on the sequence morphological correlation, records the sequence error distribution and the potential error pattern library, checks the base quality score threshold range, marks the non-conforming entries, and generates a filtering mark list.

2. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 1, wherein It also includes: A two-way verification module, which performs numerical statistics on the base quality score variance and skewness based on the filtering mark list, generates potential error distribution measurement parameters, performs reverse alignment, and checks the differences between the sample batch number and the pathogen source parameters, and generates a sequence verification result; A result evaluation module, which reviews the potential error distribution measurement parameters and the adaptive abundance threshold based on the sequence verification result, compares the pathogen source parameters with the microbiota abundance parameters, checks the fluctuations of the sequence length values and the base quality scores, marks the entries outside the deviation range, and generates a comprehensive evaluation parameter.

3. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 1, wherein The data acquisition module includes: A sequence acquisition sub-module, which parses sequences and records sequence numbers based on clinical sample characteristics and sequencing fragments, proofreads quality scores and confirms base positions, screens length values and microbiota abundance parameters, compares pathogen source parameters with sequence number marks, and generates sequence retrieval information; A parameter screening sub-module, which checks the length values and base quality scores and calibrates the pathogen source parameters based on the sequence retrieval information, records the correspondence between the microbiota abundance parameters and the acquisition timestamp, compares the sample batch number with the base position sequence, and generates a quality screening record; A timing recording sub-module, which performs an association test between the acquisition timestamp and the sample batch number and summarizes the microbiota abundance parameters based on the quality screening record, proofreads the distributions of the sequence length values and the pathogen source parameters, compares the base quality scores with the pathogen source parameter marks, and generates an acquisition element set.

4. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 1, characterized in that, The threshold adjustment module includes: An abundance monitoring sub-module, which performs statistics on the microbiota abundance parameters and pathogen source parameters and confirms the median interval based on the acquisition element set, proofreads the fluctuations of the sequence length values and the timestamp, marks the microbiota abundance segmentation information, and generates abundance median elements; The fluctuation comparison sub-module records the change of the sequence length value at each acquisition timestamp based on the median abundance factor, compares the pathogen source parameter with the segmented information of the microbial flora abundance, monitors the association of the base quality score interval, and generates a fluctuation distribution control. The threshold recording sub-module compares the base quality score distribution based on the fluctuation distribution control, extracts the threshold intervals of the microbial flora abundance parameter and the pathogen source parameter, calibrates the change curve of the sequence length value at the timestamp, and generates a threshold regulation parameter.

5. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 1, characterized in that, The morphology mapping module includes: The morphology screening sub-module compares the difference between the sequence length value and the base quality score based on the threshold regulation parameter, marks the suspicious sections, extracts the comparison result of the acquisition timestamp and the range of the pathogen source parameter, and generates a difference comparison record. The parameter distribution sub-module compares the distribution of the microbial flora abundance parameter and the pathogen source parameter based on the difference comparison record, organizes the acquisition timestamp marks, checks the associated sections of the sequence length value and the base quality score, and generates a distribution mapping record. The morphology correction sub-module calibrates the adaptive abundance threshold based on the distribution mapping record, monitors the mutual influence of the microbial flora abundance parameter and the pathogen source parameter, checks the difference sections of the base quality score within the acquisition timestamp range, and generates a sequence morphology association.

6. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 1, wherein The anomaly filtering module includes: The error detection sub-module cross-compares the distribution of the pathogen source parameter and the microbial flora abundance parameter based on the sequence morphology association, searches for the abnormal range, locates the abnormal points of the sequence length value and the base quality score, and generates an error section record. The pattern recording sub-module summarizes the sequence error distribution based on the error section record, organizes the potential error pattern library, cross-checks the abnormal points of the base quality score and the pathogen source parameter, and compares the distribution of the microbial flora abundance parameter and the sequence length value, and generates a potential distribution record. The marking output sub-module checks the threshold range of the base quality score based on the potential distribution record, marks the abnormal positions, compares the sequence error distribution with the difference of the microbial flora abundance parameter, and identifies the items where the pathogen source parameter exceeds the range, and generates a filtering mark list.

7. The algorithm evaluation system using the metagenomic sequencing data as training data according to claim 2, characterized in that The two-way verification module includes: The statistics sub-module cross-calculates the variance and skewness values of the base quality score based on the filtering mark list, organizes the abnormal information of the microbial flora abundance parameter, records the abnormal sections of the pathogen source parameter, and generates an error measurement factor. The reverse verification sub-module locates the difference of the pathogen source parameter based on the error measurement factor, re-checks the abnormal information of the microbial flora abundance parameter, calibrates the deviation points of the base quality score and the difference distribution of the sample batch number, and generates a batch verification record. The difference comparison sub-module performs segmented detection of the sample batch number and the pathogen source parameter based on the batch verification record, summarizes the abnormal positions of the base quality score, and compares the microbial flora abundance parameter with the sequence length value information, and generates a sequence verification result.

8. The algorithm evaluation system using the metagenomic sequencing data according to claim 2 as training data, characterized in that, The result evaluation module includes: The error review sub-module counts the parameters for calculating the potential error distribution based on the sequence verification result, cross-checks the pathogen source parameter and the microbial flora abundance parameter, and compares the basic information of the sequence length value and the base quality score, and generates an error calculation record. A parameter comparison sub-module, based on the error measurement record, compares the pathogen source parameter with the flora abundance parameter, extracts deviation information, verifies the correspondence between the base quality score and the sequence length value, correlates the potential error distribution measurement parameter, and generates a parameter comparison record; A fluctuation marking sub-module, based on the parameter comparison record, conducts a fluctuation check on the sequence length value and the base quality score, confirms the deviation section, verifies whether the pathogen source parameter and the flora abundance parameter meet the measurement reference range, and generates a comprehensive evaluation parameter.