A specimen biological information analysis and verification system and method based on big data
Patent Information
- Application Number
- CN202610858873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-18
AI Technical Summary
但现有共识方法多采用固定阈值或简单多数投票原则,未考虑不同流程在历史表现中可信度的差异,导致基准构建结果易受低质量流程的错误输出干扰,共识标准的精准度和代表性不足
[0040]By collecting the actual outputs of multiple gene variation analysis processes and constructing an internal dynamic benchmark using weighted consensus frequency, the traditional method is completely free from its reliance on manually labeled dynamic consensus benchmark features, significantly reducing the human and time costs of bioinformatics analysis process verification.
Smart Images

Figure CN122778064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a specimen bioinformatics analysis and verification system and method based on big data. Background Technology
[0002] When faced with the same sample, different institutions or research teams typically develop and use their own bioinformatics analysis workflows for gene variant identification. These workflows differ in algorithm design, parameter configuration, and database dependencies, resulting in a lack of uniformity and comparability in their variant detection outputs. How to objectively and efficiently evaluate and verify the reliability and consistency of these gene variant analysis workflows is a key technical challenge currently facing the field of bioinformatics.
[0003] Currently, the validation of gene variation analysis workflows typically relies on manually constructed external reference sets. This involves expert teams using rigorous manual review, experimental verification, or intersection analysis of multiple workflow outputs to determine a set of variant sites considered to be real and existing as reference standards. The outputs of the workflow to be validated are then compared with these reference standards to calculate metrics such as sensitivity and specificity. However, this approach has significant limitations: firstly, constructing reference sets is time-consuming, costly, and highly dependent on the experience and judgment of domain experts, leading to significant differences in reference sets between different institutions and poor reproducibility; secondly, with the rapid iteration of sequencing technologies and analytical algorithms, static reference standards struggle to adapt to the constantly evolving validation requirements of workflows, easily becoming outdated and ineffective. Furthermore, this validation method lacks a unified framework, making it difficult to compare the performance evaluation results of different workflows across different systems.
[0004] Some studies have attempted to build consensus benchmarks based on the consistency of outputs from multiple processes to alleviate reliance on external reference standards. However, existing consensus methods often employ fixed thresholds or simple majority voting principles, failing to consider the differences in the credibility of different processes in their historical performance. This makes the benchmark construction results susceptible to interference from erroneous outputs of low-quality processes, resulting in insufficient accuracy and representativeness of the consensus standard. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a specimen bioinformatics analysis and verification system and method based on big data, which solves the problems mentioned in the background.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for bioinformatics analysis and verification of specimens based on big data, comprising the following steps:
[0007] Step 1: Multi-source data collection:
[0008] The system acquires gene variation detection results from multiple gene variation analysis processes for multiple samples. For each feature in each sample, i.e. the detected object of the gene variation analysis process in the sample, it generates an indicator value representing whether the gene variation analysis process has given a detection value. At the same time, it assigns a confidence weight coefficient to each gene variation analysis process.
[0009] Step 2: Construction of the dynamic consensus benchmark verification set:
[0010] For each feature of each specimen, the weighted consensus frequency of that feature in that specimen is calculated based on the indicator value and the confidence weight coefficient. The weighted consensus frequency is the ratio of the weighted sum of the indicator value and the confidence weight coefficient of that feature in all gene variation analysis processes to the sum of the confidence weight coefficients of all gene variation analysis processes. At the same time, the weighted consensus frequency values of all features of all specimens are constructed into a set, and the upper quartile of this set is taken as the adaptive screening threshold. For each specimen, features with a weighted consensus frequency greater than or equal to the adaptive screening threshold are selected to form the dynamic consensus benchmark feature set of that specimen.
[0011] Step 3: Calculation of the reproducibility risk index:
[0012] Using the dynamic consensus benchmark feature set, the gene variation analysis process to be validated is benchmarked, the number of true positives, false positives and false negatives are counted, and the Jaccard similarity coefficient is calculated as a comprehensive metric to measure the degree of conformity between the gene variation analysis process to be validated and the dynamic consensus benchmark features.
[0013] Step 4: Dynamic Consensus Benchmark Update
[0014] When any of the following triggering conditions are met, the weighted consensus frequency is recalculated, and the adaptive screening threshold and the corresponding dynamic consensus benchmark feature set are adjusted based on the recalculated weighted consensus frequency.
[0015] The triggering conditions include: the addition of a new gene variation analysis process; an update of the weights of an existing gene variation analysis process; reaching a preset update cycle; or the cumulative number of newly acquired samples reaching a preset threshold.
[0016] As a further aspect of the present invention: the upper quartile refers to the weighted consensus frequency value located at the 75th percentile after arranging all weighted consensus frequency values in the set in ascending order;
[0017] If the upper quartile position is not an integer, then take the two weighted consensus frequency values before and after that position and perform linear interpolation, and use the interpolation result as the upper quartile.
[0018] As a further aspect of the present invention: the number of true positives is the number of features contained in the intersection of the detection feature set of the gene variation analysis process to be verified and the dynamic consensus benchmark feature set;
[0019] The false positive count is the number of features in the feature set detected by the gene variation analysis process to be validated that do not belong to the dynamic consensus benchmark feature set;
[0020] The false negative count is the number of features in the dynamic consensus benchmark feature set that were not detected by the gene variation analysis process to be verified.
[0021] As a further aspect of the present invention: the Jaccard similarity coefficient is obtained by dividing the number of true positives by the sum of the number of true positives, false positives, and false negatives;
[0022] As a further aspect of the present invention, the formula for calculating the Jaccard similarity coefficient is as follows:
[0023]
[0024] In the formula, J k,p TP is the Jaccard similarity coefficient. k,p The number of true positives, FP k,p The number of false positives, FN k,p The number of false negatives FN k,p ;
[0025] As a further aspect of the present invention: calculate the arithmetic mean of the Jaccard similarity coefficients of the gene variation analysis process to be verified on all specimens, and subtract the arithmetic mean from 1 to obtain the reproducibility risk index of the gene variation analysis process to be verified.
[0026] As a further aspect of this invention: when a new gene variation analysis process is added, its initial confidence weight coefficient... Calculated using the following formula: ,in, The overall reproducibility risk index for the newly added process.
[0027] As a further aspect of the present invention: the weight update of the existing gene variation analysis process is performed using the following formula: ;
[0028] In the formula, The updated credibility weight coefficients, The trusted weight coefficients before the update are given, and S represents the relationship between this process and the current dynamic consensus base. Let λ be the average Jaccard similarity coefficient of gene variation analysis process i across all specimens relative to the current dynamic consensus benchmark feature, where λ is a smoothing parameter and 0 < λ < 1;
[0029] As a further aspect of the present invention: the feature is a variant site on the genome, including any one or more of single nucleotide variants, small fragment insertions or deletions, structural variants, or copy number variants.
[0030] A big data-based specimen bioinformatics analysis and validation system, which is used to execute a big data-based specimen bioinformatics analysis and validation method, includes:
[0031] The data acquisition module is used to acquire the gene variation detection results of multiple samples from multiple gene variation analysis processes, and generate an indicator value for each feature in each sample for each gene variation analysis process.
[0032] The weighting module is used to assign a reliable weight coefficient to each gene variation analysis process based on its historical performance.
[0033] The consensus frequency calculation module is used to calculate the weighted consensus frequency of each feature on each sample;
[0034] The threshold determination module is used to aggregate the weighted consensus frequency of all features of all specimens and take the upper quartile of the sorted values as the adaptive screening threshold.
[0035] The dynamic consensus benchmark construction module is used to form a dynamic consensus benchmark feature set by combining features in each sample whose weighted consensus frequency is not lower than the adaptive screening threshold.
[0036] The testing and evaluation module is used to compare the gene variation analysis process to be validated using a dynamic consensus benchmark feature set, count the number of true positives, false positives, and false negatives, and calculate the Jaccard similarity coefficient as a measure of consistency.
[0037] As a further aspect of the present invention: the test evaluation module is also used to calculate a reproducibility risk index based on the Jaccard similarity coefficient, wherein the reproducibility risk index is defined as 1 minus the Jaccard similarity coefficient, in order to quantify the non-reproducibility risk of the gene variation analysis process to be verified.
[0038] As a further aspect of the present invention, the system also includes a dynamic update module, which is used to update the confidence weight coefficients of each gene variation analysis process when the triggering conditions are met, and to trigger the recalculation of the consensus frequency and the automatic update of the adaptive screening threshold and the dynamic consensus benchmark feature set.
[0039] This invention provides a specimen bioinformatics analysis and verification system and method based on big data. Compared with existing technologies, it has the following advantages:
[0040] By collecting the actual outputs of multiple gene variation analysis processes and constructing an internal dynamic benchmark using weighted consensus frequency, the traditional method is completely free from its reliance on manually labeled dynamic consensus benchmark features, significantly reducing the human and time costs of bioinformatics analysis process verification.
[0041] Using the upper quartile of the weighted consensus frequency set of all features across all samples as the screening threshold can automatically adjust with data distribution, avoiding the problem of poor adaptability of fixed thresholds across different projects or batches, and ensuring that features included in the dynamic consensus benchmark feature set always have a high overall consensus.
[0042] By using the Jaccard similarity coefficient to measure the consistency between the process to be verified and the dynamic benchmark, subjective evaluation is transformed into objective numerical values, which facilitates horizontal comparison and iterative optimization between different processes.
[0043] When a new process is added or the weight of an existing process changes, the system automatically recalculates the weighted consensus frequency, adjusts the adaptive threshold, and updates the benchmark set, ensuring that the verification criteria are always synchronized with the latest state of the current process group and avoiding outdated benchmarks. Attached Figure Description
[0044] Figure 1 This is a system block diagram of a specimen bioinformatics analysis and verification system based on big data according to the present invention.
[0045] Figure 2 This is a flowchart illustrating a big data-based bioinformatics analysis and verification method for specimens according to the present invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Please see Figure 1 and Figure 2 As shown, the embodiments of the present invention provide the following technical solutions:
[0048] As an embodiment of the present invention:
[0049] This invention relates to a method for bioinformatics analysis and verification of specimens based on big data, comprising:
[0050] Step 1: Multi-source data collection:
[0051] The system acquires gene variation detection results from multiple gene variation analysis processes for multiple samples. For each feature in each sample, it generates an indicator value representing whether the gene variation analysis process has detected it.
[0052] We collected bioinformatics analysis results of the same set of samples from gene variation analysis workflows distributed across multiple laboratories, using different sequencing platforms, different versions of analysis software, and different parameter combinations.
[0053] Each specimen is labeled with an index k, where k = 1, 2, ..., K, and K is the total number of specimens; a specimen is a biological sample.
[0054] For each specimen k, select N known gene variation analysis procedures, and label each gene variation analysis procedure with index i, i=1,2,…,N;
[0055] For any gene variation analysis workflow i, the gene variation detection result output for sample k is identified by feature number j, where j=1,2,…,M. k M k The total number of unique features after deduplication and merging, as reported by the analysis procedure for all gene variations on specimen K.
[0056] For each specimen k, the features that have been detected at least once in all gene variation analysis procedures i are summarized as the complete set of candidate features for that specimen;
[0057] Among them, the characteristic is a variant site on the genome, including any one or more of single nucleotide variants, small fragment insertions and deletions, structural variants or copy number variants;
[0058] For each specimen, each gene variation analysis procedure, and each feature, a corresponding indicator value d is constructed. k,i,j Indicator value d k,i,j Used to record whether feature j was detected in specimen k in gene variation analysis process i;
[0059] If gene variation analysis procedure i detects feature j in specimen K, then d k,i,j =1;
[0060] If gene variation analysis procedure i does not detect feature j in specimen K, then d k,i,j =0;
[0061] All indicator values d k,i,j This is transformed into an indicator value matrix, forming the foundational data for constructing dynamic consensus benchmark features;
[0062] Before constructing the indicator value matrix, the raw variant detection results output from different gene variant analysis workflows are standardized to ensure that the same biological variant is correctly identified as the same feature across different workflows. Specifically, this includes:
[0063] The variation detection results output from each gene variation analysis workflow are unified to the same reference genome version, and a genome coordinate transformation tool is used to perform coordinate mapping between versions;
[0064] For single nucleotide variants and small fragment insertions and deletions, a left-aligned and simplified algorithm is used to standardize the allele representation, eliminating the differences in coordinates and representations caused by inconsistencies in reference genome versions or allele writing between processes;
[0065] For each standardized variant site, a globally unique feature identifier is generated based on its chromosome, start position, and standardized reference and alternative allele sequences;
[0066] All variants detected on specimen k by the complete gene variation analysis process are processed as described above, and then deduplicated and merged according to the feature identifier to form the complete set of candidate features for the specimen.
[0067] At the same time, a confidence weight coefficient w is set for each gene variation analysis process i. i Confidence weight coefficient w i This indicates the reliability of the historical performance of the gene variation analysis process;
[0068] The confidence weight coefficients for all gene variation analysis procedures are set to w. i =1, meaning that all gene variation analysis processes are given the same initial level of confidence;
[0069] Step 2: Construction of the dynamic consensus benchmark verification set:
[0070] pass: Calculate the weighted consensus frequency F corresponding to feature j in specimen k. k,j ;
[0071] In the formula:
[0072] The molecule is the detection indicator value d for all gene variation analysis procedures. k,i,j With weight w i The result of weighted summation;
[0073] The denominator is the sum of the weights of all N gene variation analysis procedures;
[0074] F k,j The value range of is limited to between 0 and 1;
[0075] When the weight of all gene variation analysis procedures is 1, F k,jThis is equivalent to the proportion of gene variation analysis processes that detected this feature out of the total number of gene variation analysis processes;
[0076] F to obtain all characteristics of all specimens k,j Then, take all F. k,j The upper quartile of the value is used as the adaptive screening threshold θ, which is used to screen out features that can be included in the dynamic consensus benchmark feature set.
[0077] Among them, take all F k,j The upper quartile of the value refers to all weighted consensus frequencies F. k,j The value located at the 75th percentile after sorting from smallest to largest;
[0078] If the obtained upper quartile position is not an integer, then take the two adjacent sort values before and after the position and perform linear interpolation, and use the interpolation result as the upper quartile.
[0079] As the number of gene variation analysis processes increases and / or the confidence weight coefficients are updated, the weighted consensus frequency F... k,j The distribution changes simultaneously, causing the adaptive screening threshold θ to be updated accordingly;
[0080] And the dynamic consensus benchmark feature set constructed from sample K is represented as In the dynamic consensus benchmark feature set middle, ;
[0081] The consensus benchmark is determined dynamically based entirely on the data distribution by using a weighted consensus frequency and an upper quartile adaptive threshold, avoiding subjective biases caused by preset fixed standards. The threshold can be automatically adjusted as the number of processes and weights change, ensuring that the benchmark set always reflects the high confidence characteristics of the current data set.
[0082] Step 3: Calculation of the reproducibility risk index:
[0083] When a new gene variation analysis workflow p needs to be evaluated, its detection results for all K samples are obtained, and a corresponding indicator value d is constructed. k,p,j And form a set of detection features of gene variation analysis p on specimen K. In the detected feature set middle, ;
[0084] Compare the detected feature set Dk,p with the current dynamic consensus benchmark feature set G of the same specimen. k Perform a step-by-step comparison and count the following three basic counts:
[0085] True positive count (TP) k,p: Indicates that it is simultaneously detected by the gene variation analysis process p and belongs to the dynamic consensus benchmark feature set G. k The number of features;
[0086] The true positive value equals the detected feature set D. k,p With dynamic consensus benchmark feature set G k The number of common elements in the intersection of two sets, that is, the number of features contained in the intersection of the two sets;
[0087] False positives (FP) k,p : Indicates that the gene variation was detected by the gene variation analysis process p but does not belong to the dynamic consensus benchmark feature set G. k The number of features;
[0088] The false positive value is equal to the value in the detected feature set D. k,p In but not in the dynamic consensus benchmark feature set G k The number of features in;
[0089] False negative number FN k,p : Indicates that it belongs to the dynamic consensus benchmark feature set G k However, the number of features missed by the gene variation analysis process p;
[0090] The false positive value is equal to the value in the dynamic consensus benchmark feature set G. k In but not in the detection feature set D k,p The number of features in.
[0091] Through true positive number TP k,p False positives (FP) k,p False negative number FN k,p To determine the degree of agreement between the gene variation analysis process p and the dynamic consensus benchmark characteristics on specimen K;
[0092] Using the Jaccard similarity coefficient J k,p As a comprehensive metric, its formula is as follows:
[0093]
[0094] Among them, the Jaccard similarity coefficient J k,p The closer the value is to 1, the smaller the difference between the results of the gene variation analysis process p on specimen K and the dynamic consensus benchmark characteristics;
[0095] When the dynamic consensus benchmark feature set is empty and the gene variation analysis process p also fails to detect any features, then J k,p =1;
[0096] When the consensus benchmark set is empty, but process p detects a feature, then J... k,p =0.
[0097] The evaluation of single specimens is aggregated into a comprehensive risk index at the level of gene variation analysis process;
[0098] Define the reproducibility risk index RI of the gene variation analysis process p to be validated. p The arithmetic mean of 1 minus the Jaccard similarity coefficients for all specimens is given by the following formula:
[0099]
[0100] Among them, RI p The theoretical range of the reproducibility risk index RI is 0 to 1; p The lower the value, the smaller the average difference between the output of the gene variation analysis process and the dynamic consensus benchmark characteristics condensed from multi-source big data, indicating higher reproducibility across platforms and research laboratories, and lower risk; conversely, a higher reproducibility risk index (RI) indicates a lower risk. p A significantly elevated value indicates a substantial risk of inconsistency in the output of this gene variation analysis process.
[0101] Based on the dynamic consensus benchmark, true positives, false positives, and false negatives are statistically analyzed. By using the Jaccard similarity coefficient and the reproducibility risk index, the overall consistency between the process and the multi-source consensus is condensed into a single indicator between 0 and 1. This can intuitively measure the reproducibility across platforms and laboratories, providing a clear basis for the selection and optimization of gene variation analysis processes.
[0102] As a second embodiment of the present invention:
[0103] In its specific implementation, compared to Embodiment 1, the technical solution of this embodiment differs from that of Embodiment 1 only in that this embodiment further includes the step of dynamically updating the dynamic consensus benchmark, as follows:
[0104] The gene variation analysis workflow p is incorporated into a pre-established gene variation analysis workflow pool, and its performance feedback is used to correct the confidence weight coefficients of each gene variation analysis workflow, as follows:
[0105] pass: Calculate the gene variation analysis process p and assign appropriate initial weights. ;
[0106] Among them, the low-risk gene variant analysis process received a relatively high weight close to 1, while the high-risk gene variant analysis process had a lower initial weight.
[0107] This method also supports an automated periodic update mechanism, which is to automatically trigger a re-evaluation of all processes in the gene variation analysis process pool based on a preset time period or when the number of newly acquired samples accumulates to a preset threshold.
[0108] Specifically as follows:
[0109] For any gene variation analysis procedure i, the recent performance score S of gene variation analysis procedure i is determined by calculating the average Jaccard similarity coefficient of gene variation analysis procedure i relative to the current dynamic consensus benchmark feature across all specimens. i The formula is as follows:
[0110]
[0111] In the formula, TP k,i FP k,i 、FN k,i It is based on the latest dynamic consensus benchmark feature G k Based on the current output of the gene variation analysis workflow, the number of true positives, false positives, and false negatives are recalculated.
[0112] The confidence weight coefficients for each gene variation analysis procedure are updated using the exponential smoothing method, as shown in the following formula:
[0113]
[0114] In the formula: λ is a smoothing parameter, which takes a value between 0 and 1. In this embodiment, its value is 0.9; w1 i The updated confidence weight coefficients for gene variation analysis workflow i;
[0115] After the trusted weight coefficients are updated, the system automatically returns to the dynamic consensus benchmark verification set construction step and utilizes all updated w1 values. i Recalculate the weighted consensus frequency F for each feature k,j Simultaneously, based on the new frequency distribution, the adaptive screening threshold θ is redefined, and the dynamic consensus benchmark feature set G for each specimen is reconstructed. k .
[0116] By incorporating the newly evaluated gene variation analysis process into the weight update system and making feedback corrections based on the average Jaccard similarity coefficient between each gene variation analysis process and the benchmark in recent times, the credible weights can dynamically reflect the true performance of the process, drive the iterative evolution of the consensus benchmark, and continuously improve accuracy over long-term operation.
[0117] As an embodiment of the present invention:
[0118] In its specific implementation, compared to Embodiment 1 and Embodiment 2, the technical solution of this embodiment is to combine the solutions of Embodiment 1 and Embodiment 2. Furthermore, this invention also provides a specimen bioinformatics analysis and verification system based on big data. This system is used to execute a specimen bioinformatics analysis and verification method based on big data. The system includes:
[0119] The data acquisition module is used to acquire the gene variation detection results of multiple samples from multiple gene variation analysis processes, and generate an indicator value for each feature in each sample for each gene variation analysis process.
[0120] The weighting module is used to assign a reliable weight coefficient to each gene variation analysis process based on its historical performance.
[0121] The consensus frequency calculation module is used to calculate the weighted consensus frequency of each feature on each sample;
[0122] The threshold determination module is used to aggregate the weighted consensus frequency of all features of all specimens and take the upper quartile of the sorted values as the adaptive screening threshold.
[0123] The dynamic consensus benchmark construction module is used to form a dynamic consensus benchmark feature set by assembling features in each sample whose weighted consensus frequency is not lower than the adaptive screening threshold.
[0124] The testing and evaluation module is used to compare the gene variant analysis process to be verified using a dynamic consensus benchmark feature set, count the number of true positives, false positives, and false negatives, and calculate the Jaccard similarity coefficient as a consistency measure; it is also used to calculate a reproducibility risk index based on the Jaccard similarity coefficient, the reproducibility risk index being defined as 1 minus the Jaccard similarity coefficient, to quantify the non-reproducibility risk of the gene variant analysis process to be verified.
[0125] The dynamic update module is used to update the confidence weight coefficients of each gene variation analysis process and trigger the recalculation of consensus frequency and automatic update of adaptive screening threshold and dynamic consensus benchmark feature set when any of the following triggering conditions are met: a new gene variation analysis process is added; the weights of existing gene variation analysis processes are updated; a preset update cycle is reached; or the number of newly acquired samples accumulates to a preset threshold.
[0126] The specific functions of each module correspond one-to-one with the method steps in Embodiment 1 and Embodiment 2, and will not be repeated here.
[0127] This invention constructs a dynamic consensus benchmark by using a weighted consensus frequency and an adaptive threshold, which can objectively quantify and analyze the reproducibility risk of the process. It also introduces a weighted feedback update mechanism based on actual performance, which enables the benchmark to continuously self-optimize as data accumulates, significantly improving the accuracy, adaptability, and long-term operational stability of cross-platform process verification.
[0128] It should be stated that all user data collected in this application was collected with the user's consent and authorization, and the use of user data is legal and compliant, and the use and processing of user data comply with the relevant laws, regulations and standards of the relevant regions.
[0129] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0130] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0131] All the above formulas use dimensionless numerical calculations, that is, by introducing a reference benchmark, the original physical quantities are transformed into dimensionless relative values, eliminating the influence of units and retaining the relative size relationship of physical quantities; the formula is the closest to the real situation obtained by software simulation of a large amount of data, and the preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.
[0132] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0133] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for bioinformatics analysis and verification of specimens based on big data, characterized in that, Includes the following steps: Step 1: Multi-source data collection: The system acquires gene variation detection results from multiple gene variation analysis processes for multiple samples. For each feature in each sample, i.e. the detected object of the gene variation analysis process in the sample, it generates an indicator value representing whether the gene variation analysis process has given a detection value. At the same time, it assigns a confidence weight coefficient to each gene variation analysis process. Step 2: Construction of the dynamic consensus benchmark verification set: For each feature of each specimen, the weighted consensus frequency of that feature in that specimen is calculated based on the indicator value and the confidence weight coefficient. The weighted consensus frequency is the ratio of the weighted sum of the indicator value and the confidence weight coefficient of that feature in all gene variation analysis processes to the sum of the confidence weight coefficients of all gene variation analysis processes. Simultaneously, the weighted consensus frequency values of all features of all specimens are constructed into a set, and the upper quartile of this set is taken as the adaptive screening threshold; for each specimen, features with a weighted consensus frequency greater than or equal to the adaptive screening threshold are selected to form the dynamic consensus benchmark feature set of the specimen. Step 3: Calculation of the reproducibility risk index: Using a dynamic consensus benchmark feature set, the gene variation analysis process to be validated is benchmarked, and the number of true positives, false positives, and false negatives are counted. Based on this, the Jaccard similarity coefficient is calculated as a comprehensive metric to measure the degree of conformity between the gene variation analysis process to be validated and the dynamic consensus benchmark features.
2. The method for bioinformatics analysis and verification of specimens based on big data according to claim 1, characterized in that: If the upper quartile position is not an integer, then take the two weighted consensus frequency values before and after that position and perform linear interpolation, and use the interpolation result as the upper quartile.
3. The method for specimen bioinformatics analysis and verification based on big data according to claim 1, characterized in that: The true positive count is the number of features contained in the intersection of the detection feature set of the gene variant analysis process to be validated and the dynamic consensus benchmark feature set; The false positive count is the number of features in the feature set detected by the gene variation analysis process to be validated that do not belong to the dynamic consensus benchmark feature set; The false negative count is the number of features in the dynamic consensus benchmark feature set that were not detected by the gene variation analysis process to be verified.
4. The method for specimen bioinformatics analysis and verification based on big data according to claim 1, characterized in that: The Jaccard similarity coefficient is derived by dividing the number of true positives by the sum of the number of true positives, false positives, and false negatives.
5. The method for specimen bioinformatics analysis and verification based on big data according to claim 1, characterized in that: When any of the following triggering conditions are met, the weighted consensus frequency is recalculated, and the adaptive screening threshold and the corresponding dynamic consensus benchmark feature set are adjusted based on the recalculated weighted consensus frequency. The triggering conditions include: the addition of a new gene variation analysis process; an update of the weights of an existing gene variation analysis process; reaching a preset update cycle; or the cumulative number of newly acquired samples reaching a preset threshold.
6. The method for bioinformatics analysis and verification of specimens based on big data according to claim 1, characterized in that: The weight update of the existing gene variation analysis process is based on the average Jaccard similarity coefficient between each gene variation analysis process and the current dynamic consensus benchmark feature set, and the credible weight coefficient is updated using the exponential smoothing method.
7. The method for specimen bioinformatics analysis and verification based on big data according to claim 1, characterized in that: The feature is a variant site on the genome, including any one or more of single nucleotide variants, small fragment insertions or deletions, structural variants, or copy number variants.
8. A specimen bioinformatics analysis and verification system based on big data, the system being used to execute the specimen bioinformatics analysis and verification method based on big data as described in any one of claims 1-7, characterized in that, The system includes: The data acquisition module is used to acquire the gene variation detection results of multiple samples from multiple gene variation analysis processes, and generate an indicator value for each feature in each sample for each gene variation analysis process. The weighting module is used to assign a reliable weight coefficient to each gene variation analysis process based on its historical performance. The consensus frequency calculation module is used to calculate the weighted consensus frequency of each feature on each sample; The threshold determination module is used to aggregate the weighted consensus frequency of all features of all specimens and take the upper quartile of the sorted values as the adaptive screening threshold. The dynamic consensus benchmark construction module is used to form a dynamic consensus benchmark feature set by combining features in each sample whose weighted consensus frequency is not lower than the adaptive screening threshold. The testing and evaluation module is used to compare the gene variation analysis process to be validated using a dynamic consensus benchmark feature set, count the number of true positives, false positives, and false negatives, and calculate the Jaccard similarity coefficient as a measure of consistency.
9. The specimen bioinformatics analysis and verification system based on big data according to claim 8, characterized in that: The test evaluation module is also used to calculate a reproducibility risk index based on the Jaccard similarity coefficient. The reproducibility risk index is defined as 1 minus the Jaccard similarity coefficient, which is used to quantify the non-reproducibility risk of the gene variation analysis process to be verified.
10. The specimen bioinformatics analysis and verification system based on big data according to claim 8, characterized in that: It also includes a dynamic update module, which is used to update the confidence weight coefficients of each gene variation analysis process when the triggering conditions are met, and to trigger the recalculation of the consensus frequency and the automatic update of the adaptive screening threshold and the dynamic consensus benchmark feature set.