Processing method of auxiliary diagnosis data, electronic device, medium and product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0008]本申请的一个目的是提供辅助诊断数据的处理方法、电子设备、介质及产品,至少用以解决现有技术中单个变异评分可能无法充分反映多个变异共同作用对样本级癌症辅助分析结果的影响,也不利于追溯
[0192]本实施例所述辅助诊断数据的处理方法通过对肿瘤体细胞变异位点构建参考—突变序列窗口并进行数值化特征编码,同时结合同基因关系、基因组距离关系和功能区域关系构建跨变异窗口关联矩阵,可实现对同一样本内不同变异位点之间关联性的联合建模,使处理结果不仅能够表征单一变异位点的局部影响,还能够反映多个变异位点之间的协同作用关系,从而提高对肿瘤体细胞变异生物学效应的表征能力;进一步地,通过基于关联强度对多个变异窗口的影响向量进行融合,避免仅对单个变异独立评分,增强癌症相关变异特征的表达能力,降低低可信变异窗口对分析结果的干扰,提高辅助诊断数据的准确性、稳定性及可靠性。此外,通过生成变异级评分、基因级评分及样本级辅助诊断评分,实现了样本级评分的可追溯性,可分别定位至具体的候选变异与候选基因。
Smart Images

Figure CN122551894A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence-assisted medical analysis technology, and in particular to methods, electronic devices, media and products for processing auxiliary diagnostic data. Background Technology
[0002] With the development of high-throughput sequencing technology, tumor genomic data has been widely used for cancer-aided analysis, candidate variant screening, tumor risk stratification, and prognostic assessment. Current tumor gene sequencing data processing typically includes steps such as data quality control, reference genome alignment, somatic variant detection, and variant annotation.
[0003] Existing methods for processing tumor gene sequence data typically treat each variant site as an independent object. For example, traditional methods may only determine whether a particular variant is located in a known cancer-related gene, or may only use information such as the variant's chromosomal coordinates, reference bases, mutated bases, sequencing depth, and variant frequency for scoring.
[0004] However, tumor samples typically contain multiple somatic variants. These variants may be located in the same gene, in adjacent regions of the same chromosome, or belong to the same functional region or regulatory relationship. If each variant is scored independently, it is difficult to reflect the combined effects of multiple variants within the same sample.
[0005] Therefore, the main shortcoming of the existing technology is that it is difficult to transform the correlation between multiple somatic cell variations within the same sample into a computable data structure and to use this correlation to fuse the effects of multiple variations.
[0006] The resulting technical problem is that a single variant score may not be able to fully reflect the combined effect of multiple variants on the results of sample-level cancer auxiliary analysis, and it is also not conducive to retrospection.
[0007] Therefore, how to provide methods, electronic devices, media, and products for processing auxiliary diagnostic data to overcome the obvious deficiencies of existing technologies has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this application is to provide methods, electronic devices, media, and products for processing auxiliary diagnostic data, at least to address the issue that in the prior art, a single variant score may not adequately reflect the combined effect of multiple variants on the results of sample-level cancer auxiliary analysis, and is also not conducive to retrospective analysis.
[0009] To achieve the above objectives, some embodiments of this application provide the following aspects:
[0010] In a first aspect, some embodiments of this application provide a method for processing auxiliary diagnostic data. The method includes: receiving gene sequencing data of a tumor sample to be analyzed and generating somatic mutation site information; constructing a reference-mutation sequence window pair for the somatic mutation site information; numerically processing the reference sequence window and the mutation sequence window in the reference-mutation sequence window pair to generate mutation window input features, and sequentially encoding the mutation window input features to obtain a mutation window vector; the mutation window input features include reference window features and mutation window features, and the mutation window vector includes a reference window vector and a mutation window vector; generating an original mutation influence vector based on the difference between the mutation window vector and the reference window vector, and concatenating the original mutation influence vector with window auxiliary features in the mutation window input features to form a set of mutation influence vectors; constructing a cross-mutation window association matrix based on the somatic mutation site information, and weighted fusing the set of mutation influence vectors based on the cross-mutation window association matrix to form a fused set of mutation influence vectors; the cross-mutation window association matrix is used to represent the association strength between two mutation windows.
[0011] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the auxiliary diagnostic data processing method described above.
[0012] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the auxiliary diagnostic data processing method described above.
[0013] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the auxiliary diagnostic data processing method described above.
[0014] Compared with related technologies, this application constructs a reference-mutation sequence window for tumor somatic mutation sites and performs numerical feature encoding. Simultaneously, it constructs a cross-mutation window association matrix by combining syngeneic relationships, genomic distance relationships, and functional region relationships. This enables joint modeling of the associations between different mutation sites within the same sample, allowing the processing results to characterize not only the local impact of a single mutation site but also the synergistic relationships between multiple mutation sites, thereby improving the characterization ability of the biological effects of tumor somatic mutations. Furthermore, by fusing the influence vectors of multiple mutation windows based on association strength, it avoids scoring individual mutations independently, enhances the expression of cancer-related mutation features, reduces the interference of low-confidence mutation windows on the analysis results, and improves the accuracy, stability, and reliability of auxiliary diagnostic data. In addition, by generating mutation-level scores, gene-level scores, and sample-level auxiliary diagnostic scores, it achieves traceability of sample-level scores, allowing for the identification of specific candidate mutations and candidate genes. Attached Figure Description
[0015] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0016] Figure 1 An exemplary flowchart of a method for processing auxiliary diagnostic data provided in some embodiments;
[0017] Figure 2 An exemplary flowchart of another method for processing auxiliary diagnostic data provided in some embodiments;
[0018] Figure 3 An exemplary structural diagram of an electronic device is provided for some embodiments. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] In this embodiment of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good morals.
[0021] First Embodiment
[0022] The first embodiment relates to a method for processing auxiliary diagnostic data. For example... Figure 1 As shown, the processing method may include the following steps:
[0023] S11 receives gene sequencing data from the tumor sample to be analyzed and generates somatic mutation site information.
[0024] S12, construct a reference-mutation sequence window pair for the somatic cell mutation site information.
[0025] S13, the reference sequence window and the mutation sequence window in the reference-mutation sequence window pair are numerically processed to generate mutation window input features, and the mutation window input features are sequence encoded to obtain a mutation window vector; the mutation window input features include reference window features and mutation window features, and the mutation window vector includes reference window vector and mutation window vector.
[0026] S14, generate an original mutation influence vector based on the difference between the mutation window vector and the reference window vector, and concatenate the original mutation influence vector with the window auxiliary features to form a set of mutation influence vectors.
[0027] S15, a cross-variation window association matrix is constructed based on the somatic cell mutation site information, and the set of mutation influence vectors is weighted and fused based on the cross-variation window association matrix to form a fused mutation influence vector set. The cross-variation window association matrix is used to represent the association strength between two mutation windows.
[0028] The core of this embodiment lies in DNA sequencing data for tumor samples. A reference-mutation sequence window is constructed for tumor somatic mutation sites. Each mutation window is numerically encoded, and a cross-mutation window association matrix is built based on syngeneic relationships, genomic distance relationships, and functional region relationships. The association strength between different mutation windows within the same sample is calculated, and the influence vectors of multiple mutation windows within the same sample are fused based on this association strength. By constructing a reference-mutation sequence window for tumor somatic mutation sites and numerically encoding its features, and simultaneously constructing a cross-mutation window association matrix based on syngeneic relationships, genomic distance relationships, and functional region relationships, joint modeling of the associations between different mutation sites within the same sample can be achieved. This allows the processing results to not only characterize the local impact of a single mutation site but also reflect the synergistic relationships between multiple mutation sites, thereby improving the characterization ability of the biological effects of tumor somatic mutations. Furthermore, by fusing the influence vectors of multiple mutation windows based on association strength, the independent scoring of individual mutations is avoided, enhancing the expression ability of cancer-related mutation features, reducing the interference of low-confidence mutation windows on the analysis results, and improving the accuracy, stability, and reliability of auxiliary diagnostic data.
[0029] Second Embodiment
[0030] The second embodiment relates to another method for processing auxiliary diagnostic data. Compared to the first embodiment, the second embodiment further improves and optimizes the data processing flow, variant association analysis method, and auxiliary diagnostic data generation mechanism to enhance the comprehensive characterization ability of tumor somatic cell variants and improve the accuracy and stability of auxiliary diagnostic results. Please refer to the following for details. Figure 2 As shown, the method for processing the auxiliary diagnostic data specifically includes the following steps:
[0031] S21: Receive gene sequencing data from the tumor sample to be analyzed and generate somatic mutation site information. In this embodiment, the gene sequencing data includes FASTQ files, BAM files, VCF files, a reference genome file, normal sample data corresponding to the tumor sample, and sample tags.
[0032] Gene sequencing data refers to DNA fragment data obtained through sequencing equipment, or data files formed after comparison and mutation detection.
[0033] FASTQ files refer to a file format that records raw DNA sequencing fragments and their base quality fractions.
[0034] A BAM file is a binary alignment file created after sequencing fragments are aligned to a reference genome. It records the location of each DNA sequence fragment on the reference genome. A read refers to a DNA sequence fragment read by the sequencing device.
[0035] VCF files refer to a file format that records information about variant sites. They typically include information such as chromosome number, variant location, reference base, mutated base, sequencing depth, and variant quality score.
[0036] The reference genome refers to the human genome sequence used as a comparison standard.
[0037] Sample labels include cancer type labels, risk level labels, or prognostic labels.
[0038] As a preferred embodiment, S21 specifically includes:
[0039] When the received gene sequencing data is a FASTQ file, quality control is performed on the DNA sequencing fragments to obtain the quality-controlled sequencing fragments. The quality-controlled sequencing fragments are compared with a reference genome file to generate a BAM file. Somatic cell variations are detected based on the differences between the tumor sample and its normal sample, and somatic cell variation site information is generated. This somatic cell variation site information is expressed in the form of a somatic cell variation site table, which includes field names and field meanings. The specific contents of the somatic cell variation site table can be found in Table 1.
[0040]
[0041] Table 1: Details of Somatic Cell Variation Sites
[0042] Among them, the frequency of variant alleles The percentage of sequencing fragments supporting the mutated base in the total sequencing depth at that site is represented by the following formula:
[0043]
[0044] in, Indicates the first The number of DNA sequence fragments supporting the mutated bases at each mutation site. This indicates the total sequencing depth at that site.
[0045] In this embodiment, step S21 outputs a somatic mutation site table, which serves as the data basis for subsequent construction of a reference—the mutation sequence window—and calculation of cross-mutation window associations.
[0046] S22, construct a reference-mutation sequence window pair for the somatic cell mutation site information.
[0047] To achieve standardized representation and batch processing of various somatic cell variations, and to make fuller use of upstream and downstream base sequence information of the variations, S22 specifically includes:
[0048] S221, at each somatic cell mutation site, with the mutation site as the center, L bases are cut upstream and L bases are cut downstream from the reference genome to construct the reference sequence window with a length of 2L+1.
[0049] Specifically, in the first Individual cell mutation sites, their chromosome coordinates are Centered on this variant site, samples were extracted upstream from the reference genome. One base, truncation downstream 4 bases, forming a length of 1000 bases Reference sequence window:
[0050]
[0051] in, Indicates the first Reference sequence window for individual cell variation sites. Indicates the bases in the reference genome. Set the preset window radius, such as 128, 256, 512, or 1024.
[0052] S222, based on the reference base in the somatic cell mutation site information and mutant bases In the reference sequence window, the reference bases (reference fragments) are replaced with mutant bases (mutant fragments) to construct the mutant sequence window:
[0053]
[0054] in, (⋅) indicates a substitution operation. In this embodiment, the reference sequence window is used to represent the original local DNA sequence state of the variant site in the reference genome. The mutated sequence window is used to represent the local DNA sequence state after the variant site has undergone mutation.
[0055] S223, determine the mutation type of the somatic cell mutation site information, and process the reference sequence window and the mutation sequence window according to the mutation type to form a reference-mutation sequence window pair. Thus, each mutation site forms a reference-mutation sequence window pair:
[0056]
[0057] As a preferred embodiment, the method for determining the mutation type of somatic cell mutation site information is as follows: the mutation type is determined based on the length relationship between the reference base and the mutated base.
[0058] Specifically, if and If both are 1 in length and are different, then it is a single nucleotide variant;
[0059] like Length greater than If the length is specified, then it is an insertion mutation;
[0060] like Length greater than Length indicates a missing variant;
[0061] like and If all nucleotides are greater than 1 and have the same length, then it is a polynucleotide variation.
[0062] The steps for processing the reference sequence window and the mutated sequence window according to the mutation type include: if the inserted fragment is too long, truncating the portion exceeding the window length; if a missing position occurs, using... <del>Mark; if the window length is insufficient, use <pad>Mark completion; if an unknown base exists in the window, use... <n>mark.
[0063] S224, All reference-mutation sequence window pairs corresponding to somatic variations are grouped into a reference-mutation sequence window pair set. The reference-mutation sequence window pair set P is represented as:
[0064]
[0065] Where n is the number of somatic cell variants in the sample after quality filtering.
[0066] S23, the reference sequence window and the mutation sequence window in the reference-mutation sequence window alignment are numerically processed to generate mutation window input features, and the mutation window input features are sequence encoded to obtain a mutation window vector. The window input feature matrix includes reference window encoding sequence, mutation window encoding sequence, reference-mutation differential marker, mutation relative position encoding, mutation type encoding, strand direction encoding, sequencing depth encoding, mutation allele frequency encoding, mutation quality fraction encoding, and / or region type encoding.
[0067] To make fuller use of the upstream and downstream base sequence information of the variation, step S23 specifically includes the following steps:
[0068] S231 maps DNA bases and special markers to integer codes, see Table 2 for details.
[0069]
[0070] Table 2: Integer Coding Table for DNA Bases and Special Markers
[0071] S232 encodes the bases at each position in the reference sequence window and the mutant sequence window, respectively, to form the reference sequence window coding sequence and the mutant sequence window coding sequence. The reference sequence window coding sequence is represented as follows: The mutation sequence window encoding sequence is represented as , Indicates the first The first reference window The encoded value at each position, Indicates the first The mutation window of the _th mutation window The encoded value of each position.
[0072] S233, compare the coding values at the same position of the reference sequence window coding sequence and the mutation sequence window coding sequence bit by bit, and generate the reference-mutation difference marker sequence of the mutation window based on the comparison result.
[0073]
[0074]
[0075] in, Indicates the first In the mutation window, the _ ... Whether a reference-mutation difference occurs at a given location is a single-point marker value. Indicates the first The entire differential marker sequence of each mutation window, consisting of all positions. composition.
[0076] S234, the relative position encoding sequence of each mutation window is composed of the relative position encoding values of each position.
[0077] Specifically, the relative position encoding sequence of the i-th mutation window ,
[0078] This represents the distance encoding of the j-th position in the i-th mutation window relative to the central mutation site. .
[0079] S235 encodes the variant type, chain direction, and region type in somatic cell variant site information, forming variant type code, chain direction code, and region type code.
[0080] Mutation type encoding The encoding rules for representing mutation types are shown in Table 3:
[0081]
[0082] Table 3: Coding Rules for Variance Types
[0083] Chain direction encoding The encoding rules for indicating the direction of the mutation chain are shown in Table 4.
[0084]
[0085] Table 4: Chain Direction Encoding Rules
[0086] Region type encoding This is used to indicate the functional region where the mutation occurs. This functional region can be represented by the somatic mutation site table. Field obtained. When this field is missing, it is obtained through variant coordinates. The coding rules for obtaining the regions are shown in Table 5, which are obtained by overlapping matching with genome annotation regions.
[0087]
[0088] Table 5: Regional Type Coding Rules
[0089] S236, Based on the sequencing depth of the somatic cell mutation site information, obtain the sequencing depth code, which is represented as:
[0090] in, To predetermine the upper limit of depth normalization. S237, based on the site sequencing depth and the number of sequencing fragments supporting mutated bases in the somatic cell mutation site information, obtain the variant allele frequency code, which is expressed as:
[0091]
[0092] S238, Based on the variation quality score in the somatic cell variation site information, obtain the variation quality score code, which is represented as:
[0093]
[0094] in, This sets the upper limit for normalizing the quality score.
[0095] S239, the reference window coding sequence, mutation window coding sequence, reference-mutation differential marker sequence, relative position coding sequence, variant type coding, strand direction coding, sequencing depth coding, variant allele frequency coding, variant quality fraction coding, and region type coding are concatenated to form the variant window input feature. The reference window feature consists of the reference window coding sequence and common auxiliary features, and the mutation window feature consists of the mutation window coding sequence and common auxiliary features. The variant window input feature is represented as follows:
[0096]
[0097] in, Indicates the reference window encoding sequence, Represents the mutation window coding sequence; Indicates reference-mutation differential marker sequence, Represents a relative position encoded sequence. , , , , , For window-level auxiliary features, These are public auxiliary features.
[0098] S240, use the same sequence encoder to perform sequence encoding on the reference window features and mutation window features (sequence encoding is represented by E()) to obtain the reference window vector and mutation window vector.
[0099] Specifically, the reference window vector is represented as The mutation window vector is represented as
[0100] S24. Generate an original mutation influence vector based on the difference between the mutation window vector and the reference window vector, and concatenate the original mutation influence vector with the window auxiliary features in the mutation window input features to form a set of mutation influence vectors, so that the analysis results no longer depend solely on the mutation coordinates or mutation types.
[0101] In a preferred embodiment, S24 specifically includes the following steps:
[0102] S241, calculate the difference between the reference window vector and the mutation window vector to obtain the original mutation influence vector.
[0103] Specifically, the original variation effect vector is represented as:
[0104] S242, the original mutation effect vector The variant influence vector is formed by concatenating the window auxiliary features in the variant window input features with the variant influence vector input features; the window auxiliary features in the variant window input features include sequencing depth encoding, variant allele frequency encoding, variant quality score encoding, variant type encoding, and region type encoding. ).
[0105] Specifically, the variation effect vector is represented as:
[0106] ,
[0107] This represents the operation of concatenating multiple values or vectors in sequence to form a longer vector.
[0108] S243, combine the mutation impact vectors corresponding to all mutation windows into a mutation impact vector set. The mutation impact vector refers to the vector obtained after numerical processing of the mutation window, which is used to represent the impact of the mutation on the local DNA sequence.
[0109] Specifically, the set of vectors affecting mutations
[0110] S25, Obtain sequencing quality weights Functional area weight and cancer-related weights .
[0111] Sequencing quality weight The formula used to represent the reliability of sequencing data for this variant site is:
[0112]
[0113] in, For preset coefficients, .
[0114] Functional area weight The location of the variant is determined based on its functional region. Specifically, the overlap between the variant coordinates and the genome annotation interval is assessed, and values are assigned according to the rules shown in Table 6.
[0115]
[0116] Table 6: Rules for Assigning Weights to Functional Areas
[0117] When a mutation site matches multiple region types simultaneously, the region type with the highest weight is taken as the functional region weight of that mutation window.
[0118] Cancer-related weights This is used to indicate the relevance of the variant window to the cancer-assisted analysis task. It is determined based on the frequency of the variant region appearing in cancer samples in the training samples and the statistical correlation between the region and the sample label.
[0119] This area appears frequently Calculation formula:
[0120]
[0121] in, This indicates the number of cancer samples in the training dataset that contain the mutated region. This represents the total number of cancer samples in the training samples.
[0122] Statistical correlation score The association weight, representing the degree of correlation between a region and a cancer type or risk label, can be obtained statistically from the training data. For example, it can be calculated based on the difference in the proportion of the region appearing in samples with different labels. The formula for calculating cancer correlation weight is as follows:
[0123]
[0124]
[0125] in, This indicates the number of high-risk samples that contain the variant region. This represents the total number of high-risk samples. This indicates the number of low-risk samples that contain the variant region. This represents the total number of low-risk samples.
[0126] S26, a cross-variation window association matrix is constructed based on the somatic cell mutation site information, and the set of mutation influence vectors is weighted and fused based on the cross-variation window association matrix to form a fused mutation influence vector set. The cross-variation window association matrix is used to represent the association strength between two mutation windows.
[0127] To ensure that the cross-variation window fusion process has a clear data source and computational path, step S26 specifically includes the following steps:
[0128] S261, construct an n×n cross-variation window association matrix, where n represents the number of mutation windows.
[0129] S262 determines the variant window weights by combining sequencing quality weights, functional region weights, and cancer relevance weights, forming a set of window weights. The variant window weights are expressed as follows:
[0130]
[0131] The set of window weights is represented as: .
[0132] S263, determine the associated elements in the variant window association matrix through syngeneic relationships, location distance relationships, and functional region relationships.
[0133] First, determine the genetic relationship.
[0134] If the mutation sites corresponding to the i-th and j-th mutation windows belong to the same gene, then: ,otherwise, .
[0135] Among them, the gene attribution is derived from the somatic cell variation site table. Field.
[0136] Next, determine the location distance relationship.
[0137] If the mutation sites corresponding to the i-th and j-th mutation windows are located on the same chromosome, then the formula for calculating the distance between them is as follows:
[0138]
[0139] in, and These represent the coordinates of the two mutation sites. This indicates the preset distance attenuation coefficient. The closer the distance, the higher the attenuation coefficient. The larger; the farther the distance, The smaller.
[0140] If the two are not on the same chromosome, then:
[0141] Then, determine the relationships between functional areas.
[0142] If the mutation sites corresponding to the i-th and j-th mutation windows are located in the same functional region type, or if the genes they belong to belong to the same predefined functional group, then: ,otherwise, .
[0143] Among them, the functional area type comes from and Functional grouping can be determined by a preset gene grouping table. A preset gene grouping table is a pre-stored set of genes used to record which genes belong to the same type of biological function.
[0144] Based on the above three types of relationships, the elements of the correlation matrix in the mutation window correlation matrix are determined as follows:
[0145]
[0146] in, These are preset weighting coefficients used to control the proportion of the three types of relationships in the association matrix.
[0147] when At that time, set This is to preserve the information of each mutation window itself.
[0148] S264, normalize each row in the cross-variation window association matrix to obtain the normalized association weights.
[0149] Specifically, the formula for calculating the normalized association weight is:
[0150]
[0151] in, Indicates the first The mutation window for the first The contribution ratio of each mutation window.
[0152] S265, determine the fusion vector by normalizing the correlation weight and the mutation window weight, and perform weighted fusion of the fusion vector and the mutation influence vector to obtain the fused mutation influence vector and the set of fused mutation influence vectors.
[0153] Specifically, the formula for calculating the fusion vector is:
[0154]
[0155] The fusion method of the fusion vector and the mutation influence vector is as follows:
[0156]
[0157] in, Indicates the fused first The vector containing the effects of the mutation. The information includes reference-mutation differences for each variant, as well as information about other related variant windows within the same sample.
[0158] The set of mutation effect vectors is represented as
[0159] S27, perform auxiliary diagnostic comprehensive scoring on the fusion variant impact vector, stratify the samples by risk based on the scoring results, perform low-confidence screening on the variant window, and generate a cancer auxiliary diagnostic data processing report.
[0160] In a preferred embodiment, S27 specifically includes the following steps:
[0161] S271, respectively determine the variant-level score, gene-level score and sample-level auxiliary diagnostic score of the fusion variant influence vector.
[0162] Specifically, first, the mutation level score of the fusion mutation impact vector is determined. For the... The fusion-induced mutation impact vector Perform mutation level scoring:
[0163] in, Indicates the first The mutation level score for each mutation window; This represents a preset or trained scoring weight vector; Indicates the bias term; This represents a normalization function that transforms the input value to the range of 0 to 1.
[0164] Secondly, according to the somatic cell variation site table This field aggregates multiple variant scores belonging to the same gene.
[0165] Let the first The set of variations contained in each gene is:
[0166] in, Indicates the first The gene to which the variant belongs. Indicates the first One gene.
[0167] Next, the gene-level score of the influence vector of fusion variants was determined.
[0168] No. Gene-level scoring The calculation is as follows:
[0169]
[0170] in, Indicates the first Gene-level score of each gene; It is a minimal constant that prevents the denominator from being zero.
[0171] Then, sample-level aggregation is performed on all the fused mutation effect vectors to obtain the sample-level fused vector:
[0172]
[0173] in, This represents the sample-level fusion vector.
[0174] Select the highest-scoring individuals from all gene-level scores. We obtain a Top-K gene score set by sorting genes from highest to lowest. Top-K refers to selecting the top K genes after ranking them from highest to lowest score. Results. Calculate the overall statistical characteristics of the sample. The overall statistical characteristics of the sample include at least one of the following: total number of somatic cell variants, number of high-weight variants, average sequencing depth, proportion of low-confidence variants, and distribution of variant numbers in different functional regions.
[0175] Sample-level fusion vector The Top-K gene score set and the overall statistical features of the sample are concatenated to obtain the sample-level input vector: .
[0176] Finally, the sample-level input vectors are scored to determine the sample-level auxiliary diagnostic score. The method for determining the sample-level auxiliary diagnostic score is as follows:
[0177] in, This represents a sample-level auxiliary diagnostic score; This represents a sample-level scoring weight vector obtained from pre-set or training methods. This indicates the bias term.
[0178] S272, based on the screening criteria for low-confidence windows, perform low-confidence screening on the variant windows and sample results. If the proportion of low-confidence variant windows in the sample results exceeds the preset proportion threshold, then mark the sample as a sample that needs to be reviewed.
[0179] Specifically, the filtering criteria for low-confidence windows are:
[0180] 1. Sequencing depth of variant sites ;
[0181] 2. Variation in mass fraction ;
[0182] 3. Frequency of variant alleles ;
[0183] 4. The proportion of unknown base N in the sequence window is higher than that of unknown base N. .
[0184] in, , , and All are preset thresholds. If any of the above screening conditions are met, the corresponding variant window will be marked as a low-confidence window.
[0185] S273, Based on the risk stratification rules, the sample-level auxiliary diagnostic scores are stratified by risk. In this embodiment, the low-risk threshold is set to... The high-risk threshold is set to And satisfy:
[0186]
[0187] The risk stratification rules are as follows:
[0188]
[0189] S274, Generate a cancer auxiliary diagnostic data processing report. The cancer auxiliary diagnostic data processing report includes sample number, sample-level auxiliary diagnostic score, risk stratification, Top-K candidate variants, Top-K candidate genes, cross-variable window association results, low-confidence data markers, and suggested review regions.
[0190] Among them, the Top-K candidate variants are those with the highest variant level scores. Within a mutation window, the Top-K candidate genes are those with the highest gene-level scores. For each gene, the cross-variant window association result is other variant windows that have a higher association weight with the high-scoring variant window.
[0191] Compared with existing technologies, the auxiliary diagnostic data processing method provided in this embodiment has the following advantages:
[0192] The auxiliary diagnostic data processing method described in this embodiment constructs a reference-mutation sequence window for tumor somatic mutation sites and performs numerical feature encoding. Simultaneously, it constructs a cross-mutation window association matrix by combining syngeneic relationships, genomic distance relationships, and functional region relationships. This enables joint modeling of the associations between different mutation sites within the same sample, allowing the processing results to characterize not only the local impact of a single mutation site but also the synergistic effects between multiple mutation sites, thereby improving the characterization ability of the biological effects of tumor somatic mutations. Furthermore, by fusing the influence vectors of multiple mutation windows based on association strength, it avoids scoring individual mutations independently, enhances the expression of cancer-related mutation features, reduces the interference of low-confidence mutation windows on the analysis results, and improves the accuracy, stability, and reliability of the auxiliary diagnostic data. In addition, by generating mutation-level scores, gene-level scores, and sample-level auxiliary diagnostic scores, it achieves traceability of sample-level scores, allowing for the identification of specific candidate mutations and candidate genes.
[0193] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0194] Third Embodiment
[0195] This embodiment provides an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, mainframe computers, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0196] The electronic device includes: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the auxiliary diagnostic data processing method provided in any one or more of the above embodiments.
[0197] Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, an input device 1103, and an output device 1104. The various components are interconnected via a bus or other means (the diagram shows an example of bus connection). The processor 1101 can be used to execute instructions stored in the memory 1102 to control the overall operation of the electronic device. The memory 1102 may include a program storage area and a data storage area, wherein the program storage area stores the operating system and applications required for at least one function; the data storage area stores data created according to the use of the electronic device, etc. The memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as disk storage devices, flash memory devices, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may also include storage resources located remotely to the processor and accessible via a network.
[0198] Input device 1103 can be used to receive input numerical or character information or user operation signals, such as a touch screen, keypad, mouse, trackpad, touchpad, indicator, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include display devices (such as liquid crystal displays, light-emitting diode displays, plasma displays, and optional touch screens), auxiliary lighting devices (such as LEDs), and haptic feedback devices (such as vibration motors), etc.
[0199] To facilitate user interaction, the electronic device may be configured to include a display device (such as an LCD or CRT monitor) and input devices such as a keyboard and pointing devices (e.g., a mouse or touchpad). Feedback can be any form of sensory feedback (e.g., visual feedback, auditory feedback); input may also be received via voice, touch, or other means.
[0200] This application also relates to a computer-readable medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of the auxiliary diagnostic data processing method provided in any one or more of the above embodiments. The computer-readable medium may be a memory included in an electronic device, or it may be a standalone storage medium not assembled into the device.
[0201] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or a combination of both. Examples include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. Specific examples of storage media may include, but are not limited to, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory, optical fibers, portable CD-ROMs, optical storage devices, magnetic storage devices, etc., or any suitable combination thereof.
[0202] Computer-readable media may store one or more programs that can be used by or in conjunction with an instruction execution system. The media may be permanent or non-permanent, removable or non-removable, and may store information by any method or technology, including computer-readable instructions, data structures, program modules, or other data.
[0203] The computer program code used to implement the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, and C++) and conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. The remote computer can be connected to the user's computer via any network (including a local area network or a wide area network) or can be connected to an external computer.
[0204] In the above embodiments, the functions can be implemented in whole or in part by software, hardware, firmware, or any combination thereof, for example, by using application-specific integrated circuits, general-purpose computers, or other similar hardware devices. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions; it can also be implemented by hardware, for example, as a circuit that works in conjunction with the processor to execute the steps or functions.
[0205] This application also provides a computer program product, including one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one storage medium to another via wired (e.g., DSL) or wireless (e.g., wireless, microwave) means. The computer-readable storage medium may be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive).
[0206] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0207] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first" and "second" are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.
[0208] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.< / n> < / pad> < / del>
Claims
1. A method for processing auxiliary diagnostic data, characterized in that, The processing method includes: Receive gene sequencing data from tumor samples to be analyzed and generate somatic mutation site information; Construct reference-mutation sequence window pairs for the somatic cell mutation site information; The reference sequence window and the mutation sequence window in the reference-mutation sequence window pair are numerically processed to generate mutation window input features, and the mutation window input features are sequence encoded to obtain a mutation window vector; the mutation window input features include reference window features and mutation window features, and the mutation window vector includes reference window vector and mutation window vector; The original mutation influence vector is generated based on the difference between the mutation window vector and the reference window vector, and the original mutation influence vector is concatenated with the window auxiliary features in the mutation window input features to form a set of mutation influence vectors. A cross-variation window association matrix is constructed based on the somatic mutation site information, and the set of mutation influence vectors is weighted and fused based on the cross-variation window association matrix to form a fused set of mutation influence vectors; the cross-variation window association matrix is used to represent the association strength between two mutation windows.
2. The method for processing auxiliary diagnostic data according to claim 1, characterized in that, The gene sequencing data includes FASTQ files, BAM files, VCF files, reference genome files, normal sample data corresponding to tumor samples, and sample labels; The steps for generating somatic cell mutation site information include: When the received gene sequencing data is a FASTQ file, the DNA sequencing fragments are subjected to quality control to obtain the quality-controlled sequencing fragments. The quality-controlled sequencing fragments are compared with the reference genome file to generate a BAM file; Somatic cell variations are detected based on the differences between tumor samples and normal samples, and somatic cell variation site information is generated; the somatic cell variation site information is expressed in the form of a somatic cell variation site table, which includes field names and field meanings.
3. The method for processing auxiliary diagnostic data according to claim 1, characterized in that, The steps for constructing a reference-mutation sequence window for the somatic mutation site information include: For each somatic cell mutation site, with the mutation site as the center, L bases are cut upstream and L bases are cut downstream from the reference genome to construct the reference sequence window with a length of 2L+1, where L is the preset window radius. Based on the reference bases and mutant bases in the somatic cell mutation site information, the reference bases are replaced with mutant bases in the reference sequence window to construct the mutant sequence window. The mutation type of somatic cell mutation site information is determined, and the reference sequence window and the mutation sequence window are processed according to the mutation type to form a reference-mutation sequence window pair; All reference-mutation sequence window pairs corresponding to somatic cell variations are combined into a reference-mutation sequence window pair set.
4. The method for processing auxiliary diagnostic data according to claim 3, characterized in that, The steps of numerically processing the reference sequence window and the mutant sequence window in the reference-mutant sequence window pair to generate mutant window input features, and then performing sequence encoding on the mutant window input features to obtain a mutant window vector include: Map DNA bases and special markers to integer codes; The bases at each position in the reference sequence window and the mutant sequence window are encoded separately to form the reference sequence window encoding sequence and the mutant sequence window encoding sequence; The coding values at the same position in the reference sequence window coding sequence and the mutation sequence window coding sequence are compared bit by bit, and a reference-mutation difference marker sequence for the mutation window is generated based on the comparison results. The relative positional encoding sequence of each mutation window is composed of the relative positional encoding values at each position; Encode the mutation type, chain direction, and region type in somatic cell mutation site information to form mutation type code, chain direction code, and region type code; Sequencing depth codes are obtained based on the site sequencing depth in somatic cell variation site information. Based on the site sequencing depth and the number of sequencing fragments supporting mutated bases in somatic cell mutation site information, the frequency coding of mutated alleles is obtained. Based on the variation quality score in somatic cell variation site information, obtain the variation quality score code; The reference window coding sequence, mutation window coding sequence, reference-mutation differential marker sequence, relative position coding sequence, variant type coding, strand direction coding, sequencing depth coding, variant allele frequency coding, variant quality fraction coding, and region type coding are concatenated to form the variant window input feature; the reference window feature is composed of the reference window coding sequence and common auxiliary features, and the mutation window feature is composed of the mutation window coding sequence and the common auxiliary features. The reference window features and mutation window features are sequentially encoded using the same sequence encoder to obtain the reference window vector and mutation window vector.
5. The method for processing auxiliary diagnostic data according to claim 4, characterized in that, The steps of generating an original mutation impact vector based on the difference between the mutation window vector and the reference window vector, and concatenating the original mutation impact vector with window auxiliary features to form a set of mutation impact vectors include: Calculate the difference between the reference window vector and the mutation window vector to obtain the original mutation effect vector; The original mutation impact vector is concatenated with the window auxiliary features in the mutation window input features to form the mutation impact vector; the window auxiliary features in the mutation window input features include sequencing depth encoding, mutation allele frequency encoding, mutation quality score encoding, mutation type encoding, and region type encoding. Combine the mutation impact vectors corresponding to all mutation windows into a mutation impact vector set.
6. The method for processing auxiliary diagnostic data according to claim 5, characterized in that, The steps of constructing a cross-variation window association matrix based on the somatic mutation site information, and weighting and fusing the set of mutation influence vectors based on the cross-variation window association matrix to form a fused set of mutation influence vectors include: Construct an n×n cross-mutation window association matrix, where n represents the number of mutation windows; The variation window weight was determined by sequencing quality weight, functional region weight, and cancer relevance weight. The associated elements in the variation window association matrix are determined by the same gene relationship, location distance relationship, and functional region relationship. Normalize each row of the cross-variation window association matrix to obtain the normalized association weights; The fusion vector is determined by normalized correlation weights and mutation window weights, and the fusion vector and the mutation influence vector are weighted and fused to obtain the fused mutation influence vector and the set of mutation influence vectors.
7. The method for processing auxiliary diagnostic data according to claim 1, characterized in that, After obtaining the fusion variant impact vector, the method further includes: performing an auxiliary diagnostic comprehensive score on the fusion variant impact vector, stratifying the samples by risk based on the score results, and generating a cancer auxiliary diagnostic data processing report.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method for processing auxiliary diagnostic data as described in any one of claims 1 to 7.
9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the auxiliary diagnostic data processing method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the method for processing auxiliary diagnostic data as described in any one of claims 1 to 7.