DNA fragmented gene detection data processing method based on artificial intelligence
Through artificial intelligence-based methods, the topological structure and sequence association patterns in DNA fragmented detection data are extracted and fused, and the problem of insufficient data processing efficiency and accuracy in the prior art is solved, and more efficient and accurate data classification results are achieved.
Patent Information
- Application Number
- CN202510488583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art is difficult to efficiently and accurately extract and utilize key information in DNA fragmented gene detection data. In addition, traditional machine learning models have high computational complexity and poor generalization capabilities when processing large-scale fragmented data, making it difficult to meet the needs of efficiency and robustness in practical applications.
Using an artificial intelligence-based method, the multi-scale topological feature matrix is generated by segmenting DNA fragmented detection data into local data blocks, and continuous co-modulation analysis is performed to generate a multi-scale topological feature matrix; the attention mechanism is used to assign dynamic weights to the rings and branch structures in the topological feature matrix; the base association mode of the base sequence is extracted through one-dimensional convolution; the weighted topological features and sequence features are tensor spliced in hidden space to generate the fused higher-order feature vector; and classification is carried out through multi-layer perceptron and lightweight neural network submodules.
It effectively improves the accuracy and interpretability of the classification results of DNA fragmentation detection data, and improves the efficient extraction and fusion ability of topological structures and sequence association patterns in DNA fragmentation data.
Smart Images

Figure CN120015134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA fragment data processing, and in particular to a method for processing DNA fragmentation gene detection data based on artificial intelligence. Background Art
[0002] With the rapid development of high-throughput sequencing technology, the application of DNA fragmentation gene detection data in personalized medicine and genomics research is becoming more and more extensive. However, due to the characteristics of high dimensionality, low signal-to-noise ratio and complex spatial structure of the fragmented data generated during the sequencing process, traditional analysis methods are difficult to efficiently and accurately extract and utilize the key information therein. Existing technologies usually rely on single-dimensional feature extraction methods, such as sequence-based base association patterns or structure-based topological features, and fail to fully explore the collaborative information of sequence and spatial structure in DNA fragmentation data, resulting in insufficient classification accuracy and interpretability. In addition, traditional machine learning models often face problems such as high computational complexity and poor generalization ability when processing large-scale fragmented data, and it is difficult to meet the requirements for efficiency and robustness in practical applications. In view of the above problems, there is an urgent need for an innovative method that can integrate multi-dimensional features and adaptively process complex data structures to improve the processing efficiency of DNA fragmentation gene detection data. Summary of the invention
[0003] The present invention overcomes the deficiencies of the prior art and provides a method for processing DNA fragmentation gene detection data based on artificial intelligence.
[0004] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: The first aspect of the present invention discloses a method for processing DNA fragmentation gene detection data based on artificial intelligence, comprising the following steps: The DNA fragmentation detection data is divided into several local data blocks, and continuous homology analysis is performed on each local data block to generate a multi-scale topological feature matrix; The attention mechanism is used to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features. The base association pattern of the base sequence in the DNA fragmentation detection data is extracted through one-dimensional convolution to obtain sequence features. The weighted topological features and sequence features are concatenated in the latent space to generate a fused high-order feature vector. The fused high-order feature vectors are coarsely classified through a multi-layer perceptron, and a preliminary classification result is output; based on the preliminary classification result, a lightweight neural network submodule is constructed to perform fine-grained classification to obtain the final classification result.
[0005] Preferably, the DNA fragmentation detection data is divided into several local data blocks, and continuous coherence analysis is performed on each local data block to generate a multi-scale topological feature matrix, specifically: Based on the preset sliding window size and step length, the DNA fragmentation detection data is divided into multiple local data blocks; the window size is determined according to the average length and mutation frequency of the fragments, and the step length is set to 50% of the window size; A point cloud representation is constructed for each local data block, where each point represents a base or mutation site, and the distance matrix between points is calculated based on the physicochemical properties of the bases (such as charge and hydrophobicity); Generate the Vietoris-Rips complex based on the distance matrix, and calculate the barcode of the Vietoris-Rips complex at different scales through continuous homology analysis, and extract the topological invariant features of the ring structure (H1) and branch connection point (H0); If the continuous length of the ring structure in the barcode exceeds a preset length threshold, it is determined that the corresponding local data block has topological features and is marked as a key area; The topological features of each local data block are integrated into a multi-scale topological feature matrix, in which each feature vector contains the persistence score of the ring structure, the number of branch points and their spatial distribution information.
[0006] Preferably, the attention mechanism is used to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features, specifically: The multi-scale topological feature matrix is used as input, where each row represents the topological feature vector of a local data block, including the persistence score of the ring structure, the number of branch points and their spatial distribution information; Initialize the query, key, and value matrices in the attention mechanism, where the query matrix is generated by the global context features, the key matrix is obtained by linear transformation of the topological feature vector, and the value matrix is directly mapped to the weighted representation of the topological feature vector; Calculate the dot product of the query matrix and the key matrix, and normalize them through the Softmax function to get the attention weight. If the persistence score of the ring structure exceeds the preset score threshold, increase its corresponding attention weight by the preset amplitude. Multiply the attention weights by the value matrix to obtain the weighted topological feature vector; if there are multiple ring structures in the topological feature matrix, a multi-head attention mechanism is further introduced to calculate the weights of different ring structures and fuse their results; The weighted topological feature vectors are integrated into weighted topological features.
[0007] Preferably, the base association pattern of the base sequence in the DNA fragmentation detection data is extracted by one-dimensional convolution to obtain the sequence features, specifically: Convert DNA fragmentation detection data into a numerical sequence, where each base (A, T, C, G) is mapped to a preset numerical code and the sequence is padded to a fixed length; Initialize the one-dimensional convolution kernel, input the digitized sequence into the one-dimensional convolution layer, and extract the local base association pattern by sliding the convolution kernel; if the convolution kernel coverage area contains a known functional mutation site, increase the weight of the convolution kernel according to a preset ratio; Perform nonlinear activation (such as ReLU) on the convolution output and reduce the feature dimension through the maximum pooling layer to retain the significant features of the base association pattern in the local area; If there are repeated segments or low-complexity regions in the sequence, a residual connection is introduced to add the original sequence features to the convolution output; The pooled feature vectors are concatenated into a complete sequence feature representation, where each feature dimension corresponds to a specific base association pattern.
[0008] Preferably, the weighted topological features and the sequence features are tensor-joined in the latent space to generate a fused high-order feature vector, specifically: The weighted topological feature matrix and the sequence feature matrix are input into the latent space mapping layer respectively; the weighted topological feature matrix contains the ring structure, branch connection points and their dynamic weight information, and the sequence feature matrix contains the base association pattern and its significant features; The weighted topological features and sequence features are mapped to latent space representations of the same dimension through a fully connected layer. If the dimension of the topological feature matrix is higher than that of the sequence feature matrix, the sequence feature matrix is zero-filled or interpolated to ensure the consistency of the dimension. Normalize the mapped latent space representation, and perform tensor splicing of the normalized topological features and sequence features in the latent space; Perform nonlinear transformation (such as activation function or feature crossover) on the concatenated high-order feature representation to generate a fused high-order feature vector.
[0009] Preferably, the fused high-order feature vectors are coarsely classified by a multi-layer perceptron, and a preliminary classification result is output, specifically: The fused high-order feature vector is input into the input layer of the multi-layer perceptron, where each feature dimension corresponds to the joint representation of the weighted topological feature and the sequence feature; Perform nonlinear transformation on feature vectors through hidden layers and use activation functions (such as ReLU) to extract complex relationships between features; The output of the hidden layer is passed to the Softmax output layer to calculate the probability distribution of each category. If the maximum probability value is lower than the preset probability threshold, the feature backtracking mechanism is triggered to readjust the fusion weights of the weighted topological features and sequence features and iteratively optimize the classification results. The probability distribution of Softmax output is used as the preliminary classification result.
[0010] Preferably, a lightweight neural network submodule is constructed based on the preliminary classification results, and fine-grained classification is performed to obtain the final classification results, specifically: The probability distribution of each category is obtained based on the preliminary classification results output by the multi-layer perceptron as the input of the lightweight neural network submodule; Construct a lightweight neural network submodule to extract fine-grained features through nonlinear transformation of the probability distribution of each category in the hidden layer; The output of the hidden layer is passed to the Softmax output layer to calculate the fine-grained probability distribution of each category, and the maximum value in the fine-grained probability distribution is taken as the highest probability value, and the second largest value is taken as the second highest probability value; if the difference between the highest probability value and the second highest probability value is lower than the preset threshold, it is determined that the classification result is ambiguous, and the fusion weights of the weighted topological features and sequence features are readjusted and the classification results are iteratively optimized; Calculate the confidence of the classification result. The confidence is the difference between the highest probability value and the second highest probability value. If the confidence is lower than the preset confidence threshold, the corresponding classification result will be marked as "pending" and output to the manual review module. Otherwise, the category corresponding to the highest probability will be used as the final classification result.
[0011] The second aspect of the present invention discloses a DNA fragmentation gene detection data processing system based on artificial intelligence, the DNA fragmentation gene detection data processing system includes a memory and a processor, the memory stores a DNA fragmentation gene detection data processing method program, when the DNA fragmentation gene detection data processing method program is executed by the processor, any one of the steps of the DNA fragmentation gene detection data processing method is implemented.
[0012] The third aspect of the present invention discloses a computer-readable storage medium, which includes a DNA fragmentation gene detection data processing method program. When the DNA fragmentation gene detection data processing method program is executed by a processor, any one of the steps of the DNA fragmentation gene detection data processing method is implemented.
[0013] The present invention solves the technical defects existing in the background technology, and has the following beneficial effects: dividing DNA fragmentation detection data into a number of local data blocks, performing continuous homology analysis on each local data block, and generating a multi-scale topological feature matrix; using the attention mechanism to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix, and obtaining weighted topological features; extracting the base association pattern of the base sequence in the DNA fragmentation detection data by one-dimensional convolution, and obtaining the sequence features; performing tensor splicing of the weighted topological features and the sequence features in the latent space, and generating a fused high-order feature vector; performing coarse-grained classification on the fused high-order feature vector by a multi-layer perceptron, and outputting a preliminary classification result; constructing a lightweight neural network submodule based on the preliminary classification result, performing fine-grained classification, and obtaining the final classification result, the present invention effectively improves the accuracy and interpretability of the classification result of the DNA fragmentation detection data. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, drawings of other embodiments can be obtained based on these drawings without paying creative work.
[0015] Figure 1 This is a flow chart of the first method of the DNA fragmentation gene detection data processing method; Figure 2 This is a flow chart of the second method of the DNA fragmentation gene detection data processing method; Figure 3 This is the system block diagram of the DNA fragmentation gene detection data processing system. DETAILED DESCRIPTION
[0016] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0017] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0018] like Figure 1 As shown, the first aspect of the present invention discloses a method for processing DNA fragmentation gene detection data based on artificial intelligence, comprising the following steps: S102, dividing the DNA fragmentation detection data into a number of local data blocks, performing continuous homology analysis on each local data block, and generating a multi-scale topological feature matrix; S104, using an attention mechanism to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features; extracting base association patterns of base sequences in DNA fragmentation detection data through one-dimensional convolution to obtain sequence features; S106, performing tensor splicing of the weighted topological features and the sequence features in the latent space to generate a fused high-order feature vector; S108, performing coarse-grained classification on the fused high-order feature vectors through a multi-layer perceptron, and outputting a preliminary classification result; constructing a lightweight neural network submodule based on the preliminary classification result, performing fine-grained classification, and obtaining a final classification result.
[0019] It should be noted that the present invention realizes the efficient extraction and fusion of topological structures and sequence association patterns in DNA fragmentation data by introducing technologies such as continuous homology analysis, attention mechanism and one-dimensional convolution, and combines the collaborative classification strategy of multi-layer perceptron and lightweight neural network sub-modules to effectively improve the accuracy and interpretability of the classification results of DNA fragmentation detection data.
[0020] Preferably, the DNA fragmentation detection data is divided into several local data blocks, and continuous homology analysis is performed on each local data block to generate a multi-scale topological feature matrix, such as Figure 2 As shown, specifically: S202, based on a preset sliding window size and step length, dividing the DNA fragmentation detection data into a plurality of local data blocks; wherein the window size is determined according to the average length and mutation frequency of the fragments, and the step length is set to 50% of the window size; S204, constructing a point cloud representation for each local data block, wherein each point represents a base or a mutation site, and calculating a distance matrix between points according to the physical and chemical properties of the bases; It should be noted that each base or mutation site in the local data block is mapped to a point in three-dimensional space, and its coordinates are determined by the physical and chemical properties of the base (such as charge, hydrophobicity, molecular weight) through preset conversion rules; secondly, based on the coordinates of each point, the Euclidean distance between points is calculated. According to the symmetry and non-negativity of the distance matrix, a complete distance matrix is constructed, in which each element represents the weighted distance between two points.
[0021] S206. Generate the Vietoris-Rips complex based on the distance matrix, and calculate the barcode of the Vietoris-Rips complex at different scales through continuous homology analysis, and extract the topological invariant features of the ring structure (H1) and branch connection point (H0); Among them, Vietoris-Rips complex is a topological structure construction method based on point cloud data, which is used to characterize the high-order spatial relationship between data points.
[0022] It should be noted that, according to the distance matrix, the Vietoris-Rips complex is constructed with scale parameters that gradually increase in a preset ratio, where each scale parameter corresponds to a specific distance threshold. When the distance between points is less than the threshold, they are connected as edges, triangles or higher-dimensional simplexes (i.e., Vietoris-Rips complexes). The topological structure changes of the complex at different scales are tracked by continuous homology analysis, and the generation and disappearance process of the ring structure (H1) and the branch connection point (H0) are recorded; then, the generation and disappearance scale parameters of each topological feature are plotted as a barcode, where each bar represents the duration interval of a topological feature; if the persistence length of the ring structure in the barcode exceeds the preset threshold, the ring structure is determined to be a significant topological feature; finally, the persistence scores and spatial distribution information of all significant ring structures and branch connection points are extracted to generate a topologically invariant feature vector.
[0023] S208, if the continuous length of the ring structure in the barcode exceeds a preset length threshold, it is determined that the corresponding local data block has a topological feature and is marked as a key area; S210, integrating the topological features of each local data block into a multi-scale topological feature matrix, wherein each feature vector contains the continuity score of the ring structure, the number of branch points and their spatial distribution information.
[0024] In summary, this method can efficiently extract the multi-scale topological feature matrix and accurately characterize the complex spatial structure in DNA fragmentation data by dividing the DNA fragmentation detection data into local data blocks and performing continuous coherence analysis.
[0025] Preferably, the attention mechanism is used to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features, specifically: The multi-scale topological feature matrix is used as input, where each row represents the topological feature vector of a local data block, including the persistence score of the ring structure, the number of branch points and their spatial distribution information; Initialize the query, key, and value matrices in the attention mechanism, where the query matrix is generated by the global context features, the key matrix is obtained by linear transformation of the topological feature vector, and the value matrix is directly mapped to the weighted representation of the topological feature vector; Calculate the dot product of the query matrix and the key matrix, and normalize them through the Softmax function to get the attention weight. If the persistence score of the ring structure exceeds the preset score threshold, increase its corresponding attention weight by the preset amplitude. The attention weight is multiplied by the value matrix to obtain a weighted topological feature vector, and the weighted topological feature vector is integrated into a weighted topological feature.
[0026] Among them, if there are multiple ring structures in the topological feature matrix, a multi-head attention mechanism is further introduced to calculate the weights of different ring structures separately and merge their results.
[0027] In summary, this method improves the representation ability of key topological features by introducing the attention mechanism to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix. It can adaptively focus on the ring structure with higher persistence scores and enhance its contribution to classification decisions. The weighted topological features finally generated provide high-dimensional and precise feature representation for subsequent classification and feature fusion, thereby improving the accuracy and interpretability of genetic testing data processing.
[0028] Preferably, the base association pattern of the base sequence in the DNA fragmentation detection data is extracted by one-dimensional convolution to obtain the sequence features, specifically: Convert DNA fragmentation detection data into a numerical sequence, in which each base is mapped to a preset numerical code, and pad the sequence to a fixed length; Initialize the one-dimensional convolution kernel, input the digitized sequence into the one-dimensional convolution layer, and extract the local base association pattern by sliding the convolution kernel; if the convolution kernel coverage area contains a known functional mutation site, increase the weight of the convolution kernel according to a preset ratio; Perform nonlinear activation (such as ReLU) on the convolution output and reduce the feature dimension through the maximum pooling layer to retain the significant features of the base association pattern in the local area; It should be noted that, according to the complexity of the base sequence and the distribution of functional mutation sites, multiple one-dimensional convolution kernels are initialized, for example, their size is set to 3 to 7 bases in length, and the number of convolution kernels is dynamically adjusted according to the complexity of the sequence. The numerical sequence is input into the one-dimensional convolution layer, and the local base association pattern is extracted by sliding the convolution kernel. If the convolution kernel coverage area contains known functional mutation sites, the weight of the convolution kernel is adjusted according to a preset ratio (such as increasing by 20%) to enhance the base pattern extraction effect in the key area; then, the convolution output is nonlinearly activated (such as ReLU) to capture the nonlinear relationship in the base association pattern; then, the activated features are reduced in dimension through the maximum pooling layer to retain the significant features of the base association pattern in the local area.
[0029] If there are repeated segments or low-complexity regions in the sequence, a residual connection is introduced to add the original sequence features to the convolution output; It should be noted that before the convolution operation, the digitized sequence is preprocessed to mark the locations of repeated segments or low-complexity regions. After the convolution layer output, the original sequence features are added element by element to the convolution output through the residual connection.
[0030] The pooled feature vectors are concatenated into a complete sequence feature representation, where each feature dimension corresponds to a specific base association pattern.
[0031] The base sequence refers to the sequence formed by converting the bases (A, T, C, G) in the DNA fragment into a numerical representation according to their arrangement order in the genome. For example, a DNA fragment "ATCG" can be encoded as a numerical sequence [0, 1,2, 3], where each number represents a specific base. This numerical sequence is the input of the subsequent one-dimensional convolution operation to extract the local correlation pattern between bases, thereby generating sequence features.
[0032] In summary, this method extracts the local correlation pattern of base sequences in DNA fragmentation detection data through one-dimensional convolution, which can efficiently capture the nonlinear relationship between bases and improve the characterization ability of sequence features.
[0033] Preferably, the weighted topological features and the sequence features are tensor-joined in the latent space to generate a fused high-order feature vector, specifically: The weighted topological feature matrix and the sequence feature matrix are input into the latent space mapping layer respectively; the weighted topological feature matrix contains the ring structure, branch connection points and their dynamic weight information, and the sequence feature matrix contains the base association pattern and its significant features; The weighted topological features and sequence features are mapped to latent space representations of the same dimension through a fully connected layer. If the dimension of the topological feature matrix is higher than that of the sequence feature matrix, the sequence feature matrix is zero-filled or interpolated to ensure the consistency of the dimension. It should be noted that the weighted topological feature matrix and the sequence feature matrix are respectively input into the fully connected layer, where the output dimension of the fully connected layer is pre-set according to the dimension of the topological feature matrix; secondly, the weighted topological features are mapped to the latent space representation of the target dimension through the fully connected layer. If the dimension of the sequence feature matrix is lower than the target dimension, the sequence feature matrix is zero-padded, zero values are added to the end of the matrix to expand its dimension, or new feature values are generated between existing features through interpolation methods (such as linear interpolation) to match the target dimension.
[0034] Normalize the mapped latent space representation, and perform tensor splicing of the normalized topological features and sequence features in the latent space; It should be noted that the normalized topological feature matrix and sequence feature matrix are converted into tensor forms of the same dimension, where the topological feature matrix contains the ring structure, branch connection points and their dynamic weight information, and the sequence feature matrix contains the base association pattern and its significant features. The two tensors are concatenated along the feature dimension, and the shape of the concatenated tensor is adjusted to meet the format requirements of the subsequent model input.
[0035] Perform nonlinear transformation (such as activation function or feature crossover) on the concatenated high-order feature representation to generate a fused high-order feature vector.
[0036] It should be noted that the concatenated high-order feature representation is input into the nonlinear transformation layer, where the activation function (such as ReLU) is used to capture the nonlinear relationship between features. If the feature representation contains topological and sequence information of multiple dimensions, a feature crossover operation is introduced to enhance the synergy between features through element-by-element multiplication or weighted combination. The activated or crossed features are subjected to dimensionality reduction, and the key feature information is retained through a fully connected layer or pooling operation. The reduced-dimensional features are normalized to eliminate the dimensional differences between features and improve the convergence speed of the model; then, the normalized feature vectors are integrated to generate a fused high-order feature vector; finally, the generated high-order feature vector is used as the input of the subsequent classification or decision model to ensure the deep fusion and efficient use of topological and sequence information.
[0037] In summary, this method achieves a deep fusion of topological structure and sequence association patterns in DNA fragmentation data by tensor splicing weighted topological features and sequence features in latent space, thereby improving the comprehensiveness and robustness of feature representation.
[0038] Preferably, the fused high-order feature vectors are coarsely classified by a multi-layer perceptron, and a preliminary classification result is output, specifically: The fused high-order feature vector is input into the input layer of the multi-layer perceptron, where each feature dimension corresponds to the joint representation of the weighted topological feature and the sequence feature; Perform nonlinear transformation on feature vectors through hidden layers and use activation functions (such as ReLU) to extract complex relationships between features; The output of the hidden layer is passed to the Softmax output layer to calculate the probability distribution of each category. If the maximum probability value is lower than the preset probability threshold, the feature backtracking mechanism is triggered to readjust the fusion weights of the weighted topological features and sequence features and iteratively optimize the classification results. The probability distribution of Softmax output is used as the preliminary classification result.
[0039] It should be noted that the Multilayer Perceptron (MLP) is a feedforward artificial neural network consisting of an input layer, a hidden layer, and an output layer. It can model complex data by learning nonlinear mapping relationships. Its working principle is as follows: the input layer receives feature vectors, the hidden layer transforms the input features through nonlinear activation functions (such as ReLU) to extract complex relationships between features, and the output layer generates classification or regression results through Softmax or Sigmoid functions. MLP optimizes network parameters through the back-propagation algorithm, can adaptively capture high-order features in the data, and is suitable for classification and regression tasks of high-dimensional, nonlinear data. It has strong fitting ability and generalization performance.
[0040] In summary, this method uses a multi-layer perceptron to perform coarse-grained classification on the fused high-order feature vectors, which can efficiently capture the complex relationship between topology and sequence information and improve the accuracy of the classification results. The final output of the preliminary classification results provides high-quality candidate categories for subsequent fine-grained classification, thereby improving the overall efficiency and accuracy of genetic testing data processing.
[0041] Preferably, a lightweight neural network submodule is constructed based on the preliminary classification results, and fine-grained classification is performed to obtain the final classification results, specifically: The probability distribution of each category is obtained based on the preliminary classification results output by the multi-layer perceptron as the input of the lightweight neural network submodule; Construct a lightweight neural network submodule to extract fine-grained features through nonlinear transformation of the probability distribution of each category in the hidden layer; The output of the hidden layer is passed to the Softmax output layer to calculate the fine-grained probability distribution of each category, and the maximum value in the fine-grained probability distribution is taken as the highest probability value, and the second largest value is taken as the second highest probability value; if the difference between the highest probability value and the second highest probability value is lower than the preset threshold, it is determined that the classification result is ambiguous, and the fusion weights of the weighted topological features and sequence features are readjusted and the classification results are iteratively optimized; It should be noted that the difference between the highest probability value and the second highest probability value in the fine-grained probability distribution generated by the Softmax output layer is calculated. If the difference between the two is lower than the preset threshold (such as 0.1), the classification result is determined to be ambiguous. According to the ambiguous classification result, go back to the feature fusion stage, adjust the fusion weights of weighted topological features and sequence features, and increase the contribution of key features according to the preset proportion. Re-input the adjusted feature vector into the lightweight neural network submodule, perform fine-grained classification and calculate the new probability distribution; repeat the above steps until the difference between the highest probability value and the second highest probability value is not lower than the preset threshold or the maximum number of iterations is reached; finally, the optimized classification result is used as the final output to ensure the rigor and accuracy of the classification process.
[0042] Calculate the confidence of the classification result. The confidence is the difference between the highest probability value and the second highest probability value. If the confidence is lower than the preset confidence threshold, the corresponding classification result will be marked as "pending" and output to the manual review module. Otherwise, the category corresponding to the highest probability will be used as the final classification result.
[0043] The final classification results include the following specific contents: first, the category corresponding to the highest probability is taken as the main classification result, indicating the category to which the model believes the sample most likely belongs; second, the confidence of the classification result, that is, the difference between the highest probability value and the second highest probability value, is used to measure the reliability of the classification result; if the confidence is lower than the preset confidence threshold, the sample is marked as "pending" and output to the manual review module for further verification; in addition, it also includes the feature dimensions that contribute most to the classification results and their corresponding topological structures and base association patterns, which are interpretable information to assist users in understanding the basis of classification decisions.
[0044] It should be noted that this method further improves the accuracy and reliability of the classification results by constructing a lightweight neural network submodule to perform fine-grained classification on the preliminary classification results. Specifically, through the nonlinear transformation of the hidden layer and the probability calculation of the Softmax output layer, it is possible to adaptively extract fine-grained features and generate high-confidence classification results. If the classification result is ambiguous or the confidence is lower than the preset threshold, the feature fusion weight adjustment mechanism or the manual review process is triggered to ensure the rigor and accuracy of the classification process. The final output classification results provide high-precision discrimination results for genetic testing data, improving the credibility and practicality of data processing.
[0045] In this embodiment, the DNA fragmentation gene detection data processing method further includes the following steps: The gene mutations or functional regions in the final classification results are extracted as query sequences. If the final classification results contain multiple mutation sites, they are sorted according to the preset priority. Secondly, the query sequence is compared with the gene function annotation database (such as GO, KEGG), and the corresponding entries of the query sequence in the database are identified through sequence alignment algorithms (such as BLAST). Then, the biological function, pathway and disease association information are extracted according to the alignment results. If the query sequence has no direct matching entries in the database, its function is inferred through homologous genes or conserved regions. Then, the extracted functional information is associated with the classification results and annotated. If the same mutation site corresponds to multiple functions or pathways, the most relevant annotation is selected according to the functional significance score (such as P value). Finally, the annotation results are integrated into a functional annotation report, which includes mutation sites, functional descriptions, pathway associations, disease risks and references.
[0046] It should be noted that this method realizes the automated functional annotation and association analysis of gene mutation sites by dynamically associating the classification results with the gene function annotation database; it integrates multidimensional information of mutation sites, functional pathways and disease risks to generate structured reports, providing indirect guidance for clinical diagnosis and scientific research.
[0047] In this embodiment, the DNA fragmentation gene detection data processing method further includes the following steps: After generating the fused high-order feature vector, epigenetic data (such as methylation and chromatin conformation capture data) are introduced; The epigenetic data are converted into a three-dimensional spatial coordinate and interaction strength matrix, where the position of each DNA fragment is determined by its genomic coordinates and the interaction frequency of neighboring fragments; It should be noted that the genomic coordinates of each DNA fragment are extracted based on epigenetic data (such as chromatin conformation capture data) and mapped to the initial position in three-dimensional space. The spatial distance between fragments is calculated by the interaction frequency of adjacent fragments (such as the interaction intensity in the Hi-C matrix). If the interaction frequency is higher than the preset threshold, the spatial distance between fragments is shortened, otherwise the distance is increased. The interaction intensity matrix is constructed based on the three-dimensional coordinates, where the matrix elements represent the spatial interaction intensity between fragments.
[0048] Constructing a graph neural network, in which nodes represent DNA fragment regions and edges represent spatial interaction strengths (obtained in the interaction strength matrix); aggregating features of neighboring nodes through a graph convolution layer to generate three-dimensional structural features; It should be noted that the DNA fragment region is used as the node of the graph neural network, and the feature vector of the node includes its genomic coordinates, epigenetic markers (such as methylation level) and the interaction strength of the adjacent fragments. The edges of the graph are constructed according to the interaction strength matrix. If the interaction strength is higher than the preset threshold, an edge is established between the corresponding nodes, and the weight of the edge is determined by the interaction strength value; then, the features of the adjacent nodes are aggregated through the graph convolution layer, where the features of each node are updated to the weighted sum of its own features and the features of the adjacent nodes, and the weight is determined by the interaction strength of the edge.
[0049] The fused high-order feature vector is cross-modally fused with the three-dimensional structural feature to obtain the cross-modal fusion feature; the attention mechanism dynamically adjusts the contribution of the three-dimensional feature in the classification decision according to its significance; It should be noted that the fused high-order feature vector and the three-dimensional structural feature are mapped to the latent space representation of the same dimension respectively. Initialize the query matrix (Query), key matrix (Key) and value matrix (Value) in the cross-modal attention mechanism. Calculate the dot product of the query matrix and the key matrix, and normalize them through the Softmax function to obtain the attention weight. If the significance score of some dimensions in the three-dimensional structural feature is higher than the preset threshold, increase the corresponding attention weight according to the preset ratio; then, multiply the adjusted attention weight with the value matrix to obtain the weighted three-dimensional structural feature; finally, concatenate the weighted three-dimensional structural feature with the high-order feature vector to generate the cross-modal fusion feature.
[0050] The cross-modal fusion features are input into a lightweight neural network to generate the final classification result, and the segment area that has the greatest impact on the classification result is located through inverse gradient saliency mapping (such as Grad-CAM); It should be noted that the contribution of each feature dimension to the classification result is calculated through the reverse gradient significance mapping. If the contribution is higher than the preset threshold, it will be marked as a key feature. According to the DNA fragment region corresponding to the key feature, the fragment region with the greatest impact on the classification result is located, and a visual significance map is generated as an interpretable output of the classification result.
[0051] The localization results are superimposed with the epigenetic data to generate a visual interpretable map, annotate the functional annotations of key fragment regions and their contribution to the classification results, and output them through the user terminal.
[0052] In summary, this method realizes multimodal feature fusion and three-dimensional space modeling, thereby effectively improving the accuracy and interpretability of classification results.
[0053] like Figure 2As shown, the second aspect of the present invention discloses a DNA fragmentation gene detection data processing system 6 based on artificial intelligence, wherein the DNA fragmentation gene detection data processing system comprises a memory 41 and a processor 52, wherein the memory 41 stores a DNA fragmentation gene detection data processing method program, and when the DNA fragmentation gene detection data processing method program is executed by the processor 52, any one of the steps of the DNA fragmentation gene detection data processing method is implemented.
[0054] The third aspect of the present invention discloses a computer-readable storage medium, which includes a DNA fragmentation gene detection data processing method program. When the DNA fragmentation gene detection data processing method program is executed by a processor, any one of the steps of the DNA fragmentation gene detection data processing method is implemented.
[0055] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0056] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0057] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0058] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0059] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods of each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0060] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A DNA fragmentation gene detection data processing method based on artificial intelligence, characterized in that: The following steps are involved: The DNA fragmentation detection data is divided into several local data blocks, and continuous homology analysis is performed on each local data block to generate a multi-scale topological feature matrix; The attention mechanism is used to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features. The base association pattern of the base sequence in the DNA fragmentation detection data is extracted through one-dimensional convolution to obtain sequence features. The weighted topological features and sequence features are concatenated in the latent space to generate a fused high-order feature vector. The fused high-order feature vectors are coarsely classified through a multi-layer perceptron, and a preliminary classification result is output; based on the preliminary classification result, a lightweight neural network sub-module is constructed to perform fine-grained classification to obtain the final classification result.
2. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: The DNA fragmentation detection data is divided into several local data blocks, and continuous homology analysis is performed on each local data block to generate a multi-scale topological feature matrix, specifically: Based on a preset sliding window size and step size, the DNA fragmentation detection data is divided into a plurality of local data blocks; A point cloud representation is constructed for each local data block, where each point represents a base or mutation site, and the distance matrix between points is calculated based on the physical and chemical properties of the bases; Generate Vietoris-Rips complex based on distance matrix, calculate barcodes of Vietoris-Rips complex at different scales through continuous homology analysis, and extract topologically invariant features of ring structure and branch connection points; If the continuous length of the ring structure in the barcode exceeds a preset length threshold, it is determined that the corresponding local data block has topological features and is marked as a key area; The topological features of each local data block are integrated into a multi-scale topological feature matrix, in which each feature vector contains the persistence score of the ring structure, the number of branch points and their spatial distribution information.
3. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: The attention mechanism is used to assign dynamic weights to the ring and branch structures in the multi-scale topological feature matrix to obtain weighted topological features, specifically: The multi-scale topological feature matrix is used as input, where each row represents the topological feature vector of a local data block, including the persistence score of the ring structure, the number of branch points and their spatial distribution information; Initialize the query, key, and value matrices in the attention mechanism, where the query matrix is generated by the global context features, the key matrix is obtained by linear transformation of the topological feature vector, and the value matrix is directly mapped to the weighted representation of the topological feature vector; Calculate the dot product of the query matrix and the key matrix, and normalize them through the Softmax function to get the attention weight. If the persistence score of the ring structure exceeds the preset score threshold, increase its corresponding attention weight by the preset amplitude. The attention weight is multiplied by the value matrix to obtain a weighted topological feature vector; the weighted topological feature vector is integrated into a weighted topological feature.
4. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: The base association pattern of the base sequence in the DNA fragmentation detection data is extracted by one-dimensional convolution to obtain the sequence features, specifically: Convert DNA fragmentation detection data into a numerical sequence, where each base (A, T, C, G) is mapped to a preset numerical code and the sequence is padded to a fixed length; Initialize the one-dimensional convolution kernel, input the digitized sequence into the one-dimensional convolution layer, and extract the local base association pattern by sliding the convolution kernel; The convolution output is nonlinearly activated and the feature dimension is reduced through the maximum pooling layer to retain the significant features of the base association pattern in the local area; If there are repeated segments or low-complexity regions in the sequence, a residual connection is introduced to add the original sequence features to the convolution output; The pooled feature vectors are concatenated into a complete sequence feature representation, where each feature dimension corresponds to a specific base association pattern.
5. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: The weighted topological features and sequence features are concatenated in the latent space to generate a fused high-order feature vector, specifically: The weighted topological feature matrix and the sequence feature matrix are input into the latent space mapping layer respectively; the weighted topological feature matrix contains the ring structure, branch connection points and their dynamic weight information, and the sequence feature matrix contains the base association pattern and its significant features; The weighted topological features and sequence features are mapped to latent space representations of the same dimension through a fully connected layer. If the dimension of the topological feature matrix is higher than that of the sequence feature matrix, the sequence feature matrix is zero-filled or interpolated to ensure the consistency of the dimension. Normalize the mapped latent space representation, and perform tensor splicing of the normalized topological features and sequence features in the latent space; The concatenated high-order feature representation is transformed nonlinearly to generate a fused high-order feature vector.
6. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: The fused high-order feature vectors are coarsely classified through a multi-layer perceptron, and the preliminary classification results are output, specifically: The fused high-order feature vector is input into the input layer of the multi-layer perceptron, where each feature dimension corresponds to the joint representation of the weighted topological feature and the sequence feature; Perform nonlinear transformation on feature vectors through hidden layers and use activation functions (such as ReLU) to extract complex relationships between features; The output of the hidden layer is passed to the Softmax output layer to calculate the probability distribution of each category. If the maximum probability value is lower than the preset probability threshold, the feature backtracking mechanism is triggered to readjust the fusion weights of the weighted topological features and sequence features and iteratively optimize the classification results. The probability distribution of Softmax output is used as the preliminary classification result.
7. The method for processing DNA fragmentation gene detection data based on artificial intelligence according to claim 1, characterized in that: Based on the preliminary classification results, a lightweight neural network submodule is constructed to perform fine-grained classification and obtain the final classification results, which are as follows: The probability distribution of each category is obtained based on the preliminary classification results output by the multi-layer perceptron as the input of the lightweight neural network submodule; Construct a lightweight neural network submodule to extract fine-grained features through nonlinear transformation of the probability distribution of each category in the hidden layer; The output of the hidden layer is passed to the Softmax output layer to calculate the fine-grained probability distribution of each category, and the maximum value in the fine-grained probability distribution is taken as the highest probability value, and the second largest value is taken as the second highest probability value; if the difference between the highest probability value and the second highest probability value is lower than the preset threshold, it is determined that the classification result is ambiguous, and the fusion weights of the weighted topological features and sequence features are readjusted and the classification results are iteratively optimized; Calculate the confidence of the classification result. The confidence is the difference between the highest probability value and the second highest probability value. If the confidence is lower than the preset confidence threshold, the corresponding classification result is marked as "pending" and output to the manual review module. Otherwise, the category corresponding to the highest probability is used as the final classification result.
8. The DNA fragmentation gene detection data processing system based on artificial intelligence is characterized by: The DNA fragmentation gene detection data processing system includes a memory and a processor, wherein the memory stores a DNA fragmentation gene detection data processing method program. When the DNA fragmentation gene detection data processing method program is executed by the processor, the DNA fragmentation gene detection data processing method steps as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a DNA fragmentation gene detection data processing method program. When the DNA fragmentation gene detection data processing method program is executed by a processor, the steps of the DNA fragmentation gene detection data processing method as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Base calling using convolutions
CN112313750A
Novel multi-granularity feature fusion method based on attention mechanism
CN115905999A
Test data management and analysis system
CN119091967A
Construction method for predicting RNA and small molecule binding site model based on RNA sequence and structure information characteristics and application of construction method
CN119274651A
Genetic testing method, signature extraction method, apparatus, device, and system
US20230170047A1
Cited By
Prediction method and system for testicular sertoli cells damaged by vomitoxin
CN120954508A
A method and system for predicting vomitoxin damage to sertoli cells
CN120954508B
Gene data analysis system based on AI
CN121122387A
An AI-based gene data analysis system
CN121122387B