Gene data fused multi-cancer screening model construction method and system
By generating a gene fusion matrix, constructing a gene expression topology map, and using graph neural network training, combining density clustering to identify gene mutation hot spots and dynamically adjusting screening weights, the problems of data fusion and platform compatibility of multiple cancer screening models are solved, and high accuracy and high efficiency multi-cancer screening is achieved.
Patent Information
- Application Number
- CN202510689707.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing gene screening models are difficult to achieve simultaneous screening of multiple cancers, and due to the inconsistent data distribution of different sequencing platforms, the model often experiences expression shifts, information redundancy or feature distortion during multi-source data fusion, affecting the generalization performance and prediction accuracy of the model.
By obtaining gene sequencing data from the multi-source sequencing platform, a gene fusion matrix is generated, the overlapping status of the gene expression characteristics of each target cancer species and the normal gene expression characteristics is determined, overlapping gene segments are masked, gene expression topology map is constructed, and a graph neural network is used for training to construct a multi-cancer screening model. At the same time, the hot spots of gene mutations were identified through density clustering algorithms, and the screening weight of gene fragments was dynamically adjusted according to the model load.
It improves the accuracy, efficiency and scalability of cancer screening, enhances the sensitivity and distinction of the model to early cancer signals, and significantly improves the accuracy and generalization performance of multi-classification screening.
Smart Images

Figure CN120199326A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of early cancer screening, and particularly to a method and system for constructing a multi-cancer screening model integrating gene data. Background Art
[0002] In recent years, the early screening and precise diagnosis of cancer have gradually become the focus of clinical oncology and bioinformatics research. Traditional cancer screening methods mainly rely on tissue biopsy, imaging examination, serum biomarker detection, etc., but have problems such as strong invasiveness, insufficient sensitivity and specificity, and limited ability to identify early tiny lesions. With the development of high-throughput sequencing technology, using genomic data for cancer detection has gradually become an emerging trend, especially showing great potential in the simultaneous screening of multiple cancer types.
[0003] Existing gene screening models usually model for a single cancer type, ignoring the shared and different information in gene expression patterns, mutation characteristics, and chromosomal positions among different cancer types, and it is difficult to achieve multi-cancer screening under a unified platform. In addition, due to the problem of inconsistent data distribution among different sequencing platforms, the model often shows expression shift, information redundancy, or feature distortion during the multi-source data fusion process, affecting the generalization performance and prediction accuracy of the model.
[0004] At the same time, existing methods often adopt full-feature input when constructing the gene expression feature space, failing to effectively filter the overlapping features between cancer samples and normal samples, resulting in insufficient discrimination of the expression features learned by the model; when processing gene mutation site data, simple frequency statistics are mostly used, failing to deeply explore the mutation aggregation patterns and hot spot region distributions of different cancer types. In addition, most existing models have static parameter settings and are difficult to dynamically optimize the screening strategy according to the system operation load, thus affecting the real-time performance and resource allocation efficiency in large-scale screening scenarios.
[0005] Therefore, there is an urgent need for a method and system for constructing a multi-cancer screening model that integrates multi-source gene sequencing data, takes into account the similarities and differences among cancer types, and has dynamic adaptive capabilities, so as to improve the accuracy, efficiency, and scalability of cancer screening. Summary of the Invention
[0006] In order to solve at least one of the above technical problems, the present invention proposes a method and system for constructing a multi-cancer screening model integrating gene data.
[0007] The first aspect of the present invention provides a method for constructing a multi-cancer screening model integrating gene data, including: Obtain gene sequencing data of multi-source sequencing platforms for the target cancer population, and fuse the gene sequencing data of the multi-source sequencing platforms to generate a gene fusion matrix; Determine the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics according to the gene fusion matrix, mask the overlapping gene segments, and construct gene expression topological maps for different cancer types; Import the gene expression topological map into a graph neural network for training to construct a multi-cancer screening model; Determine the gene mutation positions of the target cancer type according to the gene sequencing data, perform cluster analysis on the gene mutation positions through a density clustering algorithm, and identify the gene mutation hot spots of each target cancer type; Obtain the gene data screening load of the multi-cancer screening model. If the screening load exceeds the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene segments according to the gene mutation hot spots.
[0008] In this solution, the gene sequencing data of the multi-source sequencing platform for the target cancer type population is obtained, and the gene sequencing data of the multi-source sequencing platform is fused to generate a gene fusion matrix. Specifically: Obtain the gene sequencing data of the multi-source sequencing platform for the target cancer type population within a preset time period, and calculate the gene expression statistics of the gene sequencing data of each sequencing platform. The gene expression statistics include the mean, variance, skewness, and kurtosis of the gene sequencing data; Draw a kernel density estimation curve of the gene expression level according to the gene expression statistics, fit the kernel density estimation curve based on the maximum likelihood estimation method, and determine the gene expression level distribution of each cancer type in each sequencing platform; Judge the difference in the gene expression level distribution of the same cancer type in different sequencing platforms according to the gene expression level distribution. If the distribution difference is greater than the preset distribution difference value, for the same cancer type, randomly select a sequencing platform as the reference platform based on the quantile matching algorithm, and the remaining sequencing platforms are used as target sequencing platforms. A quantile space is constructed based on the gene expression level distribution of the reference platform's gene sequencing data; Calculate the quantile interval of the quantile space based on the gene expression level distribution of the reference platform, map the gene sequencing data of the target sequencing platform to the quantile space of the reference platform, and perform quantile interpolation on the gene expression level of the target sequencing platform through a linear interpolation algorithm to generate a mapping function corresponding to the reference platform's quantile space; Perform a non-linear transformation on the gene expression level of the target sequencing platform according to the mapping function to adjust the deviation of the gene expression level of the target sequencing platform from the reference platform in terms of distribution skewness and kurtosis, and generate normalized gene expression data; Align the gene segments of the same cancer type from different sequencing platforms according to the normalized gene expression data to generate a gene fusion matrix.
[0009] In this solution, determining the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics according to the gene fusion matrix, masking the overlapping gene segments, and constructing the gene expression topological maps of different cancer types are specifically as follows: Extract the gene expression feature vectors of the target cancer type samples based on the gene fusion matrix. The gene expression feature vectors segment the gene sequences through a sliding window, calculate the mean expression level, variance of all gene fragments within each gene segment, and the gradient change relationship of the expression levels between adjacent genes, and generate a multi-dimensional feature representation by fusing the chromosomal position information of the gene segments; Simultaneously extract the gene expression feature vectors of normal samples, construct a normal gene expression feature benchmark set, and calculate the Mahalanobis distance between the target cancer type feature vectors and the normal feature vectors in the benchmark set according to the spatial distribution of the target cancer type gene expression feature vectors and the normal gene expression feature vectors in the benchmark set; Calculate the feature similarity between each gene segment according to the Mahalanobis distance, generate a gene segment similarity score, perform a binary judgment on the gene segment similarity score according to a preset similarity threshold, and screen out the gene segments with scores higher than the similarity threshold as overlapping gene segments; Replace the gene expression data corresponding to the overlapping gene segments in the gene fusion matrix with random noise masks, and perform Gaussian smoothing interpolation on the gene expression data at the mask boundaries; Divide the gene fusion matrix after mask processing into continuous gene segments according to a preset window, and extract the node feature vectors of each gene segment. The node feature vectors include gene expression statistics and mutation frequency information; Construct an adjacency relationship matrix of the gene segments based on the chromosomal position information of the gene segments. According to the adjacency relationship matrix and the node feature vectors, perform dimensionality reduction processing on the spatial correlation of the gene segments through a graph embedding algorithm to generate a gene expression topological map.
[0010] In this solution, importing the gene expression topological map into a graph neural network for training to construct a multi-cancer type screening model is specifically as follows: Input the node feature vectors and the adjacency relationship matrix of the gene expression topological map into a graph convolutional network with a preset number of layers, and update the node embedding representations of each gene segment by aggregating the feature information of adjacent nodes; Introduce a residual connection mechanism between multiple graph convolutional layers, adjust the update weights of the node feature vectors in combination with the mutation frequency information of the gene segments, and adopt a graph attention mechanism to adaptively learn the association strength between nodes to generate a gene expression topological embedding representation with spatial dependence; The updated node feature vectors are mapped into graph-level feature vectors through a global max pooling layer, and the graph-level feature vectors are input into a fully connected classifier to output the predicted probability distribution of the target sample belonging to different cancer types; The cross-entropy loss function is calculated based on the predicted probability distribution and the true cancer type labels, and the parameters of the graph neural network are iteratively optimized using the backpropagation algorithm to obtain a trained multi-cancer screening model.
[0011] In this solution, to determine the gene mutation positions of the target cancer types based on the gene sequencing data and perform clustering analysis on the gene mutation positions through a density clustering algorithm to identify the gene mutation hotspots of each target cancer type, specifically: Extract the chromosomal coordinate information of all gene mutation sites based on the gene sequencing data of multiple samples of each target cancer type, and construct a gene mutation position matrix containing sample ID, chromosome number, start site, and end site according to the chromosome number and base position; Introduce a density clustering algorithm to calculate the spatial distribution density of mutation sites among different samples of the target cancer type, set the neighborhood radius parameter and the minimum sample number threshold of the density clustering algorithm, calculate the number of mutant samples contained within the neighborhood radius of each mutation site, and mark a mutation site as a core point if the number of neighborhood samples of a mutation site exceeds the minimum sample number threshold; Use the core points as seed points for region expansion, traverse all mutation sites within the neighborhood of the core points and calculate their reachability, and classify the directly density-reachable mutation sites into the same clustering cluster; Iteratively execute the region expansion process until no new associated mutation sites can be added, output the clustering clusters, and determine the gene mutation situation of each clustering cluster to obtain the clustering result; Determine the gene mutation hotspots of each target cancer type according to the clustering result.
[0012] In this solution, to obtain the gene data screening load situation of the multi-cancer screening model, if the screening load exceeds the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the gene mutation hotspots, specifically: Real-time monitor the gene data screening queue length and the screening response time of a single gene fragment of the multi-cancer screening model, and determine the current screening load level of the multi-cancer screening model according to the screening queue length and the screening response time; If the current screening load level is greater than the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the mutation hotspots.
[0013] In a second aspect of the present invention, there is also provided a system for constructing a multi-cancer screening model integrating gene data, the system comprising: a memory, a processor, wherein the memory includes a program for the method of constructing a multi-cancer screening model integrating gene data, and when the program for the method of constructing a multi-cancer screening model integrating gene data is executed by the processor, the following steps are implemented: Obtain gene sequencing data of a multi-source sequencing platform for a target cancer population, and fuse the gene sequencing data of the multi-source sequencing platform to generate a gene fusion matrix; Determine the overlap between the gene expression characteristics of each target cancer and the normal gene expression characteristics according to the gene fusion matrix, mask the overlapping gene segments, and construct a gene expression topological map for different cancers; Import the gene expression topological map into a graph neural network for training to construct a multi-cancer screening model; Determine the gene mutation positions of the target cancer according to the gene sequencing data, and perform clustering analysis on the gene mutation positions by a density clustering algorithm to identify the gene mutation hot regions of each target cancer; Obtain the gene data screening load of the multi-cancer screening model, and if the screening load exceeds a preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the gene mutation hot regions.
[0014] The present invention discloses a method and a system for constructing a multi-cancer screening model integrating gene data. The method fuses gene sequencing data of a multi-source sequencing platform to generate a gene fusion matrix, analyzes the expression overlap regions between cancers and normal samples, masks them, and constructs a gene expression topological map. It uses a graph neural network for training to establish a multi-cancer screening model; at the same time, it identifies gene mutation hot regions through density clustering. When the screening load of the model exceeds the threshold, the screening weights of gene fragments are adjusted according to the mutation hot regions to improve the screening efficiency and accuracy. This method is applicable to the early detection and individualized screening of multiple cancers. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Shows a flowchart of a method for constructing a multi-cancer screening model integrating gene data according to the present invention; Figure 2 Shows a flowchart of constructing a multi-cancer screening model according to the present invention; Figure 3 Shows a flowchart of adjusting the screening weights of the multi-cancer screening model for different gene fragments according to the present invention; Figure 4 Shows a block diagram of a system for constructing a multi-cancer screening model integrating gene data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0017] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0018] Figure 1 The flowchart of a method for constructing a multi-cancer screening model integrating gene data of the present invention is shown.
[0019] As Figure 1 shown, in the first aspect of the present invention, a method for constructing a multi-cancer screening model integrating gene data is provided, including: S102, obtaining gene sequencing data of a multi-source sequencing platform for a target cancer population, fusing the gene sequencing data of the multi-source sequencing platform to generate a gene fusion matrix; S104, determining the overlap between the gene expression characteristics of each target cancer and the normal gene expression characteristics according to the gene fusion matrix, masking the overlapping gene segments, and constructing a gene expression topological map for different cancers; S106, importing the gene expression topological map into a graph neural network for training to construct a multi-cancer screening model; S108, determining the gene mutation positions of the target cancer according to the gene sequencing data, performing clustering analysis on the gene mutation positions by a density clustering algorithm, and identifying the gene mutation hot spots of each target cancer; S110, obtaining the gene data screening load situation of the multi-cancer screening model. If the screening load exceeds a preset load threshold, adjusting the screening weights of the multi-cancer screening model for different gene fragments according to the gene mutation hot spots.
[0020] It should be noted that by fusing gene data from multiple sequencing platforms to generate a gene fusion matrix, technical biases and data distribution differences between different sequencing platforms can be effectively eliminated, the compatibility and integrated analysis ability of cross-platform data can be improved, and a unified high-quality data foundation for multi-cancer screening can be provided. Secondly, by identifying and masking the overlapping segments of cancer types and normal gene expressions, non-specific interference regions can be excluded, the focusing ability of the topological map on cancer type-specific variation characteristics can be enhanced, and the sensitivity and discrimination of the model to early canceration signals can be improved. Then, by training the gene expression topological map using graph neural networks, the spatial correlation and mutation propagation paths between gene segments can be fully explored, dynamic modeling of complex expression patterns of multiple cancer types can be achieved, and the accuracy and generalization performance of multi-classification screening can be significantly improved. Subsequently, by identifying gene mutation hotspots through density clustering, the spatial aggregation characteristics of high-frequency mutation sites can be accurately located, the core regions of cancer type-specific driver mutations can be revealed, and key site bases for targeted screening and molecular typing can be provided. Finally, by dynamically adjusting the screening weights of gene fragments according to the model load, gene data related to mutation hotspots can be preferentially processed in high-concurrency scenarios, the allocation efficiency of computing resources can be optimized, and the dynamic balance between screening speed and accuracy can be achieved to meet the real-time requirements of large-scale clinical screening.
[0021] According to an embodiment of the present invention, the obtaining of gene sequencing data of multiple sequencing platforms for a target cancer population and the fusion of the gene sequencing data of the multiple sequencing platforms to generate a gene fusion matrix are specifically as follows: Obtain gene sequencing data of multiple sequencing platforms for a target cancer population within a preset time period, and calculate the gene expression statistics of the gene sequencing data of each sequencing platform, where the gene expression statistics include the mean, variance, skewness, and kurtosis of the gene sequencing data; Draw a kernel density estimation curve of the gene expression level according to the gene expression statistics, fit the kernel density estimation curve based on the maximum likelihood estimation method, and determine the gene expression level distribution of each cancer type in each sequencing platform; Judge the difference in gene expression level distribution of the same cancer type in different sequencing platforms according to the gene expression level distribution. If the distribution difference is greater than a preset distribution difference value, for the same cancer type, randomly select a sequencing platform as a reference platform based on the quantile matching algorithm, and the remaining sequencing platforms are used as target sequencing platforms, and a quantile space is constructed based on the gene expression level distribution of the gene sequencing data of the reference platform; Calculate the quantile interval of the quantile space based on the gene expression level distribution of the reference platform, map the gene sequencing data of the target sequencing platform to the quantile space of the reference platform, and perform quantile interpolation on the gene expression level of the target sequencing platform through a linear interpolation algorithm to generate a mapping function corresponding to the quantile space of the reference platform; Perform a non-linear transformation on the gene expression levels of the target sequencing platform according to the mapping function, adjust the deviation of the gene expression levels of the target sequencing platform from the reference platform in terms of distribution skewness and kurtosis, and generate normalized gene expression data; Align the gene fragments of different sequencing platforms for the same cancer type according to the normalized gene expression data to generate a gene fusion matrix.
[0022] It should be noted that during the multi-cancer screening process, due to technical principle differences (such as sequencing depth, library preparation methods) among different sequencing platforms (such as Illumina, BGI, etc.), there are systematic deviations in the distribution patterns (such as skewness, kurtosis) of gene expression data for the same cancer type. This inter-platform heterogeneity can lead to incomparability of cross-platform data, causing false positive or false negative risks in gene fusion analysis and affecting the reliability of multi-cancer model screening. By mapping the data of the target sequencing platform to the quantile space of the reference platform through the quantile matching algorithm, the differences in distribution patterns between platforms can be eliminated while retaining the relative ranking relationship of gene expression. Specifically, the quantile space constructed based on the reference platform defines a standardized statistical distribution framework, and the data of the target platform is aligned by linear interpolation to make the expression levels of data from different platforms comparable within the same quantile interval. At the same time, non-linear transformation corrects the skewness (eliminating the one-sided tail phenomenon) and kurtosis (adjusting the sharpness of the distribution) of the target platform data, making the normalized data consistent with the reference platform in statistical characteristics. This process not only eliminates the baseline shift between platforms but also retains the cancer type-specific expression patterns, and finally achieves precise spatial matching of cross-platform gene fragments through chromosome coordinate alignment to generate a highly consistent fusion matrix. The target cancer types include multiple cancer types; the gene sequencing data includes whole-genome sequencing and RNA sequencing data; the gene expression level distribution refers to the statistical distribution characteristics presented by the expression levels (such as FPKM / TPM values) measured for genomic regions in different samples, including parameters such as mean, variance, skewness, and kurtosis, reflecting the dynamic range of gene expression.
[0023] According to an embodiment of the present invention, determining the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics based on the gene fusion matrix, and masking the overlapping gene segments to construct a gene expression topological map for different cancer types, specifically: Extract the gene expression feature vectors of the target cancer type samples based on the gene fusion matrix. The gene expression feature vectors segment the gene sequence through a sliding window, calculate the mean value, variance of the expression levels of all gene fragments within each gene segment, and the change relationship of the expression level gradient between adjacent genes, and fuse the chromosome position information of the gene segments to generate a multi-dimensional feature representation; Synchronously extract the gene expression feature vectors of normal samples, construct a normal gene expression feature benchmark set, and calculate the Mahalanobis distance between the target cancer type gene expression feature vector and the normal gene expression feature vectors in the benchmark set according to the spatial distribution; Calculate the feature similarity between each gene segment according to the Mahalanobis distance, generate a gene segment similarity score, and perform a binary judgment on the gene segment similarity score according to a preset similarity threshold, and screen out the gene segments with scores higher than the similarity threshold as overlapping gene segments; Replace the gene expression data corresponding to the overlapping gene segments in the gene fusion matrix with random noise masks, and perform Gaussian smoothing interpolation on the gene expression data at the mask boundaries; Divide the gene fusion matrix after mask processing into continuous gene segments according to a preset window, and extract the node feature vectors of each gene segment, where the node feature vectors include gene expression statistics and mutation frequency information; Construct an adjacency relationship matrix of gene segments based on the chromosomal position information of gene segments, and perform dimensionality reduction processing on the spatial correlation of gene segments through a graph embedding algorithm according to the adjacency relationship matrix and node feature vectors to generate a gene expression topological map.
[0024] It should be noted that in cancer type screening, the gene expression features of cancerous samples often have some overlapping segments with normal samples (such as housekeeping genes or low-variation regions). These non-specific expressions will mask the abnormal signals of cancer type driver genes, making it difficult to focus on key variant regions during the training of multi-cancer type screening models and reducing the screening specificity. By extracting multi-dimensional features (mean, variance, gradient change, and chromosomal position) of gene segments through a sliding window and combining with the calculation of the Mahalanobis distance of the normal sample benchmark set, the similarity between cancer types and normal expressions can be accurately quantified, and overlapping segments (such as conserved functional regions) can be screened out and replaced with noise masks to eliminate their interference with the topological structure; the Gaussian smoothing processing at the mask boundaries can avoid the mutation jumps of gene expression data and maintain chromosomal continuity; the graph embedding algorithm (such as Node2Vec) based on the adjacency relationship matrix encodes the spatial correlation of gene segments (such as co-expression, cis-regulation) into a low-dimensional topological map, retaining the propagation path of cancer type specific mutation clusters (such as the linkage variation pattern of the EGFR gene cluster), enabling the model to distinguish subtle cancerous expression patterns, improving the screening sensitivity and reducing false positives.
[0025] Figure 2 Shows the flowchart of constructing a multi-cancer type screening model according to the present invention.
[0026] According to an embodiment of the present invention, the importing the gene expression topological map into a graph neural network for training to construct a multi-cancer type screening model is specifically as follows: S202, input the node feature vectors and adjacency relation matrices of the gene expression topological map into a graph convolutional network with a preset number of layers, and update the node embedding representations of each gene segment by aggregating the feature information of adjacent nodes; S204, introduce a residual connection mechanism between multiple graph convolutional layers, adjust the update weights of the node feature vectors by combining the mutation frequency information of the gene segments, and adopt a graph attention mechanism to adaptively learn the association strength between nodes, so as to generate a gene expression topological embedding representation with spatial dependence; S206, map the updated node feature vectors to graph-level feature vectors through a global max pooling layer, input the graph-level feature vectors into a fully connected classifier, and output the predicted probability distribution of the target sample belonging to different cancer types; S208, calculate the cross-entropy loss function according to the predicted probability distribution and the true cancer type label, and use the backpropagation algorithm to iteratively optimize the parameters of the graph neural network to obtain a trained multi-cancer screening model.
[0027] It should be noted that by aggregating the feature information of adjacent nodes of gene segments layer by layer through the graph convolutional network, the co-expression rules and mutation propagation paths in the spatially adjacent regions of chromosomes can be deeply mined, enhancing the model's ability to model the spatial association of local cancer signals; the residual connection mechanism effectively alleviates the gradient attenuation problem in the training of deep networks. Combining the dynamic weighting strategy of mutation frequency, the feature representations of highly variable gene segments maintain significant weights during multi-layer transmission, strengthening the model's focusing ability on driver mutation sites; the graph attention mechanism accurately quantifies the influence weight of long-range regulatory interactions across chromosome segments on cancer type classification by adaptively learning the association strength coefficient between nodes, improving the interpretability of complex expression patterns; the global max pooling layer compresses the node-level information into a global expression fingerprint at the map level by screening the significant features of each gene segment, retaining the biomarker combination with the strongest discriminability across cancer types; the fully connected classifier constructs a high-dimensional non-linear decision boundary based on the multi-level feature fusion results to realize the differential prediction of the multi-cancer probability distribution; through the cross-entropy loss function and backpropagation optimization, the model synchronously learns the common and specific expression rules among multiple cancer types during the training process, and finally, while maintaining high generalization, realizes the improvement of the accuracy of cross-cancer screening and the reduction of the misjudgment rate between categories.
[0028] According to an embodiment of the present invention, determining the gene mutation positions of the target cancer type according to the gene sequencing data, and performing clustering analysis on the gene mutation positions through a density clustering algorithm to identify the gene mutation hot regions of each target cancer type specifically includes: Extract the chromosome coordinate information of all gene mutation sites based on the gene sequencing data of multiple samples of each target cancer type, and construct a gene mutation position matrix containing sample ID, chromosome number, start site, and end site according to the chromosome number and base position; The density clustering algorithm is introduced to calculate the spatial distribution density of mutation sites among different samples of the target cancer type. The neighborhood radius parameter and the minimum sample number threshold of the density clustering algorithm are set. Taking each mutation site as the center, the number of mutant samples contained within its neighborhood radius is calculated. If the number of neighborhood samples of a certain mutation site exceeds the minimum sample number threshold, it is marked as a core point; Taking the core points as seed points for region expansion, all mutation sites within the neighborhood of the core points are traversed and their reachability is calculated, and the mutation sites that are directly density-reachable are grouped into the same clustering cluster; The region expansion process is iteratively executed until no new associated mutation sites can be added, the clustering clusters are output, and the gene mutation situation of each clustering cluster is determined to obtain the clustering result; According to the clustering result, the gene mutation hot regions of each target cancer type are determined.
[0029] It should be noted that through the density clustering algorithm to perform spatial density analysis on the gene mutation sites of the target cancer type, it can effectively filter out low-frequency noise mutations and identify mutation hot regions that are highly aggregated across samples. Screening core points based on the neighborhood radius and the minimum sample number threshold can exclude the interference of randomly scattered mutations and accurately locate the driver mutation clusters that co-occur in multiple samples. Through region expansion and reachability calculation, mutation sites that are spatially adjacent and density-connected are merged into the same clustering cluster, which can capture mutation enrichment regions that are continuously or discretely distributed on the chromosome and reveal the distribution law of cancer type-specific mutations. The mutation hot map generated based on the clustering result can guide the screening model to focus on the driver gene enrichment region preferentially, reduce redundant calculations for non-critical regions, and improve the screening efficiency.
[0030] Figure 3 The flowchart showing the adjustment of the screening weights of different gene fragments by the multi-cancer screening model of the present invention is shown.
[0031] According to an embodiment of the present invention, the gene data screening load situation of the multi-cancer screening model is obtained. If the screening load exceeds the preset load threshold, the screening weights of different gene fragments of the multi-cancer screening model are adjusted according to the gene mutation hot region, specifically: S302, the gene data screening queue length and the screening response time of a single gene fragment of the multi-cancer screening model are monitored in real time, and the current screening load degree of the multi-cancer screening model is determined according to the screening queue length and the screening response time; S304, if the current screening load degree is greater than the preset load threshold, the screening weights of different gene fragments of the multi-cancer screening model are adjusted according to the mutation hot region.
[0032] It should be noted that in the multi-cancer screening scenario, with the rapid increase in gene data volume and the improvement of real-time screening requirements, the model often faces the computational resource bottleneck under high-concurrency tasks, resulting in screening delays or reduced efficiency. The traditional static weight allocation strategy cannot adapt to dynamic load changes, easily causing insufficient analysis of key mutation regions or wasting resources in non-critical regions, affecting the timeliness and accuracy of screening. By real-time monitoring the screening queue length and response time to evaluate the model load, it is possible to dynamically adjust the screening weights of gene fragments when the load exceeds the limit, and preferentially allocate computing resources to mutation hotspots (such as high-frequency driver gene clusters). This strategy significantly improves the throughput efficiency of the model in high-load scenarios and shortens the screening response time for hotspots; at the same time, by focusing on key mutation sites, it retains the analysis accuracy of core cancer signals, avoids the risk of feature omission caused by resource contention, and achieves a dynamic balance between screening speed and accuracy.
[0033] According to an embodiment of the present invention, in the process of masking the gene fusion matrix, it further includes: Extract the local gradient vector field of gene expression levels on both sides of the mask boundary through a sliding window, and calculate the gradient direction consistency coefficient to detect the gradient direction reversal of adjacent unmasked segments; When a reversal is detected, use a direction-constrained Gaussian kernel function to adjust the interpolation weight distribution, generate a gradient-preserving smooth transition zone, and mark the phase conflict region as a secondary masked area; Based on the dynamic mask expansion algorithm, perform secondary noise replacement and boundary expansion on the secondary masked area; In the stage of constructing the gene expression topological map, dynamically adjust the boundary offset of the gene segment division window according to the chromosomal coordinates of the secondary masked area, and reconstruct the chromosomal position encoding of the node feature vector; Synchronously add dynamic connection edges between the secondary masked area and the original masked area in the adjacency relationship matrix, and perform joint representation learning on the topological association between the extended masked area and the normal segment through the graph embedding algorithm to generate a gene expression topological map resistant to phase interference.
[0034] It should be noted that in the gene data masking method, when directly performing noise substitution and Gaussian interpolation on overlapping gene segments, due to the lack of consideration of the consistency of the gene expression gradient direction in adjacent unmasked segments, the interpolation operation at the boundary is prone to cause phase conflicts in gene expression levels (such as gradient direction reversal), thereby destroying the continuity of gene expression in the chromosomal space and affecting the construction accuracy of subsequent topological maps and the reliability of model screening. Through the gradient direction consistency detection and dynamic mask extension mechanism, a Gaussian interpolation kernel function with direction constraints is introduced in the mask processing stage to accurately identify and reprocess the phase conflict regions. At the same time, the gene segment division rules and topological connection relationships are reconstructed to achieve a smooth transition of the gradient between the masked area and the normal segment and the dynamic coupling of spatial topological associations. This solution not only effectively eliminates the distortion of gene expression patterns caused by phase conflicts, but also enhances the model's ability to identify gene mutation hotspots and cancer type characteristics through an anti-interference topological map, significantly improving the classification robustness and screening accuracy of the multi-cancer type screening model in complex gene expression scenarios.
[0035] Figure 4 Fig. shows a block diagram of a multi-cancer type screening model construction system for fusing gene data according to the present invention.
[0036] In a second aspect of the present invention, there is also provided a multi-cancer type screening model construction system 4 for fusing gene data, the system includes: a memory 41, a processor 42, and a multi-cancer type screening model construction method program is included in the memory. When the multi-cancer type screening model construction method program is executed by the processor, the following steps are implemented: Obtain gene sequencing data of a multi-source sequencing platform for a target cancer type population, and fuse the gene sequencing data of the multi-source sequencing platform to generate a gene fusion matrix; Determine the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics according to the gene fusion matrix, mask the overlapping gene segments, and construct gene expression topological maps for different cancer types; Import the gene expression topological map into a graph neural network for training to construct a multi-cancer type screening model; Determine the gene mutation positions of the target cancer type according to the gene sequencing data, and perform clustering analysis on the gene mutation positions through a density clustering algorithm to identify the gene mutation hotspot regions of each target cancer type; Obtain the gene data screening load situation of the multi-cancer type screening model. If the screening load exceeds a preset load threshold, adjust the screening weights of the multi-cancer type screening model for different gene fragments according to the gene mutation hotspot regions.
[0037] The present invention discloses a method and system for constructing a multi-cancer screening model by integrating gene data. This method generates a gene fusion matrix by integrating gene sequencing data from multiple sequencing platforms, analyzes the expression overlap regions between cancer types and normal samples, and constructs a gene expression topological map after masking them. It uses a graph neural network for training to establish a multi-cancer screening model; at the same time, it identifies hotspots of gene mutations through density clustering. When the screening load of the model exceeds the threshold, the screening weights of gene fragments are adjusted according to the mutation hotspots to improve the screening efficiency and accuracy. This method is applicable to the early detection and individualized screening of multiple cancer types.
[0038] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0039] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0040] In addition, in each embodiment of the present invention, the various functional units can all be integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0041] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media that can store program codes such as mobile storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs.
[0042] Alternatively, if the above integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0043] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for constructing a multi-cancer screening model integrating gene data, characterized in that, It includes the following steps: Obtain the gene sequencing data of the multi-source sequencing platform for the population of the target cancer type, fuse the gene sequencing data of the multi-source sequencing platform, and generate a gene fusion matrix; Determine the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics according to the gene fusion matrix, mask the overlapping gene segments, and construct the gene expression topological maps of different cancer types; Import the gene expression topological maps into a graph neural network for training to construct a multi-cancer screening model; Determine the gene mutation positions of the target cancer type according to the gene sequencing data, perform clustering analysis on the gene mutation positions by the density clustering algorithm, and identify the gene mutation hot spots of each target cancer type; Obtain the gene data screening load of the multi-cancer screening model. If the screening load exceeds the preset load threshold, adjust the screening weights of different gene segments of the multi-cancer screening model according to the gene mutation hot spots; 2. The method for constructing a multi-cancer screening model integrating gene data according to claim 1, wherein, The obtaining the gene sequencing data of the multi-source sequencing platform for the population of the target cancer type, fusing the gene sequencing data of the multi-source sequencing platform, and generating a gene fusion matrix is specifically as follows: Obtain the gene sequencing data of the multi-source sequencing platform for the population of the target cancer type within a preset time period, and calculate the gene expression statistics of the gene sequencing data of each sequencing platform. The gene expression statistics include the mean, variance, skewness, and kurtosis of the gene sequencing data; Draw a kernel density estimation curve of the gene expression level according to the gene expression statistics, fit the kernel density estimation curve based on the maximum likelihood estimation method, and determine the gene expression level distribution of each cancer type in each sequencing platform; Judge the difference in the gene expression level distribution of the same cancer type in different sequencing platforms according to the gene expression level distribution. If the distribution difference is greater than the preset distribution difference value, for the same cancer type, randomly select a sequencing platform as the reference platform based on the quantile matching algorithm, and the remaining sequencing platforms are used as the target sequencing platforms, and construct a quantile space based on the gene expression level distribution of the gene sequencing data of the reference platform; Calculate the quantile intervals of the quantile space based on the gene expression level distribution of the reference platform, map the gene sequencing data of the target sequencing platform to the quantile space of the reference platform, and perform quantile interpolation on the gene expression level of the target sequencing platform through the linear interpolation algorithm to generate a mapping function corresponding to the quantile space of the reference platform; Perform a non-linear transformation on the gene expression level of the target sequencing platform according to the mapping function to adjust the deviation of the gene expression level of the target sequencing platform from the reference platform in terms of distribution skewness and kurtosis, and generate the normalized gene expression data; Align the spatial coordinates of the gene segments of the same cancer type from different sequencing platforms according to the normalized gene expression data to generate a gene fusion matrix.
3. The method for constructing a multi-cancer screening model integrating gene data according to claim 1, wherein The determining the overlap between the gene expression characteristics of each target cancer type and the normal gene expression characteristics according to the gene fusion matrix, masking the overlapping gene segments, and constructing the gene expression topological maps of different cancer types is specifically as follows: Extract the gene expression feature vector of the target cancer sample based on the gene fusion matrix. The gene expression feature vector segments the gene sequence through a sliding window, calculates the mean value, variance of the expression levels of all gene fragments within each gene segment, and the relationship of the gradient change in expression levels between adjacent genes, and fuses the chromosomal position information of the gene segment to generate a multi-dimensional feature representation; Simultaneously extract the gene expression feature vector of the normal sample, construct a normal gene expression feature benchmark set, and calculate the Mahalanobis distance between the target cancer gene expression feature vector and the normal gene expression feature vector in the benchmark set according to the spatial distribution; Calculate the feature similarity between each gene segment according to the Mahalanobis distance, generate a gene segment similarity score, and perform a binary judgment on the gene segment similarity score according to a preset similarity threshold, and screen out the gene segments with scores higher than the similarity threshold as overlapping gene segments; Replace the gene expression data corresponding to the overlapping gene segments in the gene fusion matrix with random noise masks, and perform Gaussian smoothing interpolation on the gene expression data at the mask boundary; Divide the gene fusion matrix after mask processing into continuous gene segments according to a preset window, and extract the node feature vector of each gene segment. The node feature vector includes gene expression statistics and mutation frequency information; Construct an adjacency relationship matrix of the gene segments based on the chromosomal position information of the gene segments, and perform dimensionality reduction processing on the spatial correlation of the gene segments through a graph embedding algorithm according to the adjacency relationship matrix and the node feature vector to generate a gene expression topological map; 4. The method for constructing a multi-cancer screening model integrating gene data according to claim 1, characterized in that Import the gene expression topological map into a graph neural network for training to construct a multi-cancer screening model. Specifically: Input the node feature vector and the adjacency relationship matrix of the gene expression topological map into a graph convolutional network with a preset number of layers, and update the node embedding representation of each gene segment by aggregating the feature information of adjacent nodes; Introduce a residual connection mechanism between multiple graph convolutional layers, adjust the update weight of the node feature vector in combination with the mutation frequency information of the gene segment, and adopt a graph attention mechanism to adaptively learn the association strength between nodes to generate a gene expression topological embedding representation with spatial dependence; Map the updated node feature vector to a graph-level feature vector through a global max pooling layer, input the graph-level feature vector into a fully connected classifier, and output the predicted probability distribution of the target sample belonging to different cancer types; Calculate the cross-entropy loss function according to the predicted probability distribution and the true cancer type label, and use the backpropagation algorithm to iteratively optimize the parameters of the graph neural network to obtain a trained multi-cancer screening model.
5. A method for constructing a multi-cancer screening model integrating gene data according to claim 1, characterized in that, Determine the gene mutation position of the target cancer according to the gene sequencing data, and perform clustering analysis on the gene mutation position through a density clustering algorithm to identify the gene mutation hot spot area of each target cancer. Specifically: Extract the chromosomal coordinate information of all gene mutation sites based on the gene sequencing data of multiple samples of each target cancer, and construct a gene mutation position matrix containing sample ID, chromosome number, start site, and end site according to the chromosome number and base position; The density clustering algorithm is introduced to calculate the spatial distribution density of mutation sites among different samples of the target cancer type. The neighborhood radius parameter and the minimum sample number threshold of the density clustering algorithm are set. Taking each mutation site as the center, the number of mutant samples contained within its neighborhood radius is calculated. If the number of neighborhood samples of a certain mutation site exceeds the minimum sample number threshold, it is marked as a core point; Taking the core points as seed points for region expansion, all mutation sites within the neighborhood of the core points are traversed and their reachability is calculated, and the mutation sites that are directly density-reachable are grouped into the same clustering cluster; The region expansion process is iteratively executed until no new associated mutation sites can be added, the clustering clusters are output, and the gene mutation situation of each clustering cluster is determined to obtain the clustering result; According to the clustering result, the gene mutation hot regions of each target cancer type are determined.
6. The method for constructing a multi-cancer screening model integrating gene data according to claim 1, wherein Obtain the gene data screening load situation of the multi-cancer screening model. If the screening load exceeds the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the gene mutation hot regions, specifically: Real-time monitor the gene data screening queue length and the screening response time of a single gene fragment of the multi-cancer screening model, and determine the current screening load level of the multi-cancer screening model according to the screening queue length and the screening response time; If the current screening load level is greater than the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the mutation hot regions.
7. A multi-cancer screening model construction system integrating gene data, characterized in that, The multi-cancer screening model construction system integrating gene data includes a storage and a processor. The storage includes a program for the method of constructing a multi-cancer screening model integrating gene data. When the program for the method of constructing a multi-cancer screening model integrating gene data is executed by the processor, the following steps are implemented: Obtain the gene sequencing data of the multi-source sequencing platform for the target cancer population, and fuse the gene sequencing data of the multi-source sequencing platform to generate a gene fusion matrix; According to the gene fusion matrix, determine the overlapping situation between the gene expression characteristics of each target cancer type and the normal gene expression characteristics, and mask the overlapping gene segments to construct the gene expression topological maps of different cancer types; Import the gene expression topological maps into a graph neural network for training to construct a multi-cancer screening model; Determine the gene mutation positions of the target cancer types according to the gene sequencing data, and perform clustering analysis on the gene mutation positions through the density clustering algorithm to identify the gene mutation hot regions of each target cancer type; Obtain the gene data screening load situation of the multi-cancer screening model. If the screening load exceeds the preset load threshold, adjust the screening weights of the multi-cancer screening model for different gene fragments according to the gene mutation hot regions.