Intelligent personalized colorectal cancer screening method and system based on multi-source fusion big data model
By constructing an individualized risk knowledge graph through a multi-source fusion big data model, a personalized colorectal cancer screening plan is generated, which solves the problem that traditional screening methods cannot take into account individual differences, and realizes the effectiveness and dynamic adaptability of personalized screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUDAN UNIVERSITY
- Filing Date
- 2026-02-27
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional colorectal cancer screening methods lack personalization and fail to fully consider individual differences, resulting in some high-risk individuals not being screened in a timely manner, while low-risk individuals undergo unnecessary over-screening, leading to a waste of medical resources and inconvenience to patients.
Based on a multi-source fusion big data model, by acquiring the full-dimensional health data of the target individual, generating an overall health representation vector, constructing an individualized risk knowledge graph, simulating the state transition process of health status nodes, generating a personalized colorectal cancer screening plan, and updating the model parameters based on clinical feedback.
It enables personalized colorectal cancer screening, improves the effectiveness and targeting of screening, dynamically reflects changes in individual disease risk, and continuously optimizes itself to adapt to updates in health status and medical knowledge.
Smart Images

Figure CN121747989B_ABST
Abstract
Description
A Personalized Intelligent Screening Method and System for Colorectal Cancer Based on Multi-Source Fusion Big Data Model Technical Field
[0001] This invention relates to the field of smart healthcare big data model technology, and more specifically, to a personalized intelligent screening method and system for colorectal cancer based on a multi-source fusion big data model. Background Technology
[0002] In the field of colorectal cancer screening, traditional methods mainly rely on single screening methods and universal screening strategies. For example, common methods such as fecal occult blood tests and colonoscopies are often based on general medical experience and statistical data, applying roughly the same screening protocols to all individuals.
[0003] However, there are significant differences between individuals, including genetic makeup, lifestyle habits, and medical history. A single screening method cannot comprehensively and accurately reflect an individual's true risk of disease. For example, gene sequencing results may show that certain gene mutations are highly associated with colorectal cancer, but traditional screening methods may not fully consider this factor; subtle lesions found in medical imaging slides may also be overlooked without comprehensive analysis in conjunction with other health information.
[0004] Meanwhile, generic screening strategies lack personalized consideration. Individuals of different ages, genders, and family histories have varying risks of developing colorectal cancer and different rates of disease progression. However, traditional screening programs often fail to fine-tune these differences, resulting in some high-risk individuals not receiving timely and effective screening, while some low-risk individuals may undergo unnecessary over-screening, leading to a waste of medical resources and inconvenience for patients. Summary of the Invention
[0005] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model, the method comprising:
[0006] Acquire a comprehensive health data set for the target individual. This comprehensive health data set includes gene sequencing fragment units, medical image slice units, clinical test indicator units, and electronic health record text units from different medical information systems.
[0007] Generate an overall health representation vector for the target individual based on a comprehensive health data set;
[0008] Based on the overall health representation vector and the risk association patterns in the pre-constructed group risk knowledge base, an individualized risk knowledge graph of the target individual is constructed by matching and association analysis. The individualized risk knowledge graph contains multiple health status nodes and risk transmission edges between health status nodes.
[0009] The individualized risk knowledge graph is input into a pre-trained feature evolution prediction model to simulate the state transition process of health status nodes under continuous time slices, thereby generating a dynamic risk feature vector for the target individual.
[0010] Based on the dynamic risk feature vector and the preset screening strategy rule base, a strategy mapping process is performed to generate a personalized colorectal cancer screening plan for the target individual. The personalized colorectal cancer screening plan includes the recommended screening technology type, the priority execution order of the screening technology type, the initial screening time point, and subsequent screening cycle parameters.
[0011] Based on the implementation feedback data and follow-up results of personalized colorectal cancer screening programs in clinical practice, the model parameters of the feature evolution prediction model and the risk association patterns in the population risk knowledge base are updated.
[0012] Furthermore, embodiments of the present invention also provide a personalized intelligent screening system for colorectal cancer based on a multi-source fusion big data model, comprising:
[0013] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model by executing the machine-executable instructions.
[0014] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of the colorectal cancer personalized intelligent screening system based on a multi-source fusion big data model reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the colorectal cancer personalized intelligent screening system based on a multi-source fusion big data model to perform the aforementioned colorectal cancer personalized intelligent screening method based on a multi-source fusion big data model.
[0015] Based on the above, by acquiring a comprehensive set of health data for the target individual, encompassing multi-source information such as gene sequencing fragments, medical image slices, clinical laboratory indicators, and electronic health record texts, the overall health representation vector generated from this comprehensive health data set can comprehensively reflect multiple aspects of an individual's health characteristics. Through matching and correlation analysis with a pre-constructed group risk knowledge base, a personalized risk knowledge graph is constructed, presenting the risk transmission relationship between individual health status nodes. The personalized risk knowledge graph is input into a feature evolution prediction model to simulate the state transition process of health status nodes under continuous time slices. The generated dynamic risk feature vector can reflect changes in an individual's disease risk in real time and dynamically. The personalized colorectal cancer screening plan generated based on the dynamic risk feature vector fully considers the unique circumstances of each individual, including the type of screening technology, execution sequence, initial time point, and subsequent screening cycle, achieving customization of the screening plan and improving the effectiveness and targeting of screening. Simultaneously, by updating model parameters and risk association patterns based on execution feedback data and follow-up results from clinical practice, it can continuously optimize and improve itself, adapting to changes in individual health status and updates in medical knowledge. Attached Figure Description
[0016] Figure 1 is a schematic diagram of the execution flow of the personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model provided in an embodiment of the present invention.
[0017] Figure 2 is a schematic diagram of exemplary hardware and software components of the colorectal cancer personalized intelligent screening system based on a multi-source fusion big data model provided in an embodiment of the present invention. Detailed Implementation
[0018] Figure 1 is a flowchart illustrating a personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model provided in an embodiment of the present invention, which will be described in detail below.
[0019] Step S110: Obtain the full-dimensional health data set of the target individual. The full-dimensional health data set includes gene sequencing fragment units, medical image slice units, clinical test indicator units, and electronic health record text units from different medical information systems.
[0020] In this embodiment, taking a target individual suspected of having colorectal cancer risk as an example, the hospital's internal medical data integration platform collects a comprehensive set of health data for the target individual from multiple independent medical information systems. Specifically, the gene sequencing fragment unit comes from the high-throughput sequencing system of a third-party gene testing institution, containing the target individual's whole exome sequencing data in FASTQ format, which includes the sequence information and corresponding quality value of each sequencing fragment; the medical image slice unit comes from the hospital's image archiving and communication system, containing the target individual's colonic endoscopy images and abdominal computed tomography images, which are in DICOM format, with each slice containing pixel matrix data and corresponding imaging parameter information; the clinical test indicator unit comes from the hospital's laboratory information system, containing the target individual's clinical test results for the past five years in HL7 standard format, covering the concentration data of tumor markers such as carcinoembryonic antigen and carbohydrate antigen, as well as the corresponding test timestamps; the electronic health record text unit comes from the hospital's electronic medical record system, containing the target individual's outpatient records, inpatient records, surgical records, and other text data in a mixed format of structured and unstructured text. During the data collection process, sensitive data involving personal privacy, such as names, ID numbers, and medical record numbers, are anonymized using a differential privacy technology. This involves adding noise values that conform to a Laplace distribution to the original data, so that the processed data retains its statistical analysis value while preventing the identification of specific individuals.
[0021] Step S120: Generate the overall health representation vector of the target individual based on the full-dimensional health data set.
[0022] After obtaining a comprehensive set of health data for a target individual, feature extraction and fusion processing of the multi-source heterogeneous data are required to generate a holistic health representation vector that fully reflects the individual's health status. This process involves feature processing and cross-modal fusion operations across multiple dimensions, effectively integrating multi-source information by converting different types of data, such as genes, images, clinical indicators, and text, into a unified feature representation space.
[0023] Step S121: Perform deep semantic encoding on gene sequencing fragment units, extract single nucleotide polymorphism site variation information and methylation level information of gene expression regulatory regions in gene fragments, and generate gene-dimensional feature vectors.
[0024] For FASTQ format data extracted from gene sequencing fragment units, the Burrows-Wheeler alignment tool was first used to align the sequencing fragments to the human reference genome, obtaining BAM format alignment results. Next, the HaplotypeCaller module in the GATK toolkit was used to detect single nucleotide polymorphism (SNP) sites, screening for variant sites located in the coding regions of colorectal cancer-related genes (such as APC, KRAS, TP53, etc.). Each variant site was represented by a binary feature indicating its presence or absence. Simultaneously, a methylation sequencing data analysis workflow was used to quantitatively analyze the methylation levels of gene expression regulatory regions (such as promoter regions and enhancer regions), calculating the methylation rate of each CpG site, i.e., the ratio of the number of methylated sequencing fragments to the total number of sequencing fragments. The single nucleotide polymorphism site variation information and methylation level information are arranged in the preset gene region order to form a high-dimensional gene feature matrix. Then, the matrix is reduced in dimension by principal component analysis, retaining the principal components with a cumulative contribution rate of more than 90%, and finally generating a gene dimension feature vector with dimension D1, where each element corresponds to a gene feature component after dimension reduction.
[0025] Step S122: Perform multi-scale structural analysis processing on the medical image slice unit to identify the texture regularity information of the colonic mucosa, the uniformity information of the colonic wall thickness, and the morphological contour information of the intestinal polyps in the medical image slice unit, and generate image dimension feature vectors.
[0026] For DICOM format image data in medical image slices, the images are first preprocessed, including grayscale normalization, contrast enhancement, and denoising. Then, a multi-scale Gaussian pyramid method is used to decompose the preprocessed images into multi-resolution features at different scales. At each scale, a texture analysis method based on the gray-level co-occurrence matrix is used to calculate the texture features of the colonic mucosa region, including indices such as energy, entropy, contrast, and correlation, to characterize texture regularity. For the uniformity of colonic wall thickness, an edge detection algorithm is used to extract the inner and outer edge contours of the colonic wall, and the distance from each point on the contour to the central axis is calculated as the thickness value. The standard deviation and coefficient of variation of these thickness values are then calculated as uniformity metrics. For the morphological contour information of intestinal polyps, a region-growing-based segmentation algorithm is used to segment the polyp region, and then morphological features such as area, perimeter, roundness, and aspect ratio of the circumscribed rectangle are extracted. The above texture features, thickness uniformity indices, and polyp morphological features are combined in a predetermined order to form a D2-dimensional image feature vector.
[0027] Step S123: Extract the time-series fluctuation pattern of the clinical test indicator unit, analyze the changing trend and periodic fluctuation information of the carcinoembryonic antigen concentration and carbohydrate antigen concentration in the historical time series of the clinical test indicator unit, and generate the feature vector of the clinical indicator dimension.
[0028] From the HL7 format data of the clinical laboratory indicator unit, the carcinoembryonic antigen (CEA) and carbohydrate antigen (CA) concentrations for the target individual over the past five years, along with the corresponding detection timestamps, were extracted. The data were arranged chronologically to construct two independent time series. For each time series, missing values were first imputed using linear interpolation. Then, a sliding window technique was used to segment the time series, calculating the mean, variance, maximum, and minimum values within each window. Next, a time series decomposition method was employed to decompose the original time series into trend, periodic, and residual components. The trend component reflects long-term trends and is obtained by fitting a linear or nonlinear trend line using the least squares method; the periodic component reflects periodic fluctuations and is extracted using Fourier transform analysis to obtain the main periods and their corresponding amplitudes and phases. The statistical features, trend component parameters, and periodic component parameters were combined to form a D3-dimensional feature vector for the clinical indicator.
[0029] Step S124: Perform natural language understanding processing on the electronic health record text unit, extract the description information of past medical history of intestinal diseases, family genetic disease history and personal lifestyle habits recorded in the electronic health record text unit, and generate text dimension feature vector.
[0030] For the mixed text data in the electronic health record text unit, text preprocessing is first performed, including word segmentation, stop word removal, and part-of-speech tagging. Then, a named entity recognition model based on a combination of bidirectional long short-term memory network and conditional random field is used to identify entities related to intestinal diseases from the text, such as disease names, symptom descriptions, and treatment methods. Next, a convolutional neural network based on an attention mechanism is used to extract features from the text, focusing on information such as past medical history of intestinal diseases (e.g., whether one has ulcerative colitis, Crohn's disease, etc.), family history of genetic diseases (e.g., whether a first-degree relative has colorectal cancer), and personal lifestyle habits (e.g., smoking history, drinking history, dietary habits, etc.). The identified entities and extracted features are converted into word vector representations, and then pooling operations are used to obtain fixed-length text feature vectors. Finally, the above feature vectors are concatenated to generate a text-dimensional feature vector of dimension D4.
[0031] Step S125: Input the gene dimension feature vector, image dimension feature vector, clinical indicator dimension feature vector and text dimension feature vector into the cross-modal feature interaction channel. The cross-modal feature interaction channel contains multiple parallel feature interaction sub-channels.
[0032] After obtaining the feature vectors from the gene dimension, image dimension, clinical indicator dimension, and text dimension, these four feature vectors from different modalities are input into the cross-modal feature interaction channel. This cross-modal feature interaction channel contains three parallel feature interaction sub-channels, each used to handle the interaction relationships between features from different modalities. Each sub-channel employs a different feature fusion strategy to fully capture the complex correlations between multimodal data.
[0033] Step S126: In each feature interaction sub-channel, the gene dimension feature vector, image dimension feature vector, clinical indicator dimension feature vector, and text dimension feature vector are standardized to obtain standardized feature vectors of each dimension. Based on the standardized feature vectors of each dimension, a bidirectional feature attention weighting operation is performed to calculate the feature semantic relevance score between any two feature vectors of different dimensions. The contribution weight ratio of different feature vectors in the feature interaction process is dynamically adjusted according to the feature semantic relevance score.
[0034] In each feature interaction sub-channel, the four-dimensional feature vectors of the input are first standardized. For each element in each feature vector, Z-Score standardization is applied, which involves subtracting the mean of the feature vector from the element's value and then dividing by the standard deviation, resulting in a standardized feature vector with a mean of 0 and a standard deviation of 1, thus eliminating the influence of dimensional differences. After standardization, a bidirectional feature attention weighting operation is performed. Taking gene-dimensional and image-dimensional feature vectors as an example, a cosine similarity matrix between the two feature vectors is first calculated, where each element represents the similarity between a feature component in the gene-dimensional feature vector and a feature component in the image-dimensional feature vector. Then, the similarity matrix is row-normalized and column-normalized to obtain an attention weight matrix. Based on the attention weight matrix, the gene-dimensional and image-dimensional feature vectors are weighted and summed separately to obtain the attention-weighted feature vector. A similar bidirectional attention weighting operation is applied to the interaction between any other two different-dimensional feature vectors. By calculating the semantic relevance score of the features, the contribution weight ratio of feature vectors of different dimensions in the feature interaction process is dynamically adjusted so that the feature components with high relevance receive higher weights.
[0035] Step S127: Perform weighted aggregation operation on the standardized feature vectors of each dimension according to the adjusted contribution weight ratio to generate a preliminary fused feature representation; perform redundant information compression operation on the preliminary fused feature representation to obtain a unified feature representation.
[0036] In each feature interaction sub-channel, a weighted aggregation operation is performed on the standardized feature vectors of each dimension based on the contribution weight ratio obtained in step S126. Specifically, each dimension feature vector is multiplied by its corresponding weight coefficient, and then feature concatenation is performed, that is, the feature vectors of the four dimensions are connected in order into a longer vector to generate a preliminary fused feature representation. The dimensions of the preliminary fused feature representation are D1+D2+D3+D4. Since there may be redundant information in the preliminary fused feature representation, a redundant information compression operation is required. Principal component analysis is used to reduce the dimensionality of the preliminary fused feature representation, retaining the principal components with a cumulative contribution rate of more than 95%, resulting in a unified feature representation with dimension D5.
[0037] Step S128: Input the unified feature representation into the multilayer perceptron network for nonlinear transformation and dimensionality increase operation to generate feature projections in a high-dimensional potential health feature space.
[0038] The unified feature representation obtained in step S127 is input into a multilayer perceptron network. This multilayer perceptron network contains three hidden layers: the first layer has twice the number of neurons as D5, the second layer has 1.5 times the number of neurons as the first layer, and the third layer has twice the number of neurons as the second layer. Each hidden layer is followed by a ReLU activation function to achieve a nonlinear transformation. Through forward propagation computation of the multilayer perceptron network, the unified feature representation is mapped from a low-dimensional space to a higher-dimensional potential health feature space, generating a feature projection of dimension D6, where D6 is greater than D5.
[0039] Step S129: Perform dimensionality reduction and normalization on the feature projection in the high-dimensional potential health feature space to obtain a fixed-dimensional overall health representation vector. Each dimension of the overall health representation vector corresponds to a comprehensive health status characterization factor that integrates multi-source data information.
[0040] The feature projections in the high-dimensional potential health feature space generated in step S128 are dimensionality-reduced using the t-distributed neighborhood embedding method, mapping them to a low-dimensional space with a fixed dimension of D7, where D7 is a pre-defined constant, such as 256. After dimensionality reduction, the resulting low-dimensional feature vectors are subjected to L2 normalization, ensuring that the feature vector's magnitude is 1 to eliminate scale differences between different samples. The final obtained overall health representation vector with a fixed dimension of D7 corresponds to a comprehensive health status characterization factor that integrates multi-source data information such as genes, images, clinical indicators, and text, comprehensively reflecting the health status of the target individual.
[0041] Step S130: Match and analyze the risk association patterns between the overall health representation vector and the pre-constructed group risk knowledge base to construct an individualized risk knowledge graph for the target individual. The individualized risk knowledge graph contains multiple health status nodes and risk transmission edges between health status nodes.
[0042] After obtaining the overall health representation vector of the target individual, it is necessary to match and analyze it with risk association patterns in a pre-constructed group risk knowledge base to build an individualized risk knowledge graph that reflects the individual's unique health risk relationships. The group risk knowledge base stores a large number of risk association patterns between health states obtained from large-scale population data mining. By matching the overall health representation vector of the target individual with these patterns, the potential health risk associations of the individual can be identified.
[0043] Step S131: Extract multiple preset dimensions of health element intensity values from the overall health representation vector. Each health element intensity value corresponds to a comprehensive health status characterization factor that integrates multi-source data information.
[0044] The overall health representation vector has 7 dimensions, with each dimension corresponding to a predefined health element. The value of each dimension is extracted from the overall health representation vector as the intensity value of the corresponding health element. For example, the first dimension might correspond to the health element "risk of abnormal colonic mucosal hyperplasia," with an intensity value of 0.75; the second dimension might correspond to the health element "risk of abnormal carcinoembryonic antigen levels," with an intensity value of 0.62, and so on. The intensity values of these health elements range from 0 to 1, with higher values indicating a higher degree of abnormality in the health element.
[0045] Step S132: Compare the intensity value of health elements in each dimension with the predefined benchmark health status range in the group risk knowledge base one by one, and identify abnormal health elements that exceed the corresponding benchmark health status range and their exceedance parameters.
[0046] The population risk knowledge base predefines a baseline health status range for each health element. For example, the baseline range for "risk of dysplasia of colonic mucosa" is 0 to 0.3, and the baseline range for "risk of abnormal carcinoembryonic antigen level" is 0 to 0.4. The intensity value of each health element extracted in step S131 is compared with its corresponding baseline health status range. If the intensity value exceeds this range, the health element is marked as an abnormal health element. Simultaneously, an exceedance parameter is calculated: (health element intensity value - upper limit of baseline range) / (1 - upper limit of baseline range). If the health element intensity value does not exceed the baseline range, the exceedance parameter is 0. For example, the intensity value of "risk of dysplasia of colonic mucosa" is 0.75, and the upper limit of the baseline range is 0.3. Therefore, the exceedance parameter is (0.75 - 0.3) / (1 - 0.3) = 0.45 / 0.7 ≈ 0.64.
[0047] Step S133: Based on the abnormal health elements and their exceedance parameters, traverse all risk association patterns in the group risk knowledge base and calculate the pattern matching score between each risk association pattern and the current overall health representation vector.
[0048] Each risk association pattern in the group risk knowledge base consists of a set of source health status nodes, target health status nodes, and risk transmission rules connecting them. For each risk association pattern, the health elements corresponding to the source health status nodes are first extracted. Then, the matching degree between the health element intensity values of these source health status nodes in the overall health representation vector of the target individual and the preset threshold of the source health status nodes in the risk association pattern is calculated. Specifically, for each source health status node, the ratio of its health element intensity value to the preset threshold is calculated. If the ratio is greater than 1, it is set to 1; otherwise, the ratio is taken. The matching degrees of all source health status nodes are multiplied together, and then multiplied by the weight coefficient corresponding to each source health status node (pre-set according to its importance in the risk association pattern) to obtain the preliminary matching score of the risk association pattern. Then, the excess amplitude parameter of abnormal health elements is used as a correction factor to correct the preliminary matching score to obtain the final pattern matching score. For example, a certain risk association pattern contains two source health status nodes with weight coefficients of 0.6 and 0.4, and preset thresholds of 0.5 and 0.3, respectively. The target individual's corresponding health element intensity values are 0.75 and 0.4, respectively, and the exceedance parameters are 0.64 and 0.25, respectively. Therefore, the initial matching score is (0.75 / 0.5)*0.6*(0.4 / 0.3)*0.4=1.5*0.6*1.333*0.4≈0.48. The correction factor is (1+0.64*0.5+0.25*0.5)=1+0.32+0.125=1.445, and the final pattern matching score is 0.48*1.445≈0.6936.
[0049] Step S134: Sort all risk association patterns in the group risk knowledge base according to the pattern matching score, and filter out multiple candidate risk association patterns whose pattern matching scores exceed the preset activation threshold.
[0050] The preset activation threshold is 0.5. All risk association patterns are sorted from highest to lowest according to their pattern matching scores. Then, risk association patterns with a pattern matching score greater than 0.5 are selected as candidate risk association patterns. For example, after sorting, there are 5 risk association patterns with pattern matching scores of 0.85, 0.72, 0.69, 0.55, and 0.48, respectively. The first 4 of these patterns have scores exceeding 0.5, so these 4 risk association patterns are selected as candidate risk association patterns.
[0051] Step S135: Parse the source health status node identifier, target health status node identifier, and risk transmission rule description connecting the source health status node identifier and the target health status node identifier contained in each candidate risk association pattern.
[0052] Each candidate risk association pattern has a specific structure, including source health status node identifiers, target health status node identifiers, and a risk transmission rule description. For example, the source health status node identifiers of a certain candidate risk association pattern are "SH1" (corresponding to "risk of abnormal colonic mucosal hyperplasia") and "SH2" (corresponding to "risk of abnormal carcinoembryonic antigen levels"), the target health status node identifier is "TH1" (corresponding to "risk of colorectal cancer occurrence"), and the risk transmission rule is described as "when the intensity value of SH1 is greater than 0.6 and the intensity value of SH2 is greater than 0.5, SH1 and SH2 together act on TH1, and the intensity of their influence is SH1 intensity value * 0.7 + SH2 intensity value * 0.3".
[0053] Step S136: Extract the current health element intensity value corresponding to the source health state node identifier from the overall health representation vector, and calculate the potential influence intensity value of the source health state node on the target health state node according to the risk transmission rule description.
[0054] Based on the source health status node identifiers in the candidate risk association patterns, the corresponding health element intensity values are extracted from the overall health representation vector of the target individual. For example, for source health status node identifiers "SH1" and "SH2", the extracted current health element intensity values are 0.75 and 0.62, respectively. Then, the potential impact intensity value is calculated according to the risk transmission rule description. According to the above rule description, the potential impact intensity value is 0.75*0.7+0.62*0.3=0.525+0.186=0.711.
[0055] Step S137: Calculate the status update value of the target health status node based on the potential impact intensity value and the current health element intensity value corresponding to the target health status node identifier.
[0056] The current health element intensity value corresponding to the target health status node identifier is extracted from the overall health representation vector. For example, the current intensity value of "TH1" is 0.45. Then, the potential impact intensity value and the current health element intensity value are fused to obtain the status update value. The fusion calculation adopts a weighted summation method, and the weight coefficient is set according to the confidence level of the risk transmission rule. For example, if the confidence level is 0.8, the status update value is current health element intensity value * (1-0.8) + potential impact intensity value * 0.8 = 0.45 * 0.2 + 0.711 * 0.8 = 0.09 + 0.5688 = 0.6588.
[0057] Step S138: Based on the source health status node identifier, the target health status node identifier, the current health element intensity value corresponding to the source health status node identifier, the potential impact intensity value, the current health element intensity value corresponding to the target health status node identifier, and the status update value, instantiate and generate the corresponding individual health status node and the individual risk transmission edge between nodes.
[0058] Based on the parameters extracted and calculated above, individual health status nodes and individual risk transmission edge individuals are instantiated. Each individual health status node includes attributes such as node identifier, current health element intensity value, and status update value. For example, the source health status node SH1 has the attributes "SH1" and a current intensity value of 0.75; the target health status node TH1 has the attributes "TH1", a current intensity value of 0.45, and a status update value of 0.6588. Individual risk transmission edge individuals connect the source health status node individuals and the target health status node individuals, and include attributes such as source node identifier, target node identifier, and potential impact intensity value. For example, the risk transmission edge individual connecting SH1 and TH1 has the attributes "SH1" as the source identifier, "TH1" as the target identifier, and an impact intensity value of 0.525 (0.75 * 0.7).
[0059] Step S139: Connect and combine all instantiated individual health status nodes and individual risk transmission edge nodes according to their topological relationships in the candidate risk association patterns to form a preliminary individualized risk knowledge graph that represents the target individual's unique health status association network.
[0060] All individual health status nodes and individual risk transmission edges instantiated from candidate risk association patterns are connected and combined according to their source-target relationships in each candidate risk association pattern. For example, if SH1 and SH2 in one candidate risk association pattern both point to TH1, and TH1 in another candidate risk association pattern points to TH2 (corresponding to "colorectal cancer progression risk"), then the above nodes and edges are connected according to the above topological relationship to form a preliminary individualized risk knowledge graph. This preliminary individualized risk knowledge graph uses nodes to represent health status and directed edges to represent risk transmission relationships.
[0061] Step S1310: Check whether there are isolated individual health status nodes in the preliminary individualized risk knowledge graph. If so, perform a secondary matching between the current health element intensity value of the isolated individual health status node and the benchmark health status range of other health status nodes in the group risk knowledge base, and attempt to supplement the missing individual risk transmission edge individuals.
[0062] The process iterates through all individual health status nodes in the preliminary individualized risk knowledge graph, checking for isolated nodes without incoming or outgoing edges. For example, a node SH3 (corresponding to "Gut Microbiota dysbiosis risk") is found, with a current health element strength value of 0.6, but it is not connected to any other nodes in the preliminary graph. In this case, a secondary matching is performed between SH3's current health element strength value and the baseline health status range of all other health status nodes in the group risk knowledge base to search for risk association patterns with SH3 as the source or target node. If a risk association pattern is found, where SH3 is the source node, TH1 is the target node, and the pattern matching score is 0.55 (exceeding the activation threshold of 0.5), then an individual risk transmission edge connecting SH3 and TH1 is instantiated and added to the preliminary individualized risk knowledge graph, eliminating isolated nodes.
[0063] Step S1311: Perform loop detection and redundant edge elimination operations on the supplemented preliminary individualized risk knowledge graph to ensure that at most one most representative individual risk transmission edge is retained between any two individual health status nodes, and finally form the structured individualized risk knowledge graph.
[0064] A depth-first search algorithm is used to perform loop detection on the supplemented preliminary individualized risk knowledge graph. If a loop is found, such as node A→node B→node C→node A, the edge with the highest influence strength value is retained, and other edges are deleted to break the loop. For cases where multiple risk transmission edges exist between any two nodes, a comprehensive score is calculated for each edge. This comprehensive score is obtained by weighting the influence strength value and the pattern matching score (weights of 0.7 and 0.3, respectively). The edge with the highest comprehensive score is retained as the most representative edge, and the remaining redundant edges are deleted. After loop detection and redundant edge elimination, the final structured individualized risk knowledge graph is obtained.
[0065] Step S140: Input the individualized risk knowledge graph into the pre-trained feature evolution prediction model to simulate the state transition process of health status nodes under continuous time slices, and generate the dynamic risk feature vector of the target individual.
[0066] The constructed individualized risk knowledge graph is input into a pre-trained feature evolution prediction model. This model can simulate the state transition process of health status nodes over a future period, thereby predicting the dynamic trend of health risk changes for the target individual and generating a dynamic risk feature vector. This dynamic risk feature vector contains information on the state evolution of health status nodes under different time slices.
[0067] Step S141: Extract the initial state feature representation of all healthy state nodes from the individualized risk knowledge graph and the connection strength parameter of the risk transmission edge between the healthy state nodes.
[0068] Each health status node in the personalized risk knowledge graph has its initial state feature representation, which is based on the state update value obtained in step S138. For example, the initial state feature representation of node TH1 is a vector of dimension D8, containing the state update value 0.6588 and encoded information of other relevant attributes. The connection strength parameter of the risk transmission edge is determined based on the potential influence strength value calculated in step S136; for example, the connection strength parameter of the risk transmission edge connecting SH1 and TH1 is 0.525. The initial state feature representations of all health status nodes and the connection strength parameters of the risk transmission edges are extracted and used as input data for the feature evolution prediction model.
[0069] Step S142: Input the initial state feature representation and connection strength parameters into the time-series slice division module of the feature evolution prediction model, and divide the future time window into multiple continuous time slice units according to the preset future time window length and time slice granularity.
[0070] The preset future time window length is T years, and the time slice granularity is t months, meaning each time slice unit is t months long. The time-series slicing module divides the future time window into N consecutive time slice units based on these two parameters, where N = T * 12 / t. For example, if the future time window length is 5 years and the time slice granularity is 3 months, then N = 5 * 12 / 3 = 20 time slice units.
[0071] Step S143: In the simulation module of the first time slice unit of the feature evolution prediction model, the initial state feature representation of the healthy state node is normalized, and the connection strength parameter of the risk transmission edge is normalized; based on the normalized initial state feature representation of the healthy state node and the normalized connection strength parameter of the risk transmission edge, the state transition probability distribution of each healthy state node at the end of the first time slice unit is calculated.
[0072] First, the initial state feature representations of all healthy nodes are normalized using the Min-Max normalization method, mapping the value of each feature component to between 0 and 1. Similarly, the connection strength parameters of all risk propagation edges are also normalized using Min-Max, mapping them to between 0 and 1. Then, the state transition probability distribution of each healthy node at the end of the first time slice is calculated.
[0073] Step S1431: Normalize the initial state feature representation of all healthy nodes and normalize the connection strength parameters of all risk propagation edges.
[0074] For each feature component in the initial state feature representation of a healthy node, the normalization formula is (feature value - minimum value) / (maximum value - minimum value), where the minimum and maximum values are the minimum and maximum values of that feature component in the training dataset. The connection strength parameter of a risk propagation edge is also processed using the same Min-Max normalization formula. For example, if the original value of a feature component in the initial state feature representation of a node is 0.6588, and the minimum value of this feature component in the training dataset is 0 and the maximum value is 1, then the normalized value is 0.6588. Similarly, if the original value of the connection strength parameter of a risk propagation edge is 0.525, and the minimum value in the training dataset is 0 and the maximum value is 1, then the normalized value is 0.525.
[0075] Step S1432: For each health state node in the individualized risk knowledge graph, identify all risk transmission edges that target the health state node. These risk transmission edges are called incoming edges.
[0076] Traverse each health status node in the individualized risk knowledge graph and find all risk propagation edges pointing to that node, i.e., incoming edges. For example, for node TH1, there may be incoming edges pointing to it from SH1, SH2, and SH3.
[0077] Step S1433: Obtain the normalized initial state feature representation of the source healthy state node corresponding to each incoming edge, and the normalized connection strength parameter of the incoming edge. Based on the normalized initial state feature representation of the source healthy state node corresponding to each incoming edge and the normalized connection strength parameter of the incoming edge, calculate the state transition tendency value of the incoming edge to the target healthy state node in the current time slice unit.
[0078] For each incoming edge, the first component of the normalized initial state feature representation of the source healthy node (assuming this component represents the core state value of the node) is multiplied by the normalized connection strength parameter of the incoming edge to obtain the state transition tendency value of the incoming edge for the target healthy node. For example, if the first component of the normalized initial state feature representation of the source node SH1 is 0.75 and the connection strength parameter of the incoming edge is 0.525, then the state transition tendency value is 0.75 * 0.525 = 0.39375.
[0079] Step S1434: Summarize the state transition tendency values generated by all incoming edges pointing to the target healthy state node to obtain the total state transition tendency value of the target healthy state node.
[0080] The total state transition tendency value of the target healthy node is obtained by summing the state transition tendency values of all incoming edges. For example, if node TH1 has three incoming edges with state transition tendency values of 0.39375, 0.2835 (state value of SH2 0.62 * connection strength 0.457), and 0.18 (state value of SH3 0.6 * connection strength 0.3), then the total state transition tendency value is 0.39375 + 0.2835 + 0.18 = 0.85725.
[0081] Step S1435: Define a set of new states that the target healthy state node may transition to within the first time slice unit, including the possibility of maintaining the current initial state feature representation; assign a basic transition probability parameter to each possible new state, the basic transition probability parameter being obtained based on prior knowledge in the group risk knowledge base.
[0082] The possible new states of a target healthy node include three types: state elevation, state maintenance, and state deterioration. Based on prior knowledge in the group risk knowledge base, basic transition probability parameters are assigned to these three states. For example, the basic probability of state elevation is 0.3, the basic probability of state maintenance is 0.5, and the basic probability of state deterioration is 0.2.
[0083] Step S1436: Convert the total state transition tendency value of the target healthy state node into a transition probability increment consistent with the additivity of the basic transition probability parameter through a preset scaling function. Then, according to a preset weighting ratio, add the converted transition probability increment to the basic transition probability parameter of each possible new state to obtain the unnormalized probability value of each possible new state.
[0084] The preset scaling function is f(x) = x * 0.5, which scales the total state transition propensity value to 0.5 times its original value, resulting in the transition probability increment. For example, if the total state transition propensity value is 0.85725, the transition probability increment is 0.85725 * 0.5 = 0.428625. The preset weighting ratio is 0.6 for state elevation, 0.3 for state maintenance, and 0.1 for state depreciation. The unnormalized probability of a state advancement is 0.3 + 0.428625 * 0.6 = 0.3 + 0.257175 = 0.557175; the unnormalized probability of a state maintenance is 0.5 + 0.428625 * 0.3 = 0.5 + 0.1285875 = 0.6285875; and the unnormalized probability of a state decline is 0.2 + 0.428625 * 0.1 = 0.2 + 0.0428625 = 0.2428625.
[0085] Step S1437: Input the unnormalized probability values of all possible new states into the probability normalization processing function. The probability normalization processing function converts the unnormalized probability values of all possible new states into standard probability values that sum to one, forming the state transition probability distribution of the healthy state node at the end of the first time slice unit.
[0086] The probability normalization function is to divide each unnormalized probability value by the sum of all unnormalized probability values. For example, the sum of the unnormalized probability values for the three states is 0.557175 + 0.6285875 + 0.2428625 = 1.428625. Therefore, the probability of a state advancement is 0.557175 / 1.428625 ≈ 0.39; the probability of a state maintenance is 0.6285875 / 1.428625 ≈ 0.44; and the probability of a state decline is 0.2428625 / 1.428625 ≈ 0.17. Thus, the state transition probability distribution is {advance: 0.39, maintain: 0.44, decline: 0.17}.
[0087] Step S144: Based on the state transition probability distribution, a state evolution simulation method based on random sampling is used to generate a set of possible state values for each healthy state node at the end of the first time slice unit. The set of possible state values contains multiple candidate state values.
[0088] Based on the state transition probability distribution obtained in step S143, the possible state values of healthy nodes at the end of the first time slice unit are simulated by random sampling.
[0089] Step S1441: Obtain the state transition probability distribution of each healthy state node at the end of the first time slice unit. The state transition probability distribution defines the probability value of the healthy state node transitioning to each of the multiple possible new states at the end of the first time slice unit.
[0090] For example, the state transition probability distribution of node TH1 is {increase: 0.39, maintain: 0.44, decrease: 0.17}, where "increase" means the state value increases by 0.1, "maintain" means the state value remains unchanged, and "decrease" means the state value decreases by 0.1.
[0091] Step S1442: Based on the state transition probability distribution of each healthy state node, define the cumulative probability interval in the probability space for each possible new state, wherein the length of the cumulative probability interval is proportional to its corresponding probability value.
[0092] For node TH1, the probability space is [0,1]. The cumulative probability interval for the "increase" state is [0,0.39], the cumulative probability interval for the "maintain" state is [0.39,0.39+0.44]=[0.39,0.83], and the cumulative probability interval for the "decrease" state is [0.83,1].
[0093] Step S1443: Based on the preset total number of simulations, initialize a simulation result record table to record the simulated state values of all healthy state nodes at the end of the first time slice unit.
[0094] The preset total number of simulations is M, for example, M=1000. The simulation result record table is an M-row, K-column matrix, where K is the number of healthy nodes, and each cell records the status value of the corresponding node for the corresponding number of simulations.
[0095] Step S1444: For each simulation, generate a random number that is uniformly distributed between zero and one for each healthy node independently.
[0096] For the i-th simulation (i from 1 to M), a random number r is generated for each healthy node, where r is uniformly distributed in the range [0,1]. For example, the random number generated for node TH1 is 0.56.
[0097] Step S1445: Map the random number generated by each healthy state node to the cumulative probability interval space of that healthy state node, and determine the specific cumulative probability interval into which the random number falls.
[0098] Comparing the random number 0.56 of node TH1 with the cumulative probability interval, 0.39≤0.56<0.83, therefore it falls into the cumulative probability interval of the "maintain" state.
[0099] Step S1446: Based on the cumulative probability interval of the random number falling into the range, determine the new state selected by the corresponding healthy state node at the end of the first time slice unit in this simulation, and use this new state as the candidate state value of the healthy state node in this simulation.
[0100] Since the random number of node TH1 falls into the "maintain" state range, its candidate state value is the current state value of 0.6588 (maintaining unchanged).
[0101] Step S1447: Record the candidate state value obtained by each healthy state node in this simulation to the corresponding position in the simulation result record table.
[0102] Record the candidate state value of node TH1, 0.6588, into the cell in the i-th row and corresponding column of the simulation result record table.
[0103] Step S1448: Repeat the operations of generating random numbers, mapping to the cumulative probability interval, determining the new state, and recording the simulation results until all simulation processes of the preset total number of simulations are completed.
[0104] Follow steps S1444 to S1447 to complete M simulations in sequence and fill in the simulation result record table.
[0105] Step S1449: Traverse the simulation result record table. For each healthy state node, collect all candidate state values generated in all simulations, remove duplicate candidate state values, and form a set of possible state values for the healthy state node at the end of the first time slice unit.
[0106] For example, in 1000 simulations, node TH1 may generate candidate state values of 0.5588 (decreasing), 0.6588 (maintaining), and 0.7588 (increasing). After removing duplicate values, the set of possible state values is {0.5588, 0.6588, 0.7588}.
[0107] Step S145: Based on the connection topology between the set of possible state values and the healthy state nodes, calculate the propagation range and cumulative propagation intensity of the risk signal in the individualized risk knowledge graph within the first time slice unit, and use it as a snapshot of the health state evolution of the first time slice unit.
[0108] Based on the set of possible state values for each healthy node and the connection topology of the individualized risk knowledge graph, the propagation process of risk signals in the graph is simulated. For each possible state value of each node, the influence strength on adjacent nodes through risk propagation edges is calculated, and the sum of the strengths within the propagation range is accumulated. For example, when the possible state value of node TH1 is 0.7588, it influences node TH2 through its outgoing edges, with an influence strength of 0.7588 * connection strength parameter. The propagation influence strengths of all nodes are accumulated to form a health state evolution snapshot of the first time slice unit. This health state evolution snapshot contains the distribution of possible state values of each node and the risk propagation situation at the end of the time slice.
[0109] Step S146: Use the health state evolution snapshot of the first time slice unit as the input of the simulation module of the second time slice unit of the feature evolution prediction model, and repeatedly perform the operations of calculating the state transition probability distribution, generating the set of possible state values and generating the health state evolution snapshot to obtain the health state evolution snapshot of the second time slice unit.
[0110] The set of possible state values of each node in the health state evolution snapshot of the first time slice unit is used as the initial state of the second time slice unit. The state transition probability distribution, the set of possible state values, and the health state evolution snapshot of the second time slice unit are calculated according to the method of steps S143 to S145.
[0111] Step S147: Iterate through the time slice unit simulation operation until all the divided time slice units are traversed to obtain a continuous time slice health status evolution snapshot sequence covering the entire future time window.
[0112] Following step S146, the subsequent time slice units are simulated sequentially until N time slice units are simulated, resulting in a sequence containing N health state evolution snapshots, with each snapshot corresponding to the health state at the end of a time slice unit.
[0113] Step S148: Perform feature aggregation and trend extraction operations on the time dimension of the continuous time slice health state evolution snapshot sequence, and calculate the slope information, curvature information and fluctuation amplitude information of the state value of each health state node as a function of time.
[0114] For each healthy state node, the mean of its possible state values in each time slice unit is extracted from the continuous time-slice health state evolution snapshot sequence to form the state value time series of that node. Then, trend analysis is performed on this time series to calculate slope information (rate of change of state values per unit time), curvature information (rate of change of slope), and fluctuation amplitude information (standard deviation of state values). For example, the state value time series of node TH1 is [0.6588, 0.70, 0.73, ...]. The slope is calculated through linear fitting, the curvature is calculated through the second derivative, and the fluctuation amplitude is obtained by calculating the standard deviation of the series.
[0115] Step S149: Integrate the slope, curvature, and fluctuation amplitude information of all health status nodes to construct a high-dimensional spatiotemporal evolution feature tensor that describes the dynamic evolution of the entire individualized risk knowledge graph.
[0116] The slope, curvature, and fluctuation amplitude information of each healthy node are used as three feature components and combined into a feature vector. The feature vectors of all nodes are arranged in node order, forming a matrix of dimension K×3, where K is the number of healthy nodes. This matrix is then combined with time dimension information to construct a high-dimensional spatiotemporal evolution feature tensor of dimension N×K×3, where N is the number of time slice units.
[0117] Step S1410: Input the high-dimensional spatiotemporal evolution feature tensor into the dynamic risk encoder of the feature evolution prediction model, extract the key discriminative features in the evolution mode through convolution and pooling operations, and compress and encode them into a dynamic risk feature vector. Each element in the dynamic risk feature vector corresponds to a risk score of a health state evolution mode.
[0118] The dynamic risk encoder employs a three-dimensional convolutional neural network structure, comprising multiple convolutional and pooling layers. First, a three-dimensional convolutional kernel is used to convolve the high-dimensional spatiotemporal evolution feature tensor to extract local spatiotemporal features. Then, a max-pooling layer reduces the dimensionality of the convolutional features, preserving key features. After processing by multiple convolutional and pooling layers, the resulting feature vector is input into a fully connected layer for nonlinear transformation and dimensionality compression, ultimately generating a dynamic risk feature vector with dimension D9. Each element in this dynamic risk feature vector corresponds to a risk score for a specific health state evolution pattern, such as a "rapidly increasing colorectal cancer risk pattern" or a "risk fluctuation pattern."
[0119] Step S150: Based on the dynamic risk feature vector and the preset screening strategy rule base, perform strategy mapping processing to generate a personalized colorectal cancer screening plan for the target individual. The personalized colorectal cancer screening plan includes the recommended screening technology type, the priority execution order of the screening technology type, the initial screening time point, and subsequent screening cycle parameters.
[0120] After obtaining the dynamic risk feature vector, it needs to be deeply matched with a pre-defined screening strategy rule base. This rule base is a structured decision-making system built on evidence-based medicine and multi-center clinical data, containing screening technology combinations, execution paths, and timeliness parameters corresponding to different risk levels. Each element in the dynamic risk feature vector represents a risk score for a specific health state evolution pattern. These scores are used as index keys to query predefined decision tree nodes in the rule base. By matching risk threshold ranges, health state node types, and evolution speed parameters layer by layer, a complete screening plan including technology selection, sequential arrangement, and time planning is finally generated. The entire mapping process adopts a weighted rule reasoning mechanism to ensure that the plan simultaneously meets the requirements of clinical effectiveness, technical feasibility, and individual tolerability.
[0121] Step S151: Analyze the risk score of each health status evolution pattern in the dynamic risk feature vector, and compare the risk score of each health status evolution pattern with the preset risk threshold range in the screening strategy rule base.
[0122] The dynamic risk feature vector contains D9 risk scores for health status evolution patterns. Each score corresponds to a specific pathophysiological evolution path, such as the "malignant transformation path of adenomatous polyps" and the "carcinoma transformation path of inflammatory bowel disease." The screening strategy rule base pre-sets three continuous risk threshold intervals for each evolution pattern, corresponding to low-risk, medium-risk, and high-risk levels, respectively. The parsing process first normalizes and verifies each score to ensure that its value is within the standard range of 0 to 1. Then, it uses a binary search method to determine the risk interval to which each score belongs. For example, for the "malignant transformation path of adenomatous polyps" score, the rule base sets the low-risk interval to 0 to A, the medium-risk interval to A to B, and the high-risk interval to B to 1. Each score is compared with the boundary value of the corresponding interval to generate a preliminary risk level matrix.
[0123] Step S152: Based on the results of the comparison operation, determine the risk level category to which each health status evolution pattern risk score belongs. The risk level categories include the third risk level category, the second risk level category, and the first risk level category.
[0124] Based on the interval comparison results in step S151, the risk score of each health status evolution pattern is mapped to its corresponding risk level category. The third risk level category corresponds to the low-risk interval, indicating that the evolution pattern currently poses a low threat to the individual's health; the second risk level category corresponds to the medium-risk interval, suggesting the need for clinical attention; and the first risk level category corresponds to the high-risk interval, indicating the need for priority intervention. Fuzzy logic is used to handle boundary value issues during the level determination process. When the score falls exactly at the interval boundary, a secondary judgment is made based on the clinical weight coefficient of the pattern. For example, if a score equals the medium-to-high risk threshold B, and the pattern is associated with high-risk factors such as a family history of colorectal cancer, it will be classified into the first risk level category; otherwise, the second risk level category determination will be maintained. The final generated risk level matrix will serve as the core basis for subsequent technology selection.
[0125] Step S153: Calculate the total number and density distribution of the health status evolution pattern risk scores belonging to the first risk level category. Based on the total number and density distribution, select the set of primary screening technology types that match them from the screening strategy rule base. The set of primary screening technology types includes fecal occult blood test technology, colonoscopy technology, and virtual colon imaging technology.
[0126] First, the scores for the first risk level category are statistically analyzed in two ways: the total number is counted by traversing the risk level matrix, recording the number N of all scores classified as first risk level; the density distribution is calculated using kernel density estimation to analyze the spatial distribution characteristics of these high-risk scores in the dynamic risk feature vector, identifying any clustered risk patterns. The screening strategy rule base stores a ternary mapping table of total number, density, and technology. For example, when the total number of high-risk scores exceeds C and exhibits a continuous distribution, the primary screening technology type set automatically triggers a combination scheme including colonoscopy. Based on the statistical results, the decision module in the rule base is dynamically invoked to select core technologies EF from D candidate primary screening technologies, forming an initial technology pool.
[0127] Step S154: Based on the specific health status node type corresponding to the health status evolution mode risk score belonging to the second risk level category, query the auxiliary screening technology type that has a preset association with the health status node type from the screening strategy rule base, and generate a set of auxiliary screening technology types.
[0128] The scores for the second risk level category are associated with specific health status nodes, such as "gut microbiota dysregulation points" and "chronic inflammation points." Semantic parsing techniques are used to extract the node identifiers corresponding to each medium-risk score, and then the "node-technology" association dictionary in the rule base is queried. This association dictionary, built based on clinical guidelines and expert consensus, records the recommended auxiliary screening technologies for each health status node; for example, "gut microbiota dysregulation points" correspond to fecal metagenomic sequencing, and "chronic inflammation points" correspond to fecal calprotectin detection. The selection of auxiliary screening technologies also needs to consider their complementarity with primary screening technologies. By calculating the information gain rate between technologies, it is ensured that auxiliary technologies can provide unique clinical information not covered by primary technologies, ultimately forming a set of auxiliary screening technology types that include GH (growth, inflammation, and toxicity) technologies.
[0129] Step S155: Perform a technology compatibility and cost-effectiveness analysis on the set of primary screening technology types and the set of auxiliary screening technology types. Based on the analysis results, determine the priority execution order of each screening technology type to form an ordered sequence of screening technology types.
[0130] The technology compatibility analysis unfolds across three dimensions: physiological compatibility assesses the individual's physical requirements for different technologies, such as avoiding invasive examinations for those with coagulation disorders; temporal compatibility ensures reasonable intervals between technologies, such as a one-week interval between colonoscopy and virtual colonography; and outcome compatibility verifies whether there is interference in the interpretation of results between technologies, such as the need to perform fecal occult blood tests after dietary control. Cost-effectiveness analysis utilizes a decision tree model to calculate the incremental cost-effectiveness ratio of each technology, comprehensively considering detection sensitivity, false positive rate, and cost per examination. The ranking process employs a multi-attribute decision algorithm, weighting and fusing compatibility scores and utility values to generate a comprehensive score. These scores are then arranged in descending order to form an ordered sequence of screening technology types, ensuring that high-priority technologies receive priority access to resources and provide crucial diagnostic information.
[0131] Step S156: Based on the health status evolution pattern risk score representing the evolution speed in the dynamic risk feature vector, calculate the theoretical time window length for the target individual to develop from the current health status to the critical state requiring clinical intervention.
[0132] A specific dimension is reserved in the dynamic risk feature vector to characterize the rate of health status evolution. This score is generated by integrating multiple time-series features to reflect the speed of disease progression. The theoretical time window length is calculated using a survival analysis-based predictive model, with the evolution rate score as the core independent variable, while also incorporating covariates such as age, gender, and underlying diseases. The model converts the evolution rate score into a probability distribution of disease progression time using a risk proportionality function, and takes the median of the distribution as a point estimate of the theoretical time window length. Sensitivity analysis is also performed during the calculation process to assess the impact of different covariate values on the time window length.
[0133] Step S157: Based on the theoretical time window length and the preset safety buffer coefficient, calculate the recommended initial screening time point, which is earlier than the start point of the theoretical time window.
[0134] The theoretical time window represents the estimated time it takes for the disease to progress naturally to a state requiring intervention. To ensure timely clinical intervention, a safety buffer mechanism needs to be introduced. The safety buffer coefficient is set based on the disease type and clinical experience. Colorectal cancer screening typically uses the JK buffer coefficient, and the safety buffer time is obtained by multiplying the theoretical time window length by the buffer coefficient. The initial screening time point is calculated using a backward extrapolation method, shifting the safety buffer time forward from the starting point of the theoretical time window. The buffer coefficient is also adjusted based on the individual's actual health condition, with a slightly increased buffer coefficient for high-risk groups to allow sufficient intervention time.
[0135] Step S158: Based on the fluctuation range and trend stability information of the health status evolution pattern risk score, assess the predictability of the target individual's health status evolution, and search for the corresponding screening cycle adjustment parameters from the screening strategy rule base based on the predictability.
[0136] The predictability assessment of health status evolution is accomplished through two core indicators: volatility information reflects the dispersion of risk scores over time, obtained by calculating the coefficient of variation of the score series; trend stability information characterizes the consistency of score change trends, measured by the coefficient of determination of linear regression. These two indicators are input into a fuzzy comprehensive evaluation model to generate a predictability index between 0 and 1, with the index closer to 1 indicating a more stable and predictable evolution process. The screening strategy rule base stores the mapping relationship between the predictability index and the cycle adjustment parameter; high predictability corresponds to a larger adjustment parameter (extended cycle), and low predictability corresponds to a smaller adjustment parameter (shortened cycle).
[0137] Step S1581: Extract the fluctuation range information and trend stability information corresponding to the risk score of each health status evolution mode from the dynamic risk feature vector.
[0138] The dynamic risk feature vector employs a composite encoding structure, with two auxiliary parameter bits appended to each risk score item. These bits store the volatility amplitude and trend stability information for that score, respectively. The volatility amplitude information is derived by calculating the maximum and minimum score difference of the health status evolution pattern in historical simulations, reflecting the possible range of fluctuations. The trend stability information is obtained through the smoothing exponent in time series analysis, characterizing the goodness of fit of the trend line. The extraction process requires parsing and decoding the vector, separating each score from its corresponding auxiliary parameters to form a triplet data structure containing the score value, volatility amplitude, and trend stability.
[0139] Step S1582: Calculate the average value of the fluctuation amplitude information of all health state evolution mode risk scores to obtain the overall fluctuation amplitude index; calculate the average value of the trend stability information of all health state evolution mode risk scores to obtain the overall trend stability index.
[0140] The overall fluctuation amplitude index reflects the overall fluctuation level of an individual's health state by performing arithmetic averaging on the fluctuation amplitude information of all health state evolution modes. Before calculation, each fluctuation amplitude value needs to be standardized to eliminate the dimension difference, and then the weighted average method is used to assign higher weights to evolution modes with high clinical importance. The calculation of the overall trend stability index is similar. The weighted average of all trend stability information is obtained to get a comprehensive index reflecting the consistency of the overall evolution trend.
[0141] Step S1583: Compare the overall fluctuation amplitude index with a preset fluctuation amplitude threshold, and compare the overall trend stability index with a preset trend stability threshold.
[0142] The system presets three fluctuation amplitude thresholds (low: ≤L, medium: L - M, high: >M) and three trend stability thresholds (high: ≥N, medium: O - P, low: <O), forming a 3×3 evaluation matrix. In the comparison process, the overall fluctuation amplitude index and the overall trend stability index are respectively compared with the corresponding thresholds to determine their respective belonging grade intervals. The threshold comparison adopts fuzzy boundary processing. When the index is close to the threshold, secondary judgment is made in combination with the individual's clinical characteristics to avoid mechanical classification.
[0143] Step S1584: According to the comparison results, determine the fluctuation amplitude grade to which the overall fluctuation amplitude index belongs, and determine the trend stability grade to which the overall trend stability index belongs.
[0144] Based on the comparison results of Step S1583, map the overall fluctuation amplitude index and the overall trend stability index to specific grades respectively. The fluctuation amplitude grades are divided into three levels: low, medium, and high, and the trend stability grades are divided into three levels: high, medium, and low. In the process of grade determination, clinical decision thresholds are introduced. For example, for high-risk populations of colorectal cancer, the determination threshold of the fluctuation amplitude grade will be appropriately reduced to improve sensitivity. The two determined grades will be used as input parameters for subsequent predictability evaluation.
[0145] Step S1585: Based on a preset mapping table of the correspondence between the fluctuation amplitude grade and the trend stability grade, map the obtained fluctuation amplitude grade and trend stability grade to obtain the health state evolution predictability grade of the target individual.
[0146] The system incorporates a three-dimensional mapping table of volatility amplitude, trend stability, and predictability, mapping nine possible combinations of levels to three predictability levels: high, medium, and low. The mapping rules are trained based on large clinical datasets; for example, a "low volatility - high stability" combination corresponds to a high predictability level, and a "high volatility - low stability" combination corresponds to a low predictability level. The mapping process employs a weighted voting algorithm, comprehensively considering the contribution of both levels, and adjusting boundary combinations by introducing a Bayesian correction factor.
[0147] Step S1586: Based on the predictability level of health status evolution, search the screening strategy rule base for the original screening cycle adjustment parameter that is directly associated with the predictability level of health status evolution.
[0148] The periodic adjustment parameter table in the screening strategy rule base is divided into three intervals based on predictability level, with each interval corresponding to a basic adjustment parameter range. High predictability level corresponds to the QR adjustment parameter range, medium predictability level corresponds to ST, and low predictability level corresponds to UV. Based on the mapped predictability level, specific parameter values are searched within the corresponding interval. The search process also considers individual age factors, appropriately lowering adjustment parameter values for the elderly population.
[0149] Step S1587: Obtain the basic screening cycle parameters of the target individual, and apply the original screening cycle adjustment parameters to the basic screening cycle parameters.
[0150] The baseline screening cycle parameters are set according to cancer screening guidelines, with the baseline cycle for colorectal cancer screening typically being WZ years. Parameters such as age and gender are retrieved from the individual's basic information module and combined with clinical guidelines to determine the initial baseline cycle. The original screening cycle adjustment parameters are multiplied onto the baseline cycle to obtain the initially adjusted screening cycle.
[0151] Step S1588: Calculate the final effective screening cycle adjustment parameters according to the preset adjustment rules. The screening cycle adjustment parameters are used to modify the time interval between two subsequent screening operations.
[0152] The final adjustment of the screening cycle must follow clinically feasible rules, including the cycle rounding rule, the minimum cycle limit, and the maximum cycle limit. First, the initial adjustment cycle is rounded to the nearest quarter, and then it is checked whether it meets the minimum cycle (not less than AA months) and the maximum cycle (not more than BB years) limits; finally, fine-tuning is made based on the individual's screening compliance history data, and the cycle can be appropriately extended for individuals with good compliance.
[0153] Step S159: Combine the screening cycle adjustment parameters with the basic screening cycle to generate subsequent screening cycle parameters for the target individual. The subsequent screening cycle parameters define the time interval between two consecutive screening operations.
[0154] The generation of subsequent screening cycle parameters employs a dynamic weighting mechanism, fusing the basic screening cycle and the screening cycle adjustment parameters through a weighted formula. Weight allocation is determined based on the predictability level: at a high predictability level, the adjustment parameter weight is CC, and the basic cycle weight is DD; at a medium predictability level, both weights are EE; and at a low predictability level, the adjustment parameter weight is FF, and the basic cycle weight is GG. The calculated results are standardized to integers in months.
[0155] Step S1510: Integrate the ordered screening technology type sequence, the initial screening time point, and subsequent screening cycle parameters, and assemble them into a structured personalized colorectal cancer screening plan that can be parsed and executed by the medical information system.
[0156] The integration process utilizes the HL7FHIR standard nursing plan resource format, mapping various parameters to standard fields. Ordered screening technology sequence is converted into a list of planned activities, each activity containing attributes such as technology code, execution order, and expected duration; the initial screening time point is converted into the planned start date and time; subsequent screening cycle parameters are converted into repetitive execution rules. Clinical decision support information is also added, such as pre-screening preparation requirements, contraindication warnings, and result interpretation guidelines. The final personalized colorectal cancer screening plan is a structured document in XML format, which can be directly imported into the hospital information system, supporting clinical workflows such as appointment scheduling, result tracking, and follow-up management.
[0157] Step S160: Based on the implementation feedback data and follow-up results data of the personalized colorectal cancer screening program in clinical practice, update the model parameters of the feature evolution prediction model and the risk association patterns in the population risk knowledge base.
[0158] Execution feedback and follow-up data constitute the key inputs for closed-loop optimization. Data such as technical parameters, individual compliance, and examination results during the screening process, as well as long-term follow-up health outcome information, are collected in real time through a medical data interface. This data, after standardization, is used in a dual-track update mechanism: firstly, incremental learning algorithms adjust the weight parameters of the feature evolution prediction model to improve prediction accuracy; secondly, a knowledge graph update engine discovers new risk association patterns and optimizes the strength parameters of existing patterns. The update process employs multi-source data fusion technology to ensure that the model and knowledge base continuously adapt to changes in clinical practice.
[0159] Step S161: Collect screening compliance record data, original screening operation result data, and post-screening pathological diagnosis confirmation data generated by the target individual during the implementation of the personalized colorectal cancer screening program.
[0160] Data collection is achieved through the hospital's information system integration platform, covering the entire screening process: Screening compliance record data includes indicators such as appointment confirmation rate, attendance rate, and examination completion rate, extracted through the outpatient appointment system and electronic medical record system; raw screening results data covers the raw outputs of various examinations, such as video images from colonoscopy, digital images of pathological slides, and raw readings from laboratory tests, collected through the examination equipment interface and LIS system; post-screening pathological diagnosis confirmation data serves as the gold standard, including the diagnostic conclusion, tumor stage, and histological type information from the pathology report, obtained through the pathology information system. All data is anonymized, and after removing personally identifiable information, it is associated with a unique case identifier.
[0161] Step S162: Standardize and structure the original screening results data, and extract the characteristic description information indicating positive findings, negative findings, and indeterminate findings.
[0162] Standardization processing adopted the SDTM standard of the Clinical Data Standardization Society (CDISC) to convert raw result data from different sources into a unified format. Structured processing was achieved through natural language processing technology. For text-based results (such as radiology reports), entity recognition and relation extraction were performed to extract key findings; for image-based results (such as colonoscopy images), computer-aided detection was used to identify suspicious lesion areas; and for numerical results (such as tumor marker concentrations), threshold judgment was performed to determine the nature of the result. The processed data was divided into three categories: positive finding feature descriptions, negative finding feature descriptions, and undetermined feature descriptions, which were stored in the corresponding fields of the structured database.
[0163] Step S163: Use the pathological diagnosis confirmation data as the gold standard label, compare and verify it with the feature description information in the original screening operation result data, and generate a true performance evaluation report of the screening technology type. The true performance evaluation report includes the true positive rate index, the false positive rate index, and the detection rate index.
[0164] Cross-validation was used for comparison and verification, performing a four-fold table analysis between the raw results of each screening technology and the pathological diagnosis results. The true positive rate was calculated as the number of correctly detected positive cases divided by the total number of pathologically confirmed positive cases; the false positive rate was calculated as the number of incorrectly detected positive cases divided by the total number of pathologically confirmed negative cases; and the detection rate was calculated as the number of positive cases detected by the technology divided by the total number of people screened using that technology. The evaluation report also included auxiliary indicators such as positive predictive value, negative predictive value, and Kappa concordance coefficient to comprehensively evaluate the clinical performance of the screening technologies. The report was generated using a standardized template and included numerical results, confidence intervals, and comparative analysis with historical data.
[0165] Step S164: Extract the true positive rate, false positive rate and detection rate from the real performance evaluation report to construct a feature vector of screening technology performance.
[0166] The performance feature vector adopts a fixed-dimensional structure, containing three core dimensions: true positive rate, false positive rate, and detection rate, as well as two extended dimensions: positive predictive value and negative predictive value. The values of each dimension are normalized to the 0-1 range, with the true positive rate and detection rate using positive normalization (higher values are better), and the false positive rate using negative normalization (lower values are better). During vector construction, technical identifier encoding and evaluation timestamps are also added, forming a structured feature vector containing six elements.
[0167] Step S165: Collect health status review data of the target individuals at preset follow-up time points. The health status review data includes new medical imaging slice units and clinical test indicator units.
[0168] Follow-up time points are set based on the natural history of the disease and the screening program. Colorectal cancer screening typically sets three key follow-up points: HH year, II year, and JJ year. Data collection employs a multimodal acquisition method: medical imaging slices, including follow-up colonoscopy images and abdominal CT images, are acquired through a PACS system; clinical laboratory indicators, including tumor marker detection, complete blood count, and biochemical indicators, are acquired through a LIS system. The collection process utilizes an automated reminder mechanism, generating a to-be-collected list KK months before the follow-up time point, which is then completed by medical staff. All follow-up data and baseline data are formatted using the same standards to ensure comparability.
[0169] Step S166: Input the health status review data into the feature evolution prediction model to obtain a new dynamic risk feature vector. Compare and analyze the new dynamic risk feature vector with the historical dynamic risk feature vector to calculate the actual change trajectory of the health status evolution pattern risk score.
[0170] Health status review data undergoes the same preprocessing and feature extraction process as baseline data, and is input into the feature evolution prediction model to generate new dynamic risk feature vectors. Comparative analysis employs a time-series comparison method, aligning the new vectors with historical vectors along corresponding dimensions to calculate the absolute change and relative rate of change of risk scores for each health status evolution pattern. The actual change trajectory is visualized through a line graph, and morphological parameters such as the slope, curvature, and fluctuation frequency of the trajectory are calculated. The analysis focuses on the changing trends of high-risk patterns, identifying key turning points where risk increases or decreases.
[0171] Step S167: Select target individual cases whose actual change trajectory and the evolution path previously predicted by the feature evolution prediction model exceed the preset deviation range, and mark them as key update samples.
[0172] The deviation range is determined by the distribution of prediction errors, using LL times the root mean square error of the model on the training set as the threshold. For each target individual, the mean absolute error between the actual trajectory and the predicted evolution path is calculated; if the error exceeds the threshold, it is marked as a key update sample. The screening process also considers the clinical significance of the deviation, prioritizing cases involving key health status nodes (such as cancer risk nodes). The key update samples form an independent update dataset, containing the individual's baseline data, predicted data, and actual outcome data, for subsequent model parameter tuning and knowledge base updates.
[0173] Step S168: Extract key feature combinations, key feature evolution patterns, and screening technology performance correlation features that lead to prediction bias from the full-dimensional health data set, historical overall health representation vector, actual change trajectory, and screening technology performance feature vector of the key updated samples.
[0174] Feature extraction employs an ensemble feature selection approach, combining recursive feature elimination and L1 regularization to identify the feature combinations that contribute most to prediction bias from high-dimensional data. Key feature combinations refer to the set of multiple features that synergistically lead to prediction bias; key feature evolution patterns refer to anomalous feature trajectories over time; and screening for technology performance-related features refers to individual features associated with false positive / false negative results. The extraction process generates a feature importance ranking table, retaining the top MM% of features for model updates.
[0175] Step S169: Input the key feature combination, key feature evolution pattern and screening technology performance related features into the parameter adaptive adjustment module of the feature evolution prediction model, calculate the gradient information of the model prediction error with respect to the model parameters through the backpropagation algorithm, and correct the gradient information by combining the screening technology performance related features.
[0176] The parameter adaptive adjustment module employs an online learning framework, using key features as new training samples input into the model. The backpropagation algorithm calculates the gradient of each layer's weight parameters with respect to the prediction error, and momentum optimization accelerates convergence. Features related to the screening technology's performance are used as correction factors to weight and adjust the gradient information; for example, higher weights are given to the feature weight gradients corresponding to technologies with high true positive rates. The adjustment process uses mini-batch gradient descent, employing N key update samples in each update. The learning rate is dynamically adjusted based on error changes, and iteration stops when the error no longer decreases.
[0177] Step S1610: Based on the corrected gradient information and the preset learning rate parameters, fine-tune and update the weight parameters of the network layer responsible for calculating the risk score of the feature evolution pattern in the feature evolution prediction model.
[0178] Fine-tuning updates target the model's output layer and the last hidden layer, which directly affect the risk score calculation. The learning rate is initially set to OO and decays with iterations using a cosine annealing scheduling strategy. Weight updates employ gradient truncation to limit the gradient norm to within PP to prevent gradient explosion. After each update, the model is evaluated, calculating the prediction accuracy on the validation set. If the accuracy does not improve after QQ consecutive iterations, the update process terminates. The updated model parameters are stored as the new version and run in parallel with the old version. A / B testing is used to verify the improvement.
[0179] Step S1611: Add the key feature combination, key feature evolution pattern, screening technology performance feature vector, and corresponding actual health status outcome as new risk association pattern entries to the population risk knowledge base, and adjust the connection strength parameters of the risk transmission edge between relevant health status nodes in combination with the true positive rate index, false positive rate index, and detection rate index.
[0180] The construction of new risk association pattern entries follows the triple structure of a knowledge graph (head node-relationship-tail node). Key feature combinations form the head node, key feature evolution patterns form the relationship, and actual health status outcomes form the tail node. Screening technology performance feature vectors are added as attributes to the relationship, affecting the calculation of the connection strength parameter. The connection strength parameter is adjusted using a Bayesian update rule, taking the true positive rate, false positive rate, and detection rate as prior information, calculating the posterior probability distribution through a likelihood function, and taking the expected value of the distribution as the new connection strength parameter. The adjustment process must maintain the consistency of the knowledge graph, and related risk transmission edges must be adjusted collaboratively.
[0181] For example, step S16111: the key feature combination is parsed into multiple specific health status node identifiers constituting the key feature combination and their respective health element intensity value sets; the key feature evolution pattern is parsed into the sequence of health status node identifiers involved, and the observed state transition direction and intensity change between adjacent health status nodes in the health status node identifier sequence; the actual health status outcome is parsed into the final health status node identifier and its final state value directly related to the actual health status outcome; the screening technology performance feature vector is parsed into specific true positive rate index values, false positive rate index values, and detection rate index values.
[0182] The parsing process employs a semantic parsing engine to convert unstructured feature descriptions into structured data. Key feature combinations are parsed into a set of key-value pairs (node identifier: intensity value); key feature evolution patterns are parsed into node sequences and transition matrices; actual health status outcomes are parsed into tuples (final node identifier: final state value); and screening technology performance feature vectors are parsed into three numerical indicators, corresponding to the true positive rate, false positive rate, and detection rate, respectively. The parsing results are stored in JSON format to ensure compatibility with the knowledge base.
[0183] Step S16112: Based on the parsed sequence of specific health status node identifiers, the direction and intensity of state transition changes, the specific index values contained in the screening technology performance feature vector, and the final health status node identifier and final state value, a new risk association pattern entry is constructed in a structured manner. This risk association pattern entry includes the source health status node identifier, the target health status node identifier, the state transition rule description, the associated screening technology performance data, and the expected outcome state information.
[0184] The new entries are constructed using a template-filling method. The source health status node identifier is taken from the starting node of the node sequence, and the target health status node identifier is the final node. The state transition rule description uses natural language generation technology to convert the state transition direction and intensity change into readable text. The associated screening technology performance data includes the parsed values of the three indicators. The expected outcome status information is a text description of the final status value.
[0185] Step S16113: Calculate the initial estimate of the connection strength parameter corresponding to the state transition rule from the source healthy state node to the target healthy state node in the new risk association pattern entry. The initial estimate is calculated based on the observed intensity change and the time window length.
[0186] The initial estimate uses the path integral method, summing the intensity changes of each segment in the state transition rule and dividing by the corresponding time window length (in years). During the calculation, the changes in different time units need to be standardized and uniformly converted into annual change rates to ensure the comparability of parameters.
[0187] Step S16114: Input the true positive rate index value, false positive rate index value and detection rate index value into the connection strength correction function to calculate a connection strength correction factor.
[0188] The connection strength correction function uses a weighted summation form, with the true positive rate (RR) as the weight, the detection rate (SS) as the weight, and the false positive rate (TT) as the weight. The function expression is: Correction factor = RR × True positive rate + SS × Detection rate + TT × False positive rate. All indicator values have been normalized to the range of 0-1, therefore the correction factor value range is within the preset interval.
[0189] Step S16115: Combine the initial estimate of the connection strength parameter with the connection strength correction factor to obtain the final estimate of the connection strength parameter that incorporates the screening technology performance information.
[0190] The synthesis operation uses a multiplicative model: final estimated connection strength parameter = initial estimated value × correction factor. This process integrates technical performance information into the connection strength. When the technical performance is good (high true positive rate, high detection rate, low false positive rate), the correction factor is close to 1, and the final parameter is close to the initial estimated value; when the technical performance is poor, the correction factor decreases, and the final parameter decreases accordingly.
[0191] Step S16116: Search the group risk knowledge base for existing risk association pattern entries that are identical to the source health status node identifier and the target health status node identifier in the new risk association pattern entry.
[0192] The search process is implemented through the knowledge graph's indexing system, performing precise matching based on the unique identifiers of the source and target nodes. First, it searches the relational table of the knowledge base for a tuple of (source node identifier, target node identifier). If a matching entry exists, the existing risk association pattern entry is returned, including its current connection strength parameter and confidence value; otherwise, an empty result is returned. The search results are used to determine whether to update an existing entry or add a new one.
[0193] Step S16117: If an existing risk association pattern entry is found, the final connection strength parameter estimate of the new risk association pattern entry is weighted and fused with the connection strength parameter of the existing risk association pattern entry, and the connection strength parameter of the corresponding risk transmission edge in the group risk knowledge base is updated with the fused value.
[0194] The weighted fusion adopts a time decay model, where the new parameter weight = 1 / (1+e^(-λ×t)), and λ is the decay coefficient, while t is the existence time of the existing item (in months). The fusion formula is: Updated parameter = New parameter × New weight + Existing parameter × (1 - New weight). Simultaneously, the confidence values of the items are updated and adjusted accordingly based on the parameter changes.
[0195] Step S16118: If no existing risk association pattern entry is found, the newly constructed risk association pattern entry and its calculated final connection strength parameter estimate are added as new entries to the group risk knowledge base. A new risk transmission edge is created in the topology of the group risk knowledge base, pointing from the source health status node identifier to the target health status node identifier, and the final connection strength parameter estimate is assigned to the new risk transmission edge.
[0196] New entries must undergo integrity verification to ensure they contain all necessary fields. When creating a new edge in the knowledge graph topology, the in-degree and out-degree information of the nodes must be updated, and a reverse index must be built to accelerate future search operations. The confidence level of new entries is initially set to UU, and gradually increases as more cases support this mode. An entry creation timestamp and version number are also generated for subsequent update management.
[0197] Step S16119: After adding a new entry or updating the connection strength parameter, check whether there is a logically conflicting circular risk transmission path in the group risk knowledge base. If so, adjust the connection strength parameter of the relevant risk transmission edge in a consistent manner according to the latest connection strength parameter.
[0198] Cycle path detection employs a depth-first search algorithm, traversing all reachable paths from the updated node. If a path returns to the starting node, it is considered a cycle. Consistency adjustment uses a minimum cut set algorithm, identifying the edge with the weakest connection strength within the cycle as the entry point and reducing its strength parameter by VV%. Simultaneously, the parameters of other edges in the cycle are adjusted proportionally to ensure the overall risk transmission strength remains conserved. After adjustment, the consistency of the knowledge graph must be re-verified until all cycle conflicts are resolved.
[0199] In an exemplary embodiment, a personalized intelligent screening system for colorectal cancer based on a multi-source fusion big data model is provided. This system can be a terminal, server, etc., and its internal structure is shown in Figure 2. The system includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, near-field communication, or other technologies. When the computer program is executed by the processor, it implements a personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, or it can be a button, trackball, or touchpad set on the shell of a colorectal cancer personalized intelligent screening system based on a multi-source fusion big data model, or it can be an external keyboard, touchpad, or mouse, etc.
[0200] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model, characterized in that, The method includes: acquiring a comprehensive health data set of the target individual, which includes gene sequencing fragment units, medical image slice units, clinical test indicator units, and electronic health record text units from different medical information systems; generating an overall health representation vector for the target individual based on the comprehensive health data set; performing matching and association analysis between the overall health representation vector and risk association patterns in a pre-constructed group risk knowledge base to construct an individualized risk knowledge graph for the target individual, which includes multiple health status nodes and risk transmission edges between health status nodes; and inputting the individualized risk knowledge graph into a pre-trained feature evolution prediction model to simulate health status nodes. During the state transition process under continuous time slices, a dynamic risk feature vector for the target individual is generated. Based on the dynamic risk feature vector and a pre-set screening strategy rule base, a personalized colorectal cancer screening plan is generated for the target individual. This personalized colorectal cancer screening plan includes recommended screening technology types, priority execution order of screening technologies, initial screening time point, and subsequent screening cycle parameters. Based on the execution feedback data and follow-up results data of the personalized colorectal cancer screening plan in clinical practice, the model parameters of the feature evolution prediction model and the risk association patterns in the population risk knowledge base are updated. The individualized risk knowledge graph is then input into the pre-trained feature evolution prediction model to simulate the health state nodes in the process. The state transition process under continuous time slices generates a dynamic risk feature vector for the target individual, including: extracting the initial state feature representations of all healthy state nodes and the connection strength parameters of the risk transmission edges between healthy state nodes from the individualized risk knowledge graph; inputting the initial state feature representations and connection strength parameters into the time slice partitioning module of the feature evolution prediction model, dividing the future time window into multiple continuous time slice units according to the preset future time window length and time slice granularity; in the simulation module of the first time slice unit of the feature evolution prediction model, normalizing the initial state feature representations of healthy state nodes and normalizing the connection strength parameters of the risk transmission edges; and based on the normalization... The initial state features of the healthy state nodes after normalization and the connection strength parameters of the risk propagation edges after normalization are used to calculate the state transition probability distribution of each healthy state node at the end of the first time slice unit. Based on the state transition probability distribution, a state evolution simulation method based on random sampling is used to generate a set of possible state values for each healthy state node at the end of the first time slice unit. The set of possible state values contains multiple candidate state values. Based on the connection topology between the set of possible state values and the healthy state nodes, the propagation range and cumulative propagation intensity of the risk signal in the individualized risk knowledge graph within the first time slice unit are calculated as a snapshot of the health state evolution in the first time slice unit.The health status evolution snapshot of the first time slice unit is used as the input to the simulation module of the second time slice unit of the feature evolution prediction model. The operations of calculating the state transition probability distribution, generating the set of possible state values, and generating the health status evolution snapshot are repeatedly executed to obtain the health status evolution snapshot of the second time slice unit. The simulation operation of the time slice units is iteratively executed until all divided time slice units are traversed, resulting in a continuous time slice health status evolution snapshot sequence covering the entire future time window. The continuous time slice health status evolution snapshot sequence is subjected to time-dimensional feature aggregation and trend extraction operations to calculate the slope, curvature, and fluctuation amplitude information of the state value of each health status node over time. The slope, curvature, and fluctuation amplitude information of all health status nodes are fused to construct a high-dimensional spatiotemporal evolution feature tensor describing the dynamic evolution of the entire individualized risk knowledge graph. This high-dimensional spatiotemporal evolution feature tensor is input into the dynamic risk encoder of the feature evolution prediction model. Key discriminative features in the evolution patterns are extracted through convolution and pooling operations and compressed and encoded into a dynamic risk feature vector. Each element in the dynamic risk feature vector corresponds to a risk score for a health status evolution pattern.
2. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, The process of generating an overall health representation vector for a target individual based on a comprehensive health data set includes: performing deep semantic encoding on gene sequencing fragment units to extract single nucleotide polymorphism (SNP) site variation information and methylation level information of gene expression regulatory regions, generating gene-dimensional feature vectors; performing multi-scale structural analysis on medical image slice units to identify the texture regularity information of colonic mucosa, the uniformity information of colonic wall thickness, and the morphological contour information of intestinal polyps in medical image slice units, generating image-dimensional feature vectors; performing time-series fluctuation pattern extraction on clinical laboratory indicator units to analyze the changing trends and periodic fluctuation information of carcinoembryonic antigen (CEA) concentration and carbohydrate antigen (CA) concentration in the historical time series of clinical laboratory indicator units, generating clinical indicator-dimensional feature vectors; performing natural language understanding processing on electronic health record text units to extract descriptions of past medical history of intestinal diseases, family genetic disease history, and personal lifestyle habits recorded in the electronic health record text units, generating text-dimensional feature vectors; and inputting the gene-dimensional feature vector, image-dimensional feature vector, clinical indicator-dimensional feature vector, and text-dimensional feature vector into a cross-modal feature interaction channel. The cross-modal feature interaction channel comprises multiple parallel feature interaction sub-channels. Within each sub-channel, feature vectors representing gene, image, clinical indicators, and text dimensions are standardized to obtain standardized feature vectors for each dimension. Based on these standardized vectors, a bidirectional feature attention weighting operation is performed to calculate the semantic relevance score between any two feature vectors of different dimensions. The contribution weights of different feature vectors in the feature interaction process are dynamically adjusted according to the semantic relevance scores. A weighted aggregation operation is then performed on the standardized feature vectors based on the adjusted contribution weights to generate a preliminary fused feature representation. Redundancy compression is applied to the preliminary fused feature representation to obtain a unified feature representation. This unified feature representation is then input into a multilayer perceptron network for nonlinear transformation and dimensionality increase operations to generate feature projections in a high-dimensional potential health feature space. Dimensionality reduction and normalization are then performed on the feature projections in the high-dimensional potential health feature space to obtain a fixed-dimensional overall health representation vector. Each dimension of this vector corresponds to a comprehensive health status characterization factor that integrates multi-source data information.
3. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, The process of mapping dynamic risk feature vectors to a pre-defined screening strategy rule base to generate personalized colorectal cancer screening plans for target individuals includes: parsing the risk score of each health status evolution pattern in the dynamic risk feature vector; comparing each health status evolution pattern risk score with a pre-defined risk threshold range in the screening strategy rule base; determining the risk level category to which each health status evolution pattern risk score belongs based on the comparison results, including the third risk level category, the second risk level category, and the first risk level category; statistically analyzing the total number and density distribution of health status evolution pattern risk scores belonging to the first risk level category; selecting a set of primary screening technology types matching the first risk level category from the screening strategy rule base based on the total number and density distribution, including fecal occult blood test, colonoscopy, and virtual colon imaging; querying the screening strategy rule base for auxiliary screening technology types with a pre-defined association relationship with the specific health status node type corresponding to the health status evolution pattern risk score belonging to the second risk level category, generating a set of auxiliary screening technology types; and performing primary screening... A technical compatibility and cost-effectiveness analysis is performed on the set of screening technology types and the set of auxiliary screening technology types. Based on the analysis results, the priority execution order of each screening technology type is determined, forming an ordered screening technology type sequence. According to the health status evolution pattern risk score representing the evolution speed in the dynamic risk feature vector, the theoretical time window length for the target individual to progress from the current health status to the critical state requiring clinical intervention is calculated. Based on the theoretical time window length and a preset safety buffer coefficient, a recommended initial screening time point is calculated, which is earlier than the start point of the theoretical time window. Based on the fluctuation amplitude and trend stability information of the health status evolution pattern risk score, the predictability of the target individual's health status evolution is assessed, and the corresponding screening cycle adjustment parameters are searched from the screening strategy rule base based on the predictability. Combining the screening cycle adjustment parameters and the basic screening cycle, subsequent screening cycle parameters for the target individual are generated, defining the time interval between two consecutive screening operations. The ordered screening technology type sequence, the initial screening time point, and the subsequent screening cycle parameters are integrated and structured into a personalized colorectal cancer screening plan that can be parsed and executed by a medical information system.
4. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, The process of updating the model parameters of the feature evolution prediction model and the risk association patterns in the population risk knowledge base based on the implementation feedback data and follow-up results of the personalized colorectal cancer screening program in clinical practice includes: collecting screening compliance record data, original screening operation result data, and post-screening pathological diagnosis confirmation data generated by target individuals during the implementation of the personalized colorectal cancer screening program; standardizing and structuring the original screening operation result data, extracting feature description information indicating positive findings, feature description information indicating negative findings, and feature description information that cannot be determined; and using the pathological diagnosis confirmation data as the gold standard label, comparing and verifying it with the feature description information in the original screening operation result data to generate a screening... The system retrieves the actual performance evaluation reports of the technology types, which include true positive rate, false positive rate, and detection rate indicators. It extracts these indicators from the reports to construct a performance feature vector for the screening technology. At predetermined follow-up time points, it collects health status re-examination data from target individuals, including new medical imaging slices and clinical laboratory indicator units. This data is then input into a feature evolution prediction model to obtain a new dynamic risk feature vector. The new dynamic risk feature vector is compared with historical dynamic risk feature vectors to calculate the actual change trajectory of the health status evolution pattern risk score. Finally, it filters out data that correlate with the actual change trajectory and the characteristic features. Individual cases whose predicted evolution paths exceed a preset deviation range are marked as key update samples. From the full-dimensional health data set, historical overall health representation vector, actual change trajectory, and screening technology performance feature vector of these key update samples, key feature combinations, key feature evolution patterns, and screening technology performance correlation features leading to prediction deviations are extracted. These key feature combinations, key feature evolution patterns, and screening technology performance correlation features are input into the parameter adaptive adjustment module of the feature evolution prediction model. The gradient information of the model prediction error with respect to the model parameters is calculated using the backpropagation algorithm, and the gradient information is corrected by combining the screening technology performance correlation features. The corrected gradient information is then compared with the preset... The learning rate parameter is used to fine-tune and update the network layer weight parameters responsible for calculating the risk score of feature evolution patterns in the feature evolution prediction model. Key feature combinations, key feature evolution patterns, screening technology performance feature vectors, and corresponding actual health status outcomes are added as new risk association pattern entries to the population risk knowledge base. The connection strength parameters of the risk transmission edges between relevant health status nodes are adjusted in conjunction with the true positive rate, false positive rate, and detection rate indicators. Based on the updated population risk knowledge base, the historically accumulated individualized risk knowledge graph is re-evaluated and enhanced to generate an upgraded version of the population risk knowledge base and feature evolution prediction model for subsequent screening scheme generation for target individuals.
5. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, The step of matching and analyzing the overall health representation vector with risk association patterns in a pre-constructed group risk knowledge base to construct an individualized risk knowledge graph for the target individual includes: extracting multiple preset dimensions of health element intensity values from the overall health representation vector, each health element intensity value corresponding to a comprehensive health status characterization factor that integrates multi-source data information; comparing each dimension of health element intensity value with a predefined benchmark health status range in the group risk knowledge base one by one to identify abnormal health elements that exceed the corresponding benchmark health status range and their exceedance parameters; and based on the abnormal health elements and their exceedance parameters, traversing all risk association patterns in the group risk knowledge base. The process involves: calculating the pattern matching score between each risk association pattern and the current overall health representation vector; sorting all risk association patterns in the group risk knowledge base based on the pattern matching score, and selecting multiple candidate risk association patterns whose pattern matching scores exceed a preset activation threshold; parsing the source health state node identifier, target health state node identifier, and risk transmission rule description connecting the source health state node identifier and the target health state node identifier contained in each candidate risk association pattern; extracting the current health element intensity value corresponding to the source health state node identifier from the overall health representation vector, and calculating the influence of the source health state node on the target health state based on the risk transmission rule description. The potential impact strength value of the target health state node is determined; based on the potential impact strength value and the current health element strength value corresponding to the target health state node identifier, the state update value of the target health state node is calculated; based on the source health state node identifier, the target health state node identifier, the current health element strength value corresponding to the source health state node identifier, the potential impact strength value, the current health element strength value corresponding to the target health state node identifier, and the state update value, corresponding individual health state nodes and individual risk transmission edges between nodes are instantiated; all instantiated individual health state nodes and individual risk transmission edges are then processed according to their topological relationships in the candidate risk association pattern. The data is connected and combined to form a preliminary individualized risk knowledge graph that represents the network of health status associations unique to the target individual. The preliminary individualized risk knowledge graph is then checked for isolated individual health status nodes. If such nodes exist, their current health element intensity values are matched a second time with the baseline health status ranges of other health status nodes in the group risk knowledge base, attempting to supplement missing individual risk transmission edges. The supplemented preliminary individualized risk knowledge graph undergoes loop detection and redundant edge elimination operations to ensure that at most one most representative individual risk transmission edge is retained between any two individual health status nodes, ultimately forming the structured individualized risk knowledge graph.
6. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, The method of generating a set of possible state values for each healthy state node at the end of the first time slice unit, based on the state transition probability distribution and using a random sampling-based state evolution simulation method, includes: obtaining the state transition probability distribution for each healthy state node at the end of the first time slice unit, wherein the state transition probability distribution defines the probability value of the healthy state node transitioning to each of the multiple possible new states at the end of the first time slice unit; defining the cumulative probability interval in the probability space for each possible new state based on the state transition probability distribution of each healthy state node, wherein the length of the cumulative probability interval is proportional to its corresponding probability value; initializing a simulation result record table based on a preset total number of simulations to record the simulated state values of all healthy state nodes at the end of the first time slice unit; and generating a random number uniformly distributed between zero and one independently for each healthy state node for each simulation. The random number generated for each healthy state node is mapped to the cumulative probability interval space of that healthy state node, and the specific cumulative probability interval into which the random number falls is determined. Based on the cumulative probability interval into which the random number falls, the new state selected by the corresponding healthy state node at the end of the first time slice unit in this simulation is determined, and this new state is used as the candidate state value of the healthy state node in this simulation. The candidate state value obtained by each healthy state node in this simulation is recorded in the corresponding position of the simulation result record table. The operations of generating random numbers, mapping to the cumulative probability interval, determining the new state, and recording simulation results are repeated until all simulation processes of the preset total number of simulations are completed. The simulation result record table is traversed, and for each healthy state node, all candidate state values generated in all simulations are collected, and duplicate candidate state values are removed to form the set of possible state values of the healthy state node at the end of the first time slice unit.
7. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 3, characterized in that, The step of assessing the predictability of an individual's health status evolution based on the fluctuation amplitude and trend stability information of the health status evolution pattern risk score, and searching for corresponding screening cycle adjustment parameters from the screening strategy rule base based on the predictability, includes: extracting the fluctuation amplitude and trend stability information corresponding to each health status evolution pattern risk score from the dynamic risk feature vector; calculating the average value of the fluctuation amplitude information of all health status evolution pattern risk scores to obtain an overall fluctuation amplitude index; calculating the average value of the trend stability information of all health status evolution pattern risk scores to obtain an overall trend stability index; comparing the overall fluctuation amplitude index with a preset fluctuation amplitude threshold, and comparing the overall trend stability index with a preset trend stability threshold; and based on the comparison results... The process involves determining the volatility level of the overall volatility index and the trend stability level of the overall trend stability index; based on a preset mapping table between volatility levels and trend stability levels, mapping the determined volatility level and trend stability level to obtain the predictability level of the target individual's health status evolution; searching the screening strategy rule base for the original screening cycle adjustment parameter directly associated with the predictability level of the health status evolution based on the predictability level of the health status evolution; obtaining the target individual's basic screening cycle parameter and applying the original screening cycle adjustment parameter to the basic screening cycle parameter; and calculating the final effective screening cycle adjustment parameter according to preset adjustment rules, whereby the screening cycle adjustment parameter is used to modify the time interval between two subsequent screening operations.
8. The personalized intelligent screening method for colorectal cancer based on a multi-source fusion big data model according to claim 1, characterized in that, In the simulation module of the first time slice unit of the feature evolution prediction model, the initial state feature representation of the healthy state node is normalized, and the connection strength parameter of the risk transmission edge is normalized. Based on the normalized initial state feature representation of the healthy state node and the normalized connection strength parameter of the risk transmission edge, the state transition probability distribution of each healthy state node at the end of the first time slice unit is calculated, including: normalizing the initial state feature representation of all healthy state nodes and normalizing the connection strength parameter of all risk transmission edges; for each healthy state node in the individualized risk knowledge graph, all risk transmission edges with that healthy state node as the target healthy state node are identified, and these risk transmission edges are called incoming edges; the normalized initial state feature representation of the source healthy state node corresponding to each incoming edge and the normalized connection strength parameter of the incoming edge are obtained, and based on the normalized initial state feature representation of the source healthy state node corresponding to each incoming edge and the normalized connection strength parameter of the incoming edge, the probability distribution of the incoming edge on the target healthy state node in the current time slice unit is calculated. The system generates a state transition tendency value for each node; it then aggregates the state transition tendency values generated by all incoming edges pointing to the target healthy state node to obtain the total state transition tendency value for the target healthy state node; it defines a set of new states that the target healthy state node may transition to within the first time slice unit, including the possibility of maintaining the current initial state characteristic representation; it assigns a basic transition probability parameter to each possible new state, the basic transition probability parameter being obtained based on prior knowledge in the group risk knowledge base; it converts the total state transition tendency value of the target healthy state node into a transition probability increment consistent with the additivity of the basic transition probability parameter through a preset scaling function, and then superimposes the converted transition probability increment onto the basic transition probability parameter of each possible new state according to a preset weight allocation ratio to obtain the unnormalized probability value of each possible new state; it inputs the unnormalized probability values of all possible new states into a probability normalization processing function, and converts the unnormalized probability values of all possible new states into standard probability values that sum to one, forming the state transition probability distribution of the healthy state node at the end of the first time slice unit.
9. A personalized intelligent screening system for colorectal cancer based on a multi-source fusion big data model, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the personalized intelligent colorectal cancer screening method based on a multi-source fusion big data model as described in any one of claims 1 to 8 by executing the machine-executable instructions.
Citation Information
Patent Citations
Artificial intelligence-driven medical diagnosis and treatment data processing method and system
CN120998475A
Cerebral stroke multi-mode early screening intelligent evaluation system based on large model
CN121148666A