Tumor early screening process analysis method and system based on big data technology
By building a multi-hop correlation network and performing screening weight analysis, the problems of incomplete data collection and lack of intelligent optimization of the screening process in traditional tumor screening methods are solved, and efficient and accurate early tumor screening is achieved.
Patent Information
- Application Number
- CN202510465524.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Traditional tumor screening methods have problems such as incomplete data collection, single analysis methods and lack of intelligent optimization of screening processes, which affects the accuracy and efficiency of screening results.
The tumor early screening process analysis method based on big data technology is adopted. By obtaining historical screening data, a multi-hop correlation network is built, screening weight analysis is carried out, data acquisition weight is determined, tumor risk assessment is carried out, and the secondary screening process is optimized based on the evaluation results.
It significantly improves the intelligence level of tumor screening, achieves efficient and accurate early tumor screening, and improves screening efficiency and accuracy.
Smart Images

Figure CN119993545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of early tumor screening, and in particular to a method and system for analyzing the process of early tumor screening based on big data technology. Background Art
[0002] Early screening of tumors is of great significance for improving the survival rate of cancer patients and reducing treatment costs. However, traditional tumor screening methods still have certain limitations, such as incomplete screening data collection, single analysis methods, and lack of intelligent optimization of the screening process, which affect the accuracy and efficiency of screening results.
[0003] At present, tumor screening mainly relies on biomarker detection, imaging scanning and pathological detection. However, the data correlation between different screening methods has not been fully explored, resulting in insufficient reliability of screening results. In addition, traditional screening processes mostly rely on fixed screening strategies and cannot be dynamically adjusted according to individual characteristics of patients and historical screening data, making the screening process lack flexibility and pertinence.
[0004] Knowledge graph technology can mine the correlations between different screening data, thereby optimizing the screening process; machine learning technology can train models based on a large amount of historical screening data to improve the accuracy of tumor risk assessment; data mining technology can discover potential tumor risk factors and improve screening efficiency.
[0005] Therefore, the research on the tumor early screening process analysis method and system based on big data technology has important application value. In summary, there is an urgent need for a tumor early screening process analysis method and system based on big data technology to overcome the shortcomings of traditional screening methods, improve the intelligence level of tumor screening, and achieve efficient and accurate early tumor screening. Summary of the invention
[0006] In order to solve at least one of the above technical problems, the present invention proposes a method and system for analyzing the process of early tumor screening based on big data technology.
[0007] The first aspect of the present invention provides a method for analyzing a tumor early screening process based on big data technology, comprising: Obtain the historical screening data of a preset number of cancer early screening personnel, and build a multi-hop association network between the historical screening data based on knowledge graph technology; Performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining a data acquisition weight for each screening data item according to the screening weight; Acquire screening data of tumor early screening personnel according to the data acquisition weights, and perform tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; The secondary screening process for early cancer screening personnel is optimized based on the tumor risk assessment results.
[0008] In this solution, the historical screening data of a preset number of tumor early screening personnel are obtained, and a multi-hop association network between the historical screening data is constructed based on the knowledge graph technology, specifically: Obtaining historical screening data of a preset number of tumor early screening personnel, the historical screening data including biomarker detection data, image scanning feature data and pathological index data, and performing structured processing on the historical screening data to obtain structured screening data, the structured screening data including screening item type, screening time series, screening results and medical record data; Based on the knowledge graph technology, the entity type and relationship type of the structured screening data are defined, wherein the entity type includes personnel identification, screening items, biomarkers, imaging features and pathological indicators, and the relationship type includes the temporal dependency between screening items, the correlation between biomarkers and imaging features, and the causal relationship between pathological indicators and screening results; Building a graph database for the entity types and relationship types based on big data technology, importing the entity types and relationship types into the graph database, and building a knowledge graph network with screening items as nodes and multi-hop association paths as edges; Traversing the knowledge graph network based on a breadth-first search algorithm to extract multiple association paths of broad entity types, wherein the multi-hop association paths include a first-hop association from a biomarker to an imaging feature, a second-hop association from an imaging feature to a pathological indicator, and a third-hop association of temporal dependency across screening items; Determining a data calling relationship between each screening item based on the occurrence frequency of the multi-hop association path, wherein the data calling relationship includes automatically calling the screening item data corresponding to the second-hop association path when the first-hop association path is triggered, and adjusting the execution order of the screening items according to the third-hop association path; The data call relationship is mapped to directed edge weights in a multi-hop association network to generate a multi-hop association network including screening items, screening item data nodes, call relationship edges and weight labels.
[0009] In this solution, the screening weight analysis is performed on the historical screening data according to the multi-hop association network, and the data acquisition weight of each screening data item is determined according to the screening weight, specifically: Constructing a decision tree for each multi-hop path in the multi-hop association network according to the multi-hop association network, and determining the number of calls and path depth of the screening data item corresponding to each multi-hop path node in the early tumor screening process according to the decision tree; Determining the data call frequency of each screening data item in the multi-hop association network according to the call count and the path depth; The screening weight of each screening data item is determined according to the data calling frequency, and the data acquisition weight of each screening data item is determined according to the screening weight.
[0010] In this solution, the screening data of the tumor early screening personnel are obtained according to the data acquisition weights, and the tumor risk assessment is performed on the tumor early screening personnel according to the screening data to obtain the tumor risk assessment results, which are specifically: Based on the PCA algorithm, the screening data features of historical screening data are extracted, the screening data features are mapped to the corresponding pathological indicators, and the screening data feature-pathological indicator data matrix is constructed; Performing a tumor risk probability assessment on each pathological indicator data, integrating the tumor risk probability with the screening data feature-pathological indicator data matrix, and constructing a screening data feature-pathological indicator data-tumor risk probability matrix; Introducing a gradient boosting tree algorithm to construct a tumor risk assessment model, importing the screening data feature-pathological index data-tumor risk probability matrix into the gradient boosting tree algorithm to construct a decision tree, and training the tumor risk assessment model according to the decision tree, wherein the tumor risk assessment model also includes a feature extraction and fusion layer constructed based on a PCA algorithm; Collecting screening data of tumor early screening personnel based on the data acquisition weights, the screening data including real-time biomarker detection data, image scanning feature data and pathological index time series data; Inputting the biomarker detection data and the image scanning feature data into the feature extraction and fusion layer of the tumor risk assessment model to perform cross-modal association analysis, extracting the association pattern between the biomarker concentration change and the image feature, generating a fusion feature vector, and constructing the fusion feature vector as an input set of the tumor risk assessment model; The input set is imported into the tumor risk assessment model to perform tumor risk assessment to obtain a tumor risk assessment result.
[0011] In this solution, the input set is introduced into the tumor risk assessment model to perform tumor risk assessment to obtain a tumor risk assessment result, specifically: Perform multi-dimensional feature decomposition based on the fused feature vector, extract a feature subset required for tumor risk probability calculation, and import the feature subset into the tumor risk assessment model to calculate a tumor risk probability value; When the tumor risk probability value is greater than a preset risk probability threshold, the biomarker features and image features with a fluctuation amplitude greater than the preset fluctuation threshold in the fusion feature vector are extracted and calibrated as fluctuation features, and the fluctuation features are matched with the feature library of confirmed cases in the historical screening data for secondary verification. When the matching degree is greater than the matching threshold, a high-risk warning signal is generated; According to the high-risk warning signal, the decision tree splitting rule of the tumor risk assessment model is adjusted, the splitting weight of the screening data item corresponding to the fluctuation feature is increased to a preset priority, and the tumor risk probability value is recalculated. When the recalculated probability value is continuously greater than the preset risk probability threshold, the tumor risk assessment result containing the high-risk feature marker is output; When the recalculated probability value falls below the preset risk probability threshold, the current screening data and risk assessment process are imported into the manual review system, and the split weights of the screening data items corresponding to the fluctuation characteristics of the tumor risk assessment model are secondary corrected according to the manually fed back tumor risk probability results to generate the final tumor risk assessment results.
[0012] In this scheme, the secondary screening process for early cancer screening personnel is optimized according to the tumor risk assessment results, specifically: Obtaining the high-risk probability value and its associated screening items marked in the tumor risk assessment result, and when the high-risk probability value continues to exceed the preset risk threshold, extracting the screening items that have a third-hop association path with the high-risk probability value from the multi-hop association network as priority screening items; Generate a screening item combination according to the execution sequence of the priority screening items, monitor the risk increase changes of the priority screening items in three consecutive assessment cycles in real time, and trigger the screening time compression strategy when it is detected that the risk increase of any priority screening item exceeds the preset increase threshold; Based on the screening time compression strategy, the execution frequency of the screening items with risk increase exceeding the threshold is optimized, and the execution interval of the screening items is shortened to a preset shortest period. At the same time, according to the correlation between the image features of the second-hop association path in the multi-hop association network and the pathological index, the acquisition resolution of the image scanning feature data is adjusted; When the risk increase of the screening items after frequency optimization falls below the preset increase threshold in the subsequent two assessment cycles, the execution sequence of the original screening items is restored and the secondary screening process data is generated.
[0013] The second aspect of the present invention further provides a tumor early screening process analysis system based on big data technology, the system comprising: a memory, a processor, the memory comprising a tumor early screening process analysis method program based on big data technology, and when the tumor early screening process analysis method program based on big data technology is executed by the processor, the following steps are implemented: Obtain the historical screening data of a preset number of cancer early screening personnel, and build a multi-hop association network between the historical screening data based on knowledge graph technology; Performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining the data acquisition weight of each screening data item according to the screening weight; Acquire screening data of tumor early screening personnel according to the data acquisition weights, and perform tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; The secondary screening process for early cancer screening personnel is optimized based on the tumor risk assessment results.
[0014] The present invention discloses a tumor early screening process analysis method and system based on big data technology, aiming to improve the efficiency and accuracy of early tumor screening. The method includes: obtaining historical screening data of tumor early screening personnel, constructing a multi-hop association network based on knowledge graph technology; performing screening weight analysis based on the association network to determine the data acquisition weight; obtaining screening data based on the weight, and performing tumor risk assessment; optimizing the secondary screening process according to the assessment results. The system includes data acquisition, knowledge graph construction, weight analysis, risk assessment and process optimization modules. The present invention combines big data with knowledge graphs to realize intelligent and personalized screening processes, significantly improve screening efficiency and accuracy, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A flowchart of a tumor early screening process analysis method based on big data technology of the present invention is shown; Figure 2 A flow chart showing the present invention for determining the data acquisition weight for each screening data item; Figure 3 A flow chart showing the optimization of the secondary screening process for early cancer screening personnel according to the present invention; Figure 4 A block diagram of a tumor early screening process analysis system based on big data technology of the present invention is shown. DETAILED DESCRIPTION
[0016] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0017] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0018] Figure 1 A flow chart of a tumor early screening process analysis method based on big data technology of the present invention is shown.
[0019] like Figure 1As shown, the first aspect of the present invention provides a method for analyzing the process of early cancer screening based on big data technology, comprising: S102, obtaining historical screening data of a preset number of cancer early screening personnel, and constructing a multi-hop association network among the historical screening data based on knowledge graph technology; S104, performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining a data acquisition weight for each screening data item according to the screening weight; S106, obtaining screening data of the tumor early screening personnel according to the data acquisition weight, and performing tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; S108, optimizing the secondary screening process for early cancer screening personnel based on the tumor risk assessment results.
[0020] It should be noted that by obtaining the historical screening data of a preset number of early cancer screening personnel and constructing a multi-hop association network between historical screening data based on knowledge graph technology, the structured storage and deep association mining of screening data are realized, which improves the availability and analytical value of screening data. The multi-hop association network is used to analyze the screening weight of historical screening data, and the data acquisition weight of each screening data item is determined according to the screening weight, so as to optimize the data call strategy, reduce redundant data collection, and improve the accuracy and utilization of screening data. Based on the optimized weight, the screening data of early cancer screening personnel are obtained, and the tumor risk assessment is performed by combining biomarker detection data, image scanning feature data and pathological index data. The PCA feature extraction and gradient boosting tree model are used to improve the accuracy and robustness of risk assessment, so as to achieve individualized risk prediction. Furthermore, the secondary screening process of high-risk populations is optimized according to the results of tumor risk assessment. By analyzing the screening items related to high-risk features in the multi-hop association network, the execution order and time interval of the screening items are adjusted, and the image scanning resolution and screening frequency are dynamically optimized to improve the pertinence and efficiency of screening, and finally achieve more accurate and efficient early cancer screening, improve the early detection rate and screening effect.
[0021] According to an embodiment of the present invention, the historical screening data of a preset number of cancer early screening personnel are obtained, and a multi-hop association network between the historical screening data is constructed based on the knowledge graph technology, specifically: Obtaining historical screening data of a preset number of tumor early screening personnel, the historical screening data including biomarker detection data, image scanning feature data and pathological index data, and performing structured processing on the historical screening data to obtain structured screening data, the structured screening data including screening item type, screening time series, screening results and medical record data; Based on the knowledge graph technology, the entity type and relationship type of the structured screening data are defined, wherein the entity type includes personnel identification, screening items, biomarkers, imaging features and pathological indicators, and the relationship type includes the temporal dependency between screening items, the correlation between biomarkers and imaging features, and the causal relationship between pathological indicators and screening results; Building a graph database for the entity types and relationship types based on big data technology, importing the entity types and relationship types into the graph database, and building a knowledge graph network with screening items as nodes and multi-hop association paths as edges; Traversing the knowledge graph network based on a breadth-first search algorithm to extract multiple association paths of broad entity types, wherein the multi-hop association paths include a first-hop association from a biomarker to an imaging feature, a second-hop association from an imaging feature to a pathological indicator, and a third-hop association of temporal dependency across screening items; Determining a data calling relationship between each screening item based on the occurrence frequency of the multi-hop association path, wherein the data calling relationship includes automatically calling the screening item data corresponding to the second-hop association path when the first-hop association path is triggered, and adjusting the execution order of the screening items according to the third-hop association path; The data call relationship is mapped to directed edge weights in a multi-hop association network to generate a multi-hop association network including screening items, screening item data nodes, call relationship edges and weight labels.
[0022] It should be noted that the construction of a multi-hop association network can achieve efficient organization of tumor screening data, intelligent association analysis and optimization of the screening process, thereby improving the accuracy of screening and the utilization of screening resources. The multi-hop association network is a data association structure built based on knowledge graph technology, in which screening items, biomarkers, imaging features, pathological indicators, etc. are used as nodes, and the temporal dependency between screening items, the association between biomarkers and imaging features, and the causal relationship between pathological indicators and screening results are used as directed edges to form a multi-level network structure. The "multi-hop" feature of the network is reflected in the fact that the association between data nodes is not a simple one-to-one direct association, but information is transmitted through multiple entity nodes to achieve logical reasoning and information expansion of screening data. The first hop association (biomarker-imaging feature): When certain biomarkers (such as alpha-fetoprotein AFP, carcinoembryonic antigen CEA) are abnormal, it may mean that the patient has a potential risk of malignant lesions, and changes in such biomarkers are often related to imaging features (such as tumor volume changes and enhancement features in MRI or CT scan results). Therefore, when the biomarker data is abnormal, the system can automatically trigger an imaging scan to improve the pertinence of imaging diagnosis and reduce unnecessary imaging examinations. Second-hop association (imaging features-pathological indicators): Imaging examinations may show suspicious features of tumors, such as unclear boundaries, vascular hyperplasia, abnormal enhancement patterns, etc., but the imaging results themselves cannot fully diagnose the malignancy of the tumor. Therefore, the system automatically associates and matches the most relevant pathological test items, such as biopsy histological examination, immunohistochemical analysis, etc., according to the abnormality of the imaging features, to ensure that pathological tests are performed preferentially for high-risk patients and avoid too many unnecessary invasive examinations. Third-hop association (temporal dependency across screening items): There is a temporal dependency between different screening items. For example, a patient had no abnormalities in the low-dose CT (LDCT) screening results 6 months ago, but the recent biomarker data showed a significant increase in alpha-fetoprotein. Based on the temporal dependency, the system can identify the change trend between the normal state of the patient's past screening results and the current abnormal markers, automatically adjust the screening plan, and arrange high-precision imaging examinations (such as PET-CT) or blood circulating tumor DNA (ctDNA) testing in advance to improve the timeliness of risk assessment. The big data technology includes building distributed computing databases, data warehouses, etc. Graph databases are databases used to store and manage tumor screening data and their complex relationships. They use a graph data structure, consisting of nodes, edges, and attributes, and can efficiently store and query entities such as screening items, biomarkers, imaging features, and pathological indicators and their relationships. Compared with traditional relational databases (such as MySQL and PostgreSQL), graph databases are more suitable for processing medical screening data with complex relationships and multi-hop associations.
[0023] Figure 2 A flow chart of determining the data acquisition weight of each screening data item according to the present invention is shown.
[0024] According to an embodiment of the present invention, the screening weight analysis is performed on the historical screening data according to the multi-hop association network, and the data acquisition weight of each screening data item is determined according to the screening weight, specifically: S202, constructing a decision tree for each multi-hop path in the multi-hop association network according to the multi-hop association network, and determining the number of calls and path depth of the screening data item corresponding to each multi-hop path node in the early cancer screening process according to the decision tree; S204, determining the data calling frequency of each screening data item in the multi-hop association network according to the calling times and the path depth; S206, determining a screening weight for each screening data item according to the data call frequency, and determining a data acquisition weight for each screening data item according to the screening weight.
[0025] It should be noted that by calculating the data call frequency of different screening data items in multi-hop paths, high-impact and high-relevance screening data can be identified, thereby improving the utilization of screening data and reducing redundant data calls. For example, when a biomarker (such as AFP) is abnormal, the system can intelligently match its associated imaging scan items (such as MRI) and pathological indicators (such as Ki-67 immunohistochemistry) to ensure the priority execution of necessary screening items. Secondly, the system can dynamically adjust the data acquisition strategy according to the screening weights of different screening items to improve the flexibility and efficiency of the screening process. For high-risk populations, the system will automatically increase the data acquisition weights of key screening items (such as ctDNA analysis, PET-CT) to enhance early diagnosis capabilities; for low-risk populations, unnecessary imaging scans or pathological tests are reduced, and only key biomarkers are periodically monitored, thereby reducing screening costs and patient burdens. In addition, this method can adaptively optimize the execution order of screening items based on real-time calculation of screening weights, avoid invalid screening, ensure that the screening process is more accurate and efficient, and improve the success rate of early tumor screening.
[0026] According to an embodiment of the present invention, the screening data of the tumor early screening personnel is obtained according to the data acquisition weight, and the tumor risk assessment is performed on the tumor early screening personnel according to the screening data to obtain the tumor risk assessment result, which is specifically: Based on the PCA algorithm, the screening data features of historical screening data are extracted, the screening data features are mapped to the corresponding pathological indicators, and the screening data feature-pathological indicator data matrix is constructed; Performing a tumor risk probability assessment on each pathological indicator data, integrating the tumor risk probability with the screening data feature-pathological indicator data matrix, and constructing a screening data feature-pathological indicator data-tumor risk probability matrix; Introducing a gradient boosting tree algorithm to construct a tumor risk assessment model, importing the screening data feature-pathological index data-tumor risk probability matrix into the gradient boosting tree algorithm to construct a decision tree, and training the tumor risk assessment model according to the decision tree, wherein the tumor risk assessment model also includes a feature extraction and fusion layer constructed based on a PCA algorithm; Collecting screening data of tumor early screening personnel based on the data acquisition weights, the screening data including real-time biomarker detection data, image scanning feature data and pathological index time series data; Inputting the biomarker detection data and the image scanning feature data into the feature extraction and fusion layer of the tumor risk assessment model to perform cross-modal association analysis, extracting the association pattern between the biomarker concentration change and the image feature, generating a fusion feature vector, and constructing the fusion feature vector as an input set of the tumor risk assessment model; The input set is imported into the tumor risk assessment model to perform tumor risk assessment to obtain a tumor risk assessment result.
[0027] It should be noted that feature extraction of historical screening data based on the PCA algorithm can effectively reduce dimensionality and remove redundant information, screen out key features that are strongly correlated with pathological indicators, and form a screening data feature-pathological indicator data matrix, which maps and associates multidimensional screening data (such as biomarker concentration, image scanning texture features) with pathological results (such as histological grade, malignancy), providing structured input for subsequent model training. The tumor risk assessment model constructed by introducing the gradient boosting tree algorithm can adaptively learn the nonlinear relationship between different screening features and tumor risk, and optimize the feature weights layer by layer using the splitting rule of the decision tree, thereby capturing the complex association pattern between the dynamic changes of biomarkers and abnormal imaging features. In the real-time evaluation stage, the model performs cross-modal association analysis through feature extraction and fusion layers (i.e., neural network modules that integrate biomarker detection data and image scanning feature data), such as jointly modeling the mutation frequency of circulating tumor DNA (ctDNA) in serum and the tumor volume growth rate in MRI images, generating a fusion feature vector (i.e., a numerical representation that comprehensively reflects the correlation of multi-dimensional screening indicators), so that the model can identify early tumor risk signals that are difficult to detect with single modality data. In addition, the model can optimize the risk assessment path for specific screening data of different patients by dynamically adjusting the splitting priority of feature subsets. For example, when a patient's imaging features show suspicious lesions but the biomarkers are not significantly abnormal, the model will automatically enhance the weight of the imaging features in the decision tree and make probability corrections based on the pathological results of similar cases in historical data to avoid missed diagnosis.
[0028] According to an embodiment of the present invention, the step of importing the input set into the tumor risk assessment model to perform tumor risk assessment and obtain a tumor risk assessment result is specifically as follows: Perform multi-dimensional feature decomposition based on the fused feature vector, extract a feature subset required for tumor risk probability calculation, and import the feature subset into the tumor risk assessment model to calculate a tumor risk probability value; When the tumor risk probability value is greater than a preset risk probability threshold, the biomarker features and image features with a fluctuation amplitude greater than the preset fluctuation threshold in the fusion feature vector are extracted and calibrated as fluctuation features, and the fluctuation features are matched with the feature library of confirmed cases in the historical screening data for secondary verification. When the matching degree is greater than the matching threshold, a high-risk warning signal is generated; It should be noted that due to individual differences and physiological fluctuations in biomarkers (such as ctDNA concentration) and imaging features (such as the degree of tumor edge enhancement), relying solely on risk probability thresholds can easily lead to false positives (such as misjudging the increase in AFP caused by inflammation as liver cancer risk) or missed detections (such as the imaging features of early tumors are blurred but the biomarkers are already abnormal). For example, a patient has a short-term increase in alpha-fetoprotein (AFP) due to short-term liver function abnormalities. If an early warning is triggered only based on the risk probability threshold, it may cause unnecessary invasive examinations; while another patient, although the overall risk probability does not reach the threshold, has a step-by-step increase in the number of circulating tumor cells (CTC) and a slight abnormality in PET-CT metabolic values. This collaborative fluctuation pattern may be ignored by traditional models. Therefore, by extracting cross-modal features with fluctuation amplitudes exceeding the threshold (such as a sudden increase of 50% in the coefficient of variation of biomarkers accompanied by enhanced image texture heterogeneity), and matching patterns with historical confirmed case libraries (such as comparing the AFP fluctuation curve of patients diagnosed with liver cancer 3 months before diagnosis with enhanced CT arterial phase enhancement features), physiological fluctuations and malignant lesion characteristics can be effectively distinguished. When a patient's fusion features match 80% of the pre-diagnosis feature trajectories in the historical liver cancer case library, the system will still generate a high-risk warning even if the initial risk probability value only slightly exceeds the threshold, thereby solving the clinical pain point that "early signals of occult tumors are easily diluted by the overall probability" and improving the ability to capture progressive malignant changes.
[0029] According to the high-risk warning signal, the decision tree splitting rule of the tumor risk assessment model is adjusted, the splitting weight of the screening data item corresponding to the fluctuation feature is increased to a preset priority, and the tumor risk probability value is recalculated. When the recalculated probability value is continuously greater than the preset risk probability threshold, the tumor risk assessment result containing the high-risk feature marker is output; It should be noted that the tumor risk assessment model may not be able to dynamically capture the cross-modal fluctuation features that emerge in the screening data of specific patients and are strongly associated with malignant tumors under static weight allocation. For example, when the abundance of ctDNA mutations in patients increases by 50% and MRI shows accelerated nodule enhancement, such coordinated fluctuations may indicate early malignant transformation, but if the model defaults to placing the weight of imaging features at a lower priority (such as a secondary split node), the risk may be underestimated. By increasing the split weights of fluctuation features (such as ctDNA mutations and MRI enhancement rates) to the preset priority (such as adjusting the Gini coefficient gain weight of node splitting from the third to the first), the model will prioritize the data division based on these features and strengthen their risk contribution. The recalculation of probability is intended to verify the persistence of feature abnormalities: if the adjusted risk value is still exceeded, it indicates that the fluctuation feature is not occasional noise (such as inflammatory interference), but is related to the stability of tumor progression. This mechanism overcomes the defect of traditional models' insufficient sensitivity to sudden malignant signals, ensures that accurate warnings are triggered at the early stage of feature abnormalities, avoids the risk of missed diagnosis due to parameter solidification, and improves the ability to capture progressive malignant transformation in dynamic screening.
[0030] When the recalculated probability value falls below the preset risk probability threshold, the current screening data and risk assessment process are imported into the manual review system, and the split weights of the screening data items corresponding to the fluctuation characteristics of the tumor risk assessment model are secondary corrected according to the manually fed back tumor risk probability results to generate the final tumor risk assessment results.
[0031] It should be noted that when the risk probability falls below the threshold after the model adjusts the weight, the initial fluctuation feature may be an occasional interference (such as detection error or short-term physiological abnormality). For example, a patient's CA19-9 biomarker suddenly increases and is accompanied by changes in pancreatic morphology on CT images. After the weight adjustment, the risk value briefly exceeds the threshold, but because the patient has a history of chronic pancreatitis, the indicator naturally falls back after reexamination. At this time, if the model results are directly adopted, it may lead to misjudgment due to over-reliance on the algorithm (mislabeling inflammatory lesions as pancreatic cancer risks), and may miss special malignant features that the model has not learned (such as the specific fluctuation pattern of rare neuroendocrine tumors). Manual review introduces the experience and judgment of clinicians. When the doctor confirms that the fluctuation feature is not related to malignant tumors, the system will reduce the split weight of the relevant features (such as reducing the weight attenuation coefficient of the association path between CA19-9 and pancreatic morphology from 0.8 to 0.3) to prevent subsequent similar interference data from being over-responded; if the doctor finds malignant signs that the model does not recognize (such as specific gene mutations combined with imaging calcification patterns), the weight of the new feature combination will be reversely enhanced. This dynamic correction mechanism avoids the accumulation of false positives caused by the algorithm's "overconfidence".
[0032] Figure 3 A flow chart showing the present invention for optimizing the secondary screening process for early cancer screening personnel.
[0033] According to an embodiment of the present invention, the secondary screening process for early cancer screening personnel is optimized according to the tumor risk assessment results, specifically: S302, obtaining the high-risk probability value and its associated screening items marked in the tumor risk assessment result, and when the high-risk probability value continues to exceed the preset risk threshold, extracting the screening items that have a third-hop association path with the high-risk probability value from the multi-hop association network as priority screening items; S304, generating a screening item combination according to the execution sequence of the priority screening items, monitoring the risk increase changes of the priority screening items in three consecutive evaluation cycles in real time, and triggering a screening time compression strategy when it is detected that the risk increase of any priority screening item exceeds a preset increase threshold; S306, optimizing the execution frequency of the screening items with risk increase exceeding the threshold value based on the screening time compression strategy, shortening the execution interval of the screening items to a preset shortest period, and adjusting the acquisition resolution of the image scanning feature data according to the correlation between the image features of the second hop association path in the multi-hop association network and the pathological index; S308, when the risk increase of the screening item after execution frequency optimization falls below the preset increase threshold in the subsequent two evaluation cycles, the execution sequence of the original screening item is restored and secondary screening process data is generated.
[0034] It should be noted that the extraction of third-hop associated screening items based on multi-hop association networks (such as automatically associating the characteristics of small nodules that were not paid attention to in the low-dose CT scan half a year ago when the circulating tumor cell count of a patient is abnormally increased) can deeply explore the hidden risk associations across time dimensions, ensure that the priority screening items accurately lock the early biological behavior of potential malignant lesions, and avoid the blindness of traditional fixed screening items. Secondly, by real-time monitoring of risk increase and triggering screening time compression strategies (such as shortening the MRI review interval from 6 months to 3 months), the tracking frequency of high-risk features is dynamically increased, combined with adaptive adjustment of image resolution, while ensuring screening sensitivity, optimizing medical resource allocation, and avoiding excessive examination of low-risk patients. Finally, when the risk increase falls back, the original time series mechanism is restored (such as stopping high-frequency puncture biopsy after the risk indicator is stable), realizing the flexible management of the screening process, which not only prevents the physical and mental burden of patients caused by long-term high-frequency screening, but also provides closed-loop feedback for model iteration by continuously generating secondary screening process data, thereby improving the clinical applicability of the tumor early screening system and the early interception ability of malignant lesions as a whole. The correlation between imaging features and pathological indicators refers to the strength of statistical or clinical correlation between lesion characteristics (such as tumor morphology, enhancement pattern) observed by medical imaging technology (such as CT, MRI) and pathological examination results (such as histological classification, degree of malignancy), which is obtained by correlation analysis through logistic regression.
[0035] According to an embodiment of the present invention, it also includes: Obtaining tumor risk assessment result data of the tumor early screening population, determining the screening urgency score of the secondary screening process for each tumor early screening person according to the tumor risk assessment result data, and constructing a dynamic priority queue according to the screening urgency score; Obtain the real-time schedule data and doctor scheduling data of the hospital's tumor early screening equipment, and establish a resource availability matrix including equipment idle time, doctor workload, and time consumption of early screening examination items based on the real-time schedule data and doctor scheduling data; A dual-objective optimization model of screening urgency and resource constraints is constructed based on a linear programming algorithm, the dynamic priority queue and resource availability matrix are imported into the dual-objective optimization model, and the dynamic priority queue and resource availability matrix input in real time are matched based on an online reinforcement learning algorithm to determine the examination time of each early screening item for each early cancer screening person, and obtain a distribution plan for early cancer screening equipment; The execution status of the assigned inspection task equipment in the early cancer screening equipment allocation plan and the newly added high-risk inspection case data of the hospital are monitored in real time. When a critical case is detected and inserted into the dynamic priority queue, the insertion position of the priority queue is determined according to the criticality of the case, and a resource adjustment plan for the early cancer screening equipment is obtained.
[0036] It should be noted that in the field of early cancer screening, the existing medical resource scheduling system generally adopts static scheduling rules, which makes it difficult to dynamically respond to the contradiction between the changes in the urgency of high-risk population screening and the real-time fluctuations of medical resources, resulting in problems such as long waiting time for high-risk case examinations, low equipment utilization, and uneven workload of radiologists. The present invention realizes efficient dynamic allocation of screening resources by constructing a dynamic priority queue driven by screening urgency score, combining multi-dimensional constraint modeling of resource availability matrix, and bi-objective optimization model based on linear programming and reinforcement learning. The scheme can quickly respond to the screening needs of high-risk cases, significantly shorten the waiting time for examinations, and ensure that critical cases get high-precision imaging equipment resources first through real-time monitoring of equipment status and critical case insertion mechanism, thereby improving equipment utilization and response speed. In addition, the resource availability matrix integrates multi-dimensional data such as equipment idle time, doctor load and examination time, and combines the online learning ability of reinforcement learning to optimize the workload distribution of radiologists, reduce the occupancy rate of key resources by low-risk cases, and realize efficient utilization of resources. The present invention significantly improves the hospital screening throughput under the premise of ensuring the timeliness of screening.
[0037] Figure 4 A block diagram of a tumor early screening process analysis system based on big data technology of the present invention is shown.
[0038] The second aspect of the present invention further provides a tumor early screening process analysis system 4 based on big data technology, the system comprising: a memory 41, a processor 42, the memory comprising a tumor early screening process analysis method program based on big data technology, the tumor early screening process analysis method program based on big data technology when executed by the processor, implements the following steps: Obtain the historical screening data of a preset number of cancer early screening personnel, and build a multi-hop association network between the historical screening data based on knowledge graph technology; Performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining the data acquisition weight of each screening data item according to the screening weight; Acquire screening data of tumor early screening personnel according to the data acquisition weights, and perform tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; The secondary screening process for early cancer screening personnel is optimized based on the tumor risk assessment results.
[0039] The present invention discloses a tumor early screening process analysis method and system based on big data technology, aiming to improve the efficiency and accuracy of early tumor screening. The method includes: obtaining historical screening data of tumor early screening personnel, constructing a multi-hop association network based on knowledge graph technology; performing screening weight analysis based on the association network to determine the data acquisition weight; obtaining screening data based on the weight, and performing tumor risk assessment; optimizing the secondary screening process according to the assessment results. The system includes data acquisition, knowledge graph construction, weight analysis, risk assessment and process optimization modules. The present invention combines big data with knowledge graphs to realize intelligent and personalized screening processes, significantly improve screening efficiency and accuracy, and has broad application prospects.
[0040] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0041] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0042] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0043] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0044] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0045] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A tumor early screening process analysis method based on big data technology, characterized in that: The following steps are involved: Obtain the historical screening data of a preset number of cancer early screening personnel, and build a multi-hop association network between the historical screening data based on knowledge graph technology; Performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining the data acquisition weight of each screening data item according to the screening weight; Acquire screening data of tumor early screening personnel according to the data acquisition weights, and perform tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; The secondary screening process for early cancer screening personnel is optimized based on the tumor risk assessment results.
2. According to claim 1, a method for analyzing tumor early screening process based on big data technology is characterized in that: The method of obtaining the historical screening data of a preset number of tumor early screening personnel and constructing a multi-hop association network among the historical screening data based on the knowledge graph technology is as follows: Obtaining historical screening data of a preset number of tumor early screening personnel, the historical screening data including biomarker detection data, image scanning feature data and pathological index data, and performing structured processing on the historical screening data to obtain structured screening data, the structured screening data including screening item type, screening time series, screening results and medical record data; Based on the knowledge graph technology, the entity type and relationship type of the structured screening data are defined, wherein the entity type includes personnel identification, screening items, biomarkers, imaging features and pathological indicators, and the relationship type includes the temporal dependency between screening items, the correlation between biomarkers and imaging features, and the causal relationship between pathological indicators and screening results; Building a graph database for the entity types and relationship types based on big data technology, importing the entity types and relationship types into the graph database, and building a knowledge graph network with screening items as nodes and multi-hop association paths as edges; Traversing the knowledge graph network based on a breadth-first search algorithm to extract multiple association paths of broad entity types, wherein the multi-hop association paths include a first-hop association from a biomarker to an imaging feature, a second-hop association from an imaging feature to a pathological indicator, and a third-hop association of temporal dependency across screening items; Determining a data calling relationship between each screening item based on the occurrence frequency of the multi-hop association path, wherein the data calling relationship includes automatically calling the screening item data corresponding to the second-hop association path when the first-hop association path is triggered, and adjusting the execution order of the screening items according to the third-hop association path; The data call relationship is mapped to directed edge weights in a multi-hop association network to generate a multi-hop association network including screening items, screening item data nodes, call relationship edges and weight labels.
3. According to claim 1, a method for analyzing tumor early screening process based on big data technology is characterized in that: The screening weight analysis is performed on the historical screening data according to the multi-hop association network, and the data acquisition weight of each screening data item is determined according to the screening weight, specifically: Constructing a decision tree for each multi-hop path in the multi-hop association network according to the multi-hop association network, and determining the number of calls and path depth of the screening data item corresponding to each multi-hop path node in the early tumor screening process according to the decision tree; Determining the data call frequency of each screening data item in the multi-hop association network according to the call count and the path depth; The screening weight of each screening data item is determined according to the data calling frequency, and the data acquisition weight of each screening data item is determined according to the screening weight.
4. The method for analyzing the process of early cancer screening based on big data technology according to claim 1, characterized in that: The step of obtaining screening data of a tumor early screening person according to the data acquisition weights, and performing tumor risk assessment on the tumor early screening person according to the screening data to obtain a tumor risk assessment result is specifically as follows: Based on the PCA algorithm, the screening data features of historical screening data are extracted, the screening data features are mapped to the corresponding pathological indicators, and the screening data feature-pathological indicator data matrix is constructed; Performing a tumor risk probability assessment on each pathological indicator data, integrating the tumor risk probability with the screening data feature-pathological indicator data matrix, and constructing a screening data feature-pathological indicator data-tumor risk probability matrix; Introducing a gradient boosting tree algorithm to construct a tumor risk assessment model, importing the screening data feature-pathological index data-tumor risk probability matrix into the gradient boosting tree algorithm to construct a decision tree, and training the tumor risk assessment model according to the decision tree, wherein the tumor risk assessment model also includes a feature extraction and fusion layer constructed based on a PCA algorithm; Collecting screening data of tumor early screening personnel based on the data acquisition weights, the screening data including real-time biomarker detection data, image scanning feature data and pathological index time series data; Inputting the biomarker detection data and the image scanning feature data into the feature extraction and fusion layer of the tumor risk assessment model to perform cross-modal association analysis, extracting the association pattern between the biomarker concentration change and the image feature, generating a fusion feature vector, and constructing the fusion feature vector as an input set of the tumor risk assessment model; The input set is imported into the tumor risk assessment model to perform tumor risk assessment to obtain a tumor risk assessment result.
5. The method for analyzing tumor early screening process based on big data technology according to claim 4, characterized in that: The step of importing the input set into the tumor risk assessment model to perform tumor risk assessment and obtain a tumor risk assessment result is specifically as follows: Perform multi-dimensional feature decomposition based on the fused feature vector, extract a feature subset required for tumor risk probability calculation, and import the feature subset into the tumor risk assessment model to calculate a tumor risk probability value; When the tumor risk probability value is greater than a preset risk probability threshold, the biomarker features and image features with a fluctuation amplitude greater than the preset fluctuation threshold in the fusion feature vector are extracted and calibrated as fluctuation features, and the fluctuation features are matched with the feature library of confirmed cases in the historical screening data for secondary verification. When the matching degree is greater than the matching threshold, a high-risk warning signal is generated; According to the high-risk warning signal, the decision tree splitting rule of the tumor risk assessment model is adjusted, the splitting weight of the screening data item corresponding to the fluctuation feature is increased to a preset priority, and the tumor risk probability value is recalculated. When the recalculated probability value is continuously greater than the preset risk probability threshold, the tumor risk assessment result containing the high-risk feature marker is output; When the recalculated probability value falls below the preset risk probability threshold, the current screening data and risk assessment process are imported into the manual review system, and the split weights of the screening data items corresponding to the fluctuation characteristics of the tumor risk assessment model are secondary corrected according to the manually fed back tumor risk probability results to generate the final tumor risk assessment results.
6. The method for analyzing tumor early screening process based on big data technology according to claim 1, characterized in that: The secondary screening process for early cancer screening personnel is optimized according to the tumor risk assessment results, specifically: Obtaining the high-risk probability value and its associated screening items marked in the tumor risk assessment result, and when the high-risk probability value continues to exceed the preset risk threshold, extracting the screening items that have a third-hop association path with the high-risk probability value from the multi-hop association network as priority screening items; Generate a screening item combination according to the execution sequence of the priority screening items, monitor the risk increase changes of the priority screening items in three consecutive assessment cycles in real time, and trigger the screening time compression strategy when it is detected that the risk increase of any priority screening item exceeds the preset increase threshold; Based on the screening time compression strategy, the execution frequency of the screening items with risk increase exceeding the threshold is optimized, and the execution interval of the screening items is shortened to a preset shortest period. At the same time, according to the correlation between the image features of the second-hop association path in the multi-hop association network and the pathological index, the acquisition resolution of the image scanning feature data is adjusted; When the risk increase of the screening items after frequency optimization falls below the preset increase threshold in the subsequent two assessment cycles, the execution sequence of the original screening items is restored and the secondary screening process data is generated.
7. A tumor early screening process analysis system based on big data technology, characterized in that: The tumor early screening process analysis system based on big data technology includes a storage device and a processor. The storage device includes a tumor early screening process analysis method program based on big data technology. When the tumor early screening process analysis method program based on big data technology is executed by the processor, the following steps are implemented: Obtain the historical screening data of a preset number of cancer early screening personnel, and build a multi-hop association network between the historical screening data based on knowledge graph technology; Performing screening weight analysis on the historical screening data according to the multi-hop association network, and determining the data acquisition weight of each screening data item according to the screening weight; Acquire screening data of tumor early screening personnel according to the data acquisition weights, and perform tumor risk assessment on the tumor early screening personnel according to the screening data to obtain a tumor risk assessment result; The secondary screening process for early cancer screening personnel is optimized based on the tumor risk assessment results.
Citation Information
Patent Citations
Computer equipment, system, readable storage medium and medical data analysis method
CN113012803A
Construction method of liver cancer diagnosis and treatment scheme recommendation system driven by knowledge graph
CN114121295A
Knowledge graph-based liver cancer postoperative recurrence risk prediction method and device
CN114822852A
Diabetes risk early warning method based on big data analysis
CN117253614A
Personalized tumor risk assessment guidance system based on interactive medical artificial intelligence
CN118352078A