Gynecological tumor disease risk intelligent prediction system
By building an intelligent prediction system for risk of gynecological tumors and using big data analysis and deep learning algorithms, the limitations of traditional diagnostic methods are solved, early risk assessment and personalized prediction are achieved, and the discovery rate and treatment effect of gynecological tumors are improved.
Patent Information
- Application Number
- CN202510592394.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional gynecological tumor diagnosis methods are mostly based on the patient's obvious symptoms, which leads to the middle and late stages of the disease, which is difficult to treat and poor prognosis. There are limitations in tumor marker detection and imaging examination methods, false positive or false negative problems.
Design an intelligent prediction system for risk prediction of gynecological tumor diseases, including data collection, data storage, data analysis, feature extraction, model training and result display and feedback modules. Through joint analysis and feature extraction of big data, representative and distinctive feature subsets are screened out, and deep learning algorithms are used to train risk prediction models to provide personalized risk assessment.
Accurately capture potential risk signals in the early stages of the disease, improve the discovery rate of gynecological tumor diseases, reduce false positives and false negatives, improve the accuracy and scientific nature of medical decisions, and enhance the possibility of cure and prognostic effect.
Smart Images

Figure CN120388743A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gynecological oncology, and particularly to an intelligent prediction system for the risk of gynecological oncology diseases. Background Art
[0002] Gynecological oncology seriously threatens women's health. Common gynecological oncology diseases include cervical cancer, ovarian cancer, endometrial cancer, etc. The risks of gynecological oncology diseases for women of different age groups are different. For example, ovarian cancer and endometrial cancer are more common in middle-aged and elderly women.
[0003] Environmental factors, lifestyle, etc. in different regions may affect the incidence risk of gynecological oncology diseases. For example, in some industrial pollution areas, there may be risk factors related to environmental tumors.
[0004] Therefore, with the progress of medical technology, the popularization of electronic medical record systems, the accumulation of medical big data, and the rapid development of artificial intelligence technology, a technical foundation has been provided for the intelligent prediction of the risk of gynecological oncology diseases. The detection of tumor markers is becoming more and more accurate. For example: Traditional diagnostic methods are mostly based on examinations after obvious symptoms appear in patients. For example, it is only discovered when symptoms such as vaginal bleeding occur, or when the abdominal mass of ovarian cancer significantly increases. At this time, the condition has developed to the middle and late stages, with great treatment difficulty and poor prognosis. Although there are some tumor marker detection and imaging examination methods, there are limitations in single indicators, insufficient sensitivity in detecting early tiny lesions, and high costs.
[0005] In view of the above problems, there is an urgent need to innovate and design on the basis of the original gynecological oncology risk prediction system. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent prediction system for the risk of gynecological oncology diseases, so as to solve the problems proposed in the above background art that traditional diagnostic methods are mostly based on examinations after obvious symptoms appear in patients. At this time, the condition has entered the middle and late stages with great treatment difficulty and poor prognosis. The monitoring using tumor marker detection and imaging examination methods has limitations in obtaining a single indicator, and false positives or false negatives are likely to occur.
[0007] To achieve the above purpose, the present invention provides the following technical solution: An intelligent prediction system for the risk of gynecological oncology diseases includes a data acquisition module, a data storage module, a data analysis module, a feature extraction module, a model training module, a risk prediction module, and a result display and feedback module, which form the main architecture of the system; The data acquisition module is connected to the hospital information system, the public health database, and mobile medical devices through a local area network, and the data acquisition module synchronously uploads patient information, environmental factor data, and patient lifestyle data to the data storage module and the data analysis module; A data warehouse is established within the data storage module, and a data index is established within the data warehouse of the data storage module. The data analysis module receives the patient information uploaded by the data collection module, and the data processed by the data analysis module is uploaded into the feature extraction module. The feature extraction module screens out representative and discriminatory feature subsets and enters them into the risk prediction module. The prediction model within the risk prediction module is provided by the model training module, and the data for training by the model training module comes from the data warehouse provided by the data storage module. The data calculated by the risk prediction module is uploaded to the result display and feedback module, and the result display and feedback module is connected to the doctor's personal terminal. The result display and feedback module receives the subsequent data feedback of the patient, and the result display and feedback module uploads the feedback data into the data warehouse of the data storage module.
[0008] By adopting the above technical solution, through the association of big data joint analysis and feature extraction, the screening and analysis of gynecological tumor disease information are increased, and the calculation of the risk of suffering from tumor diseases is improved.
[0009] Preferably, the data collection module performs preliminary cleaning and sorting on the collected data, removes obviously incorrect or severely missing data, and fills in the missing data reasonably.
[0010] By adopting the above technical solution, the obtained data adopts the same data format, and at the same time, the missing data of the patient is quickly supplemented, which is convenient for the evaluation of the risk model.
[0011] Preferably, the data warehouse of the data storage module adopts a distributed storage architecture, and classifies and stores data according to different data types.
[0012] By adopting the above technical solution, the classification storage method is convenient for quickly searching relevant content data in the data warehouse.
[0013] Preferably, a specified data format and transmission protocol are set for the connection between the data collection module and the data storage module.
[0014] By adopting the above technical solution, the data transfer between the data collection module and the data storage module is more stable, and data loss is avoided.
[0015] Preferably, the association rule mining algorithm is adopted in the data analysis module, and the algorithm training of the data analysis module uses the data in the data warehouse of the data storage module.
[0016] By adopting the above technical solution, it is convenient to quickly and jointly calculate the data related to tumors.
[0017] Preferably, the feature extraction module filters the data uploaded by the data collection module, and the feature subsets of the feature extraction module and the patient information features analyzed by the data analysis module are fed back into the data warehouse of the data storage module.
[0018] By adopting the above technical solution, the necessary feature data of the patient data can be quickly selected and extracted, reducing the amount of data and facilitating the analysis and calculation of the data.
[0019] Preferably, the model training module receives the data in the data warehouse of the data storage module to train the model, and the model training module adopts a deep learning algorithm.
[0020] By adopting the above technical solution, a model more suitable for gynecological tumor risk prediction is provided for the risk prediction module.
[0021] Preferably, the risk prediction module receives the original patient data from the data collection module and the data from the data analysis module and the feature extraction module.
[0022] By adopting the above technical solution, all the obtained data are input into the model by the risk prediction module for calculation.
[0023] Preferably, the result display and feedback module receives the prediction result calculated by the risk prediction module, and the result display and feedback module presents the risk information in the form of a data report.
[0024] By adopting the above technical solution, it is used to display the prediction result data to doctors and patients.
[0025] Preferably, the result display and feedback module receives the doctor's opinions and the new information of the patient and feeds them back to the data collection module and the data storage module.
[0026] By adopting the above technical solution, by updating the new data on the patient's status, a more accurate risk prediction value can be further provided for the patient's physical condition.
[0027] Compared with the prior art, the beneficial effects of the present invention are: the intelligent risk prediction system for gynecological tumor diseases: Adopting multi-dimensional data such as the patient's clinical symptoms, family history, living habits, and various examination indicators, the data related to gynecological tumor diseases are quickly extracted by the data analysis module and the feature extraction module. Among them, the detection of hormone levels such as estrogen, progesterone, and androgen is mainly recorded. When the hormone balance is disrupted, the risk of endometrial cancer may increase. In addition, the detection results of tumor markers such as CA125, CA153, and HE4 can be recorded. These data can reflect the symptoms of gynecological tumors. Therefore, the prediction system can record the individual's information based on the recorded physical examination data and family habits, and at the same time can perform risk prediction.
[0028] By comparing data related to normal range values, it is possible to accurately capture potential risk signals when there are no obvious clinical symptoms or signs in the early stage of the disease. By increasing associated data, the detection rate of early-stage gynecological tumor diseases can be improved, winning precious treatment time for patients and greatly enhancing the possibility of cure and prognosis. By collecting unique data information for each patient and fully considering the impact of individual differences such as age, reproductive history, genetic factors, etc. on risks, a personalized risk report for gynecological tumor diseases is generated. The risk assessment results provided by the system assist doctors in a more comprehensive understanding of the patient's condition. At the same time, the result display and feedback module provides a feedback channel for the patient's new physical condition, facilitating data modification for a new round of risk prediction calculation, making the risk assessment data more in line with the patient's physical condition, helping to improve the accuracy and scientific nature of medical decisions, and reducing the occurrence of adverse medical events caused by subjective judgment errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the risk prediction process of the present invention; Figure 2 It is a schematic diagram of the system process architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an intelligent risk prediction system for gynecological tumor diseases, including: a data collection module, a data storage module, a data analysis module, a feature extraction module, a model training module, a risk prediction module, and a result display and feedback module, which form the main architecture of the system; The data collection module is connected to the hospital information system, the public health database, and mobile medical devices through a local area network. The data collection module synchronously uploads patient information, environmental factor data, and patient lifestyle data to the data storage module and the data analysis module. The data collection module performs preliminary cleaning and sorting on the collected data, removes obviously incorrect or severely missing data, and fills in missing data reasonably. The connection between the data collection module and the data storage module is set with a specified data format and transmission protocol; Combined with the accompanying drawings of the specification Figure 1 - Figure 2As shown, the data acquisition module conducts information interaction via the network, connecting to the information system within the hospital, the external public health database, and in-hospital mobile medical devices, i.e., the computer devices used by doctors. It obtains information such as the patient's basic name, age, ethnicity, marital status, etc., the past medical history, current medical history, surgical history, etc. in clinical diagnoses, and inspection information such as relevant blood test results, ultrasound examination reports, and pathological examination results from the hospital information system; The doctor inputs data on the patient's lifestyle, including data information such as smoking history, alcohol consumption, exercise frequency, eating habits, reproductive history, menstrual cycle, etc., into the mobile medical device; The public health database receives data on the local environmental pollution situation, water source quality, and the incidence of complex tumors within the region related to the patient's visiting area; The data acquisition module performs preliminary cleaning and sorting on the collected data, removes data with obvious errors or severe missing values, and uses the mean filling algorithm to supplement the missing data to ensure the integrity and availability of the data. Then, it stores the data in a specific database table or data file in the data storage module according to the established data format and transmission protocol, providing a data basis for the data analysis and feature extraction module; A data warehouse is established within the data storage module, and a data index is established within the data warehouse of the data storage module. The data warehouse of the data storage module adopts a distributed storage architecture and stores data by different data types; Combined with the attached drawings of the specification Figure 1 - Figure 2 As shown, the data warehouse of the data storage module classifies the received data, divides the data into types according to themes such as the patient basic information database, inspection and test result database, gynecological oncology case database, etc., and establishes a data index to improve the data query efficiency; The data analysis module receives the patient information uploaded by the data acquisition module, and the data processed by the data analysis module is uploaded into the feature extraction module. The feature extraction module screens out representative and discriminative feature subsets and enters them into the risk prediction module. The association rule mining algorithm is used in the data analysis module, and the algorithm training of the data analysis module uses the data in the data warehouse of the data storage module. The feature extraction module screens from the data uploaded by the data acquisition module, and the feature subsets of the feature extraction module and the patient information features analyzed by the data analysis module are fed back into the data warehouse of the data storage module; Combined with the attached drawings of the specification Figure 1 - Figure 2As shown, the data analysis module and the feature extraction module sequentially receive the normalized data. The data analysis module uses the association rule mining algorithm to take the patient's clinical diagnosis information, examination and test information, and lifestyle information as the transaction data set, and sets the minimum support and minimum confidence thresholds. Using the Apriori algorithm process, by scanning the data in the data warehouse of the data storage module, counting the frequency of each single occurrence, and finding the single items that meet the minimum support threshold, these single items form the frequent 1-item set L. The candidate (k + 1)-item set (Ck+1) is generated through the frequent k-item set (Lk). The specific method is to combine the item sets in Lk pairwise and remove the combinations that cannot become the frequent (k + 1)-item set according to a certain pruning strategy; Scan the data warehouse of the data storage module again, calculate the support of the candidate (k + 1)-item set, find the frequent (k + 1)-item set Lk+1 that meets the minimum support threshold, and repeat this process until no new frequent item sets can be generated. For each frequent item set, generate association rules by splitting the item set, and screen out the association rules that meet the threshold requirements according to the previously set confidence threshold. Finally, evaluate the screened association rules and consider their practical significance for the patient's disease probability in the prediction of gynecological tumor diseases, so as to obtain the main analysis data of the patient's data; The feature extraction module selects the ground cabinet feature elimination algorithm for feature selection, inputs all features into the vector machine for training, calculates the importance score of each feature, and then gradually eliminates the feature with the lowest importance score. After each elimination, retrain the vector machine and calculate the importance score of the remaining features. Repeat this process until the predetermined number of features is reached or the stop condition is met. Select the most representative 10 - 15 features from the initial dozens of features including age, tumor marker level, lifestyle factors, etc., such as: age, CA125 level, smoking years, family tumor history, etc. as the final feature subset for model training; The prediction model in the risk prediction module is provided by the model training module, and the data used for training by the model training module comes from the data warehouse of the data storage module. The model training module receives the data in the data warehouse of the data storage module to train the model, and the model training module uses the deep learning algorithm. The risk prediction module receives the patient's original data from the data collection module and the data from the data analysis module and the feature extraction module; The model in the model training module is trained using the data in the data warehouse of the data storage module. When obtaining the training data, the collected data is processed for format unification, and the data in different formats is converted into a standard format suitable for analysis; For the numerical missing data in the dataset, the mean value of the counts of patients of the same age and gender is used for filling. For the missing categorical data, the most common mode among the patients in the region is used for filling. At the same time, outliers are detected and processed. Using a method based on statistical distribution, data outside the range of plus or minus 3 times the standard deviation is marked as abnormal, and the abnormal data is replaced using trimming or deletion methods; After that, feature processing is performed on the data, and the processed data is divided into a training set, a validation set, and a test set according to a certain proportion; The training set data is input into the logistic regression model, and the loss function value between the predicted value of the model and the true label is calculated. The calculation of this loss function value usually uses the cross-entropy loss function: where is the number of training samples, is the true label of the th sample, is the probability that the model predicts the th sample has a gynecological tumor; According to the loss value, the weight coefficients of the model are updated and optimized using the gradient descent algorithm. The gradient descent formula is: where is the th weight coefficient, and is the learning rate, which controls the step size of each weight update; Then the data is sent into the decision tree model for training. Starting from the root node, for each node, the best feature is selected for splitting according to metrics such as information gain or Gini impurity; The information gain calculation formula is: where is the dataset of the current node, is the feature to be selected, is the number of values of the feature is the subset where the feature takes the value of is the entropy of the dataset is the entropy of the subset; The Gini impurity calculation formula is: where is the set of categories, is the probability that the data in the dataset belongs to the category ; Repeatedly perform the operation of splitting nodes until the predetermined maximum depth is reached or the number of node samples is less than the set threshold; Based on the set number of decision trees, each decision tree grows based on a random subset of the training set (and a randomly selected subset of features). The training process of each tree is similar to the above decision tree training; After all decision trees are trained, input the validation set data into the random forest model and calculate evaluation metrics such as the accuracy rate and recall rate of the model on the validation set; If the model performance does not meet expectations, adjust hyperparameters such as the number of decision trees in the random forest, the maximum depth of each tree, and the feature selection ratio, and retrain until a model with better and stable performance is obtained; Finally, input the training set data into the neural network, calculate the output value of the network through forward propagation, and then calculate the loss function value between the output value and the true label. Commonly used loss functions include the mean squared error loss function: where is the output value of the model prediction for the i-th sample; According to the loss value, use the backpropagation algorithm to calculate the gradients of the connection weights and biases of each layer of neurons, and then use the gradient descent algorithm to update and optimize the weights and biases. During the training process, set the learning rate to gradually decrease as the number of training rounds increases to balance the convergence speed and accuracy of the model; After training each batch of data, calculate the loss value and evaluation metrics of the model on the validation set, and adjust model parameters such as the number of hidden layer nodes, learning rate, and regularization coefficient according to the metrics to prevent overfitting; The training process continues until the performance of the model on the validation set no longer improves or reaches the predetermined number of training rounds. After each round of training, calculate evaluation metrics such as the accuracy rate, recall rate, and F1 value of the model on the validation set, and judge whether the model has overfitting phenomenon according to these metrics. If the performance of the model on the validation set no longer improves and begins to decline, stop training, save the optimal parameters of the current model, and obtain a preliminary model; When evaluating the model's metrics, use the test set data and calculate evaluation metrics including the accuracy rate: where is true positive, is true negative, is false positive, is false negative; The recall rate is calculated as: The F1 value is calculated as: Among them is as follows A variety of evaluation metrics are used to comprehensively evaluate the prediction performance of the model. In addition to accuracy, recall, and F1 value, it also includes the Receiver Operating Characteristic curve, i.e., the ROC curve, and the area under the curve, i.e., the AUC value. This metric can comprehensively measure the classification performance of the model at different thresholds. The closer the AUC value is to 1, the better the prediction performance of the model; After the calculation formula evaluation model is completed, the complete model is input into the risk prediction module. The risk prediction module inputs the patient data to be evaluated into the model. Through model calculation, the risk probability prediction value of the patient suffering from gynecological tumors is obtained, and the evaluation data upload result is displayed and fed back to the module; The data upload result display and feedback module calculates and obtains the data of the risk prediction module. The result display and feedback module is connected to the doctor's personal terminal. The result display and feedback module receives the subsequent data feedback of the patient, and the result display and feedback module uploads the feedback data into the data warehouse of the data storage module. The result display and feedback module receives the doctor's opinions and the patient's new information and feeds them back to the data collection module and the data storage module; Thus, the module predicts the gynecological cancer risk data of this individual based on the recorded relevant data of gynecological tumors. And this module has a self-correction function to prevent large errors in the judgment of tumor data.
[0032] In the medical information system at the doctor's end, a bar chart is used to display the risk level distribution of patients. Different colors are used to distinguish low, medium, and high risks. At the same time, the various risk factors of patients and their corresponding weights are listed in detail in the form of a table. When the gynecologist explains the patient's current physical condition to the patient, the actual condition obtained from the further diagnosis of the patient is compared with the prediction result. The missing data collected by the original data collection module is filled according to the patient's description, and the original mean data is replaced to facilitate the model to re-evaluate the physical condition, so as to obtain a risk assessment prediction result that is more in line with the patient's physical data. During the subsequent diagnosis and treatment of the patient, new symptoms and new examination data are uploaded, and the relevant data is uploaded by the doctor's mobile medical device. The feedback information will be stored in the database and regularly used for the re-training and optimization of the model to continuously improve the prediction accuracy and reliability of the system.
[0033] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
Claims
1. An intelligent prediction system for the risk of gynecological tumor diseases, characterized in that, It includes: A data acquisition module, a data storage module, a data analysis module, a feature extraction module, a model training module, a risk prediction module, and a result display and feedback module form the main architecture of the system; The data acquisition module is connected to the hospital information system, the public health database, and mobile medical devices through a local area network, and the data acquisition module synchronously uploads patient information, environmental factor data, and patient lifestyle data to the data storage module and the data analysis module; A data warehouse is established in the data storage module, and a data index is established in the data warehouse of the data storage module; The data analysis module receives the patient information uploaded by the data acquisition module, and the data processed by the data analysis module is uploaded to the feature extraction module. The feature extraction module filters out representative and discriminative feature subsets and enters the risk prediction module; The prediction model in the risk prediction module is provided by the model training module, and the data for training by the model training module comes from the data warehouse provided by the data storage module; The data calculated by the risk prediction module is uploaded to the result display and feedback module, and the result display and feedback module is connected to the doctor's personal terminal. The result display and feedback module receives subsequent data feedback from patients, and the result display and feedback module uploads the feedback data to the data warehouse of the data storage module.
2. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, wherein: The data acquisition module performs preliminary cleaning and sorting on the collected data, removes data with obvious errors or severe missing values, and fills in the missing data reasonably.
3. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: The data warehouse of the data storage module adopts a distributed storage architecture to classify and store data according to different data types.
4. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: A specified data format and transmission protocol are set for the connection between the data acquisition module and the data storage module.
5. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: An association rule mining algorithm is adopted in the data analysis module, and the algorithm training of the data analysis module uses the data in the data warehouse of the data storage module.
6. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, wherein: The feature extraction module filters from the data uploaded by the data acquisition module, and the feature subsets of the feature extraction module and the patient information features analyzed by the data analysis module are fed back into the data warehouse of the data storage module.
7. An intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: The model training module receives the data in the data warehouse of the data storage module to train the model, and the model training module adopts a deep learning algorithm.
8. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: The risk prediction module receives the original patient data from the data acquisition module and the data from the data analysis module and the feature extraction module.
9. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, wherein: The result display and feedback module receives the prediction results calculated by the risk prediction module, and the result display and feedback module displays the risk information in the form of a data report.
10. The intelligent prediction system for the risk of gynecological tumor diseases according to claim 1, characterized in that: The result display and feedback module receives the doctor's opinions and new patient information and feeds them back to the data acquisition module and the data storage module.