High-grade cervical cell lesion risk prediction method and system
By applying machine learning and statistical analysis technology in the early diagnosis of cervical cancer, combined with integrated learning methods and prediction threshold fine-tuning, the problem of limited accuracy and specificity of early diagnosis of cervical cancer in the prior art has been solved, and higher prediction accuracy and personalized risk prediction are achieved.
Patent Information
- Application Number
- PCT/CN2024/113554
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-08-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems with limited accuracy and specificity in the early diagnosis of cervical cancer, and there are limitations in feature extraction and selection, so the correlation between different features cannot be fully explored.
Machine learning and statistical analysis technology are used to extract and analyze a large number of patient data in a multi-dimensional feature, and the integrated learning method is used to combine the prediction results of multiple basic models to achieve personalized risk prediction through prediction threshold fine-tuning.
It improves the accuracy and stability of early prediction of cervical cancer, provides medical professionals with more reliable decision-making basis, and achieves personalized risk prediction and intervention.
Smart Images

Figure CN2024113554_30052025_PF_FP_ABST
Abstract
Description
Method and system for predicting risk of high-grade cervical cell lesions Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for predicting the risk of high-grade cervical cell lesions. Background Art
[0002] The pathogenesis of cervical cancer typically involves a long precancerous stage, characterized by high-grade lesions in cervical cells. This stage of lesions can be controlled and treated through effective screening and early diagnosis, thereby reducing the morbidity and mortality of cervical cancer. In current clinical practice, cervical cancer screening primarily relies on cervical cytology (TCT) and high-risk human papillomavirus (HPV) testing. However, single TCT or HPV testing has limitations in the early diagnosis of cervical cancer, with limited accuracy and specificity.
[0003] To improve the accuracy of early diagnosis of cervical cancer, many studies have focused on using machine learning techniques to comprehensively analyze and predict multiple features. These features include cervical cell morphology, HPV subtype test results, and patient clinical information. However, some existing studies still have limitations in feature extraction and selection, failing to fully explore the correlations between different features, thus limiting the performance of prediction models.
[0004] Summary of the Invention
[0005] To address these limitations, the present invention proposes a method and system for predicting the risk of high-grade cervical cell lesions. This system leverages machine learning and statistical analysis techniques to analyze large amounts of patient data, extracting multi-dimensional feature information. This system then utilizes an analytical prediction model for integrated prediction, and fine-tunes prediction thresholds to achieve personalized risk prediction. This system can improve the accuracy of early-stage cervical cancer predictions, providing medical professionals with a more reliable basis for decision-making.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for predicting the risk of high-grade cervical cell lesions, comprising the following steps:
[0008] Step 1: Collect the original test results of the patient and perform data desensitization to obtain the test result data;
[0009] Step 2: preprocess the test result data to obtain original analysis data;
[0010] Step 3: inputting the original analysis data into the first analysis model, the second analysis model, and the third analysis model respectively to obtain a first prediction result, a second prediction result, and a third prediction result respectively;
[0011] Step 4: Extract the feature values corresponding to the key feature types from the original analysis data as target analysis data;
[0012] Step 5: Input the first prediction result, the second prediction result, the third prediction result and the target analysis data into a fourth prediction model to obtain a fourth prediction result; the fourth prediction result is a prediction result of the risk of high-grade cervical cell lesions;
[0013] The fourth prediction model is a prediction model trained by an ensemble learning method based on the first analysis model, the second analysis model and the third analysis model.
[0014] Furthermore, the first prediction model, the second prediction model, the third prediction model and the fourth prediction model are obtained in the following manner:
[0015] S1. Obtain the original training data for training from the database module and divide it into a training set and a validation set;
[0016] The original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values;
[0017] S2. Using a machine learning algorithm to train the first prediction model, the second prediction model, and the third prediction model based on the training set, and performing performance evaluation and optimization of each model pair based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements;
[0018] The types of machine learning algorithms selected when training the first prediction model, the second prediction model, and the third prediction model are different;
[0019] S3, using the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combining the prediction results to obtain a basic training prediction result;
[0020] S4. Calculate the degree of association between each feature type in the original training data and the original data label according to the first correlation calculation rule; filter and obtain key feature types according to the first data screening rule; extract the key feature types and corresponding feature values from the original training data as target feature data; the target feature data and the corresponding original data label constitute target original data;
[0021] S5. Create a meta-classifier, and perform model training using the basic prediction results and the target feature data as input and the original data label corresponding to the target feature data as output;
[0022] S6, fine-tuning the meta-classifier trained in S5 by adjusting the prediction threshold to the preset threshold;
[0023] S7, using the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combining the prediction results to obtain a basic validation prediction result;
[0024] S8. Using the basic verification prediction results, the meta-classifier trained in S6 is subjected to performance evaluation and optimization to obtain a fourth prediction model.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] (1) By integrating multiple cervical test data, including patient age, TCT test results, HPV13 virus test results and viral load information, the risk of high-grade cervical lesions is comprehensively assessed to improve the accuracy of risk prediction;
[0027] (2) The ensemble learning method was used to effectively combine the prediction results of multiple basic models, thereby improving the accuracy and stability of the prediction of high-grade cervical lesions;
[0028] (3) When processing and analyzing data, the output data of the basic model is integrated with the original data that is highly relevant to the target; thereby integrating data from different sources to provide a more comprehensive and accurate data view;
[0029] (4) By retaining valid data in the original data, it can ensure that the model can learn and train in a larger data space, thereby improving the performance of the model; by increasing the data dimension and improving the overall quality of the data, the overall prediction effect of the model is improved;
[0030] (4) By fine-tuning the prediction threshold, it is possible to balance the accuracy and sensitivity of the prediction according to the needs of the specific application scenario and achieve personalized risk prediction and intervention.
[0031] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a flow chart of a method for predicting the risk of high-grade cervical cell lesions provided by an embodiment of the present invention.
[0033] FIG2 is a structural diagram of a high-grade cervical cell lesion risk prediction system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following specific embodiments illustrate the embodiments of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. In order to further understand the present invention, the present invention is further described in detail below in conjunction with the best embodiment.
[0035] The inventive point of the present invention is to provide a method and system for predicting the risk of high-grade cervical cell lesions. It uses machine learning and statistical analysis techniques to analyze a large amount of patient data, extracts feature information from multiple dimensions, uses analytical prediction models for integrated prediction, and realizes personalized risk prediction through fine-tuning of prediction thresholds, thereby improving the accuracy of early prediction of cervical cancer.
[0036] One aspect of the present invention is a method for predicting the risk of high-grade cervical cell lesions. Referring to FIG1 , the method comprises the following steps:
[0037] Step 1: Collect the original test results of the patient and perform data desensitization to obtain the test result data;
[0038] The data desensitization processing refers to: performing data operations according to the desensitization strategy corresponding to the desensitization level of each data field;
[0039] Step 2: preprocess the test result data to obtain original analysis data;
[0040] The data preprocessing includes data cleaning, data filling, and data conversion;
[0041] Step 3: inputting the original analysis data into the first analysis model, the second analysis model, and the third analysis model respectively to obtain a first prediction result, a second prediction result, and a third prediction result respectively;
[0042] The first analysis model, the second analysis model, and the third analysis model are all prediction models trained based on machine learning algorithms;
[0043] Step 4: Extract the feature values corresponding to the key feature types from the original analysis data as target analysis data;
[0044] The key feature types are determined during the model training process;
[0045] Step 5: Input the first prediction result, the second prediction result, the third prediction result and the target analysis data into a fourth prediction model to obtain a fourth prediction result; the fourth prediction result is a prediction result of the risk of high-grade cervical cell lesions;
[0046] The fourth prediction model is a prediction model trained by an ensemble learning method based on the first analysis model, the second analysis model and the third analysis model.
[0047] Another aspect of the present invention is a system for predicting the risk of high-grade cervical cell lesions. Referring to FIG2 , the system is composed of the following modules:
[0048] The data collection module is used to collect the original test results of patients and perform data desensitization processing on the collected data;
[0049] Database module, used to store test result data and provide data operation functions;
[0050] A data processing module is used to pre-process the test result data;
[0051] Model training module, used for training, adjusting and optimizing prediction models;
[0052] The prediction analysis module is used to predict the risk of high-grade cervical cell lesions with the help of a prediction model and obtain prediction results;
[0053] User interface module, used to present analysis results.
[0054] As an embodiment, the patient's original test results consist of personal basic information, personal life data, TCT test data, and HPV subtype test data.
[0055] The basic personal information includes: name, age, date of birth, ID number, ethnicity, occupation, nationality, user ID, current residence, telephone number, email address, height, weight, driver's license information, social security card information, and residence permit information.
[0056] The personal life data include: age of first sexual intercourse, number of years of sexual intercourse, number of previous sexual partners, reproductive system diseases of spouse, HPV infection status of spouse, family history of cancer, exercise and fitness status, menstrual cycle, age of menarche, last menstruation, and whether menopause has occurred.
[0057] The TCT test data is the TCT examination result, specifically including the following index values: NILM (Negative for Intraepithelial Lesion or Malignancy), ASC-US (Atypical Squamous Cells of Undetermined Significance), ASC-H (Atypical Squamous Cells-High-grade), LSIL (Low-grade Squamous Intraepithelial Lesion), HSIL (High-grade Squamous Intraepithelial Lesion), SCC (Squamous Cell Carcinoma), AGC-NOS (Atypical Glandular Cells-Not Otherwise Specified), CGIN (Cervical Glandular Intraepithelial Neoplasia), and AIS (Adenocarcinoma In Situ).
[0058] The HPV subtype test data is the HPV test results, and the HPV test results are the viral load data of 14 high-risk HPV types, specifically including the viral load data of the following HPV types: HPV16, HPV18, HPV 31, HPV 33, HPV 35, HPV 39, HPV 45, HPV 51, HPV 52, HPV 56, HPV 58, HPV 59, HPV 66 and HPV 68.
[0059] As an embodiment, the TCT assay data is collected in the following manner:
[0060] (1) Obtain the TCT test report image;
[0061] (2) Perform text recognition on the TCT test report image using a text recognition algorithm to obtain the TCT test report recognition result;
[0062] (3) Use regular matching to match the TCT test report recognition results to obtain TCT test data.
[0063] The steps for collecting the HPV subtype test data are the same as those for collecting the TCT test data, and will not be repeated here.
[0064] As an embodiment, the desensitization levels include high level, low level, and no desensitization; the corresponding desensitization strategies are data deletion, data replacement, and no operation, respectively.
[0065] Specifically, data fields with a high level of redaction include: name, ID number, current residence, phone number, email address, driver's license information, social security card information, and residence permit information;
[0066] Data fields with a low level of redaction include: date of birth, ethnicity, occupation, and nationality;
[0067] The desensitization level of other data fields is not desensitized.
[0068] For date of birth, replace the month and day with the corresponding numbers; for ethnicity, occupation and nationality, replace them with the corresponding numerical codes.
[0069] It is understandable that the above-mentioned desensitization strategy can protect patient privacy while ensuring that data with analytical value is not ignored.
[0070] As an embodiment, the fourth prediction result consists of prediction classification result data and result confidence data;
[0071] The predicted classification result data is of integer type, and is used to indicate whether a cervical biopsy is required for the patient;
[0072] The confidence data of the results are decimal type; the closer the value is to zero, the smaller the reference significance of the classification result is, and the closer the value is to 1, the greater the reference significance of the classification result is.
[0073] It's understandable that confidence is the degree of certainty a model has about its results. When the model's confidence is low, the model's results are of little use, and more weight should be placed on the model's judgment. When model results differ from human judgment, further medical evidence is needed to demonstrate the reliability of the human judgment.
[0074] As an embodiment, the prediction model training performed by the model training module specifically includes training of a first prediction model, a second prediction model, a third prediction model, and a fourth prediction model.
[0075] The first prediction model, the second prediction model, the third prediction model and the fourth prediction model are obtained in the following manner:
[0076] S1. Obtain the original training data for training from the database module and divide it into a training set and a validation set;
[0077] The original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values;
[0078] S2. Using a machine learning algorithm to train the first prediction model, the second prediction model, and the third prediction model based on the training set, and performing performance evaluation and optimization of each model pair based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements;
[0079] The types of machine learning algorithms selected when training the first prediction model, the second prediction model, and the third prediction model are different;
[0080] S3, using the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combining the prediction results to obtain a basic training prediction result;
[0081] S4. Calculate the degree of association between each feature type in the original training data and the original data label according to the first correlation calculation rule; filter and obtain key feature types according to the first data screening rule; extract the key feature types and corresponding feature values from the original training data as target feature data; the target feature data and the corresponding original data label constitute target original data;
[0082] S5. Create a meta-classifier, and perform model training using the basic prediction results and the target feature data as input and the original data label corresponding to the target feature data as output;
[0083] S6, fine-tuning the meta-classifier trained in S5 by adjusting the prediction threshold to the preset threshold;
[0084] S7, using the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combining the prediction results to obtain a basic validation prediction result;
[0085] S8. Using the basic verification prediction results, the meta-classifier trained in S6 is subjected to performance evaluation and optimization to obtain a fourth prediction model.
[0086] It should be noted that the purpose of fine-tuning the model by adjusting the prediction threshold in S6 is to enable the trained model to balance the accuracy and sensitivity of prediction according to the needs of specific application scenarios, thereby achieving personalized risk prediction and intervention.
[0087] As an embodiment, in S2, the machine learning algorithm may be any one of a decision tree, a random forest, a support vector machine, and a gradient boosting algorithm.
[0088] As an embodiment, the first correlation calculation rule may be implemented by a combination of one or more of the following methods: Pearson correlation coefficient, Spearman rank, Euclidean distance, cosine similarity, and Jaccard correlation coefficient.
[0089] As an embodiment, the first data screening rule is:
[0090] All feature types are sorted by correlation, and the feature types with the top N correlations are selected as key feature types.
[0091] As an embodiment, the first data screening rule may also be:
[0092] Feature types with a correlation greater than a preset correlation threshold are selected as key feature types.
[0093] As an embodiment, the meta-classifier can be implemented using any one of logistic regression, random forest, support vector machine, neural network, gradient boosting algorithm, and K-nearest neighbor algorithm.
[0094] As an embodiment, the method of the present invention may be implemented in software and / or a combination of software and hardware, for example, by using an application specific integrated circuit (ASIC), a general-purpose computer or any other similar hardware device.
[0095] The method of the present invention can be implemented in the form of a software program that can be executed by a processor to perform the steps or functions described above. Similarly, the software program (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like.
[0096] In addition, some steps or functions of the method of the present invention may be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0097] In addition, a portion of the method described in the present invention may be implemented as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for invoking the method described in the present invention may be stored in a fixed or removable recording medium, and / or transmitted via a data stream in a broadcast or other signal-carrying medium, and / or stored in a working memory of a computer device that operates according to the program instructions.
[0098] As an embodiment, the present invention also provides a device comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to run the methods and / or technical solutions based on the aforementioned multiple embodiments.
[0099] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0100] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0101] In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0102] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Any technician familiar with the present profession can make slight changes or modifications to equivalent embodiments using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for predicting the risk of high-grade cervical cell lesions, characterized in that: The method comprises the following steps: Step 1: Collect the original test results of the patient and perform data desensitization to obtain the test result data; Step 2: preprocess the test result data to obtain original analysis data; Step 3: input the original analysis data into the first analysis model, the second analysis model and the third analysis model respectively, and obtain the first prediction result, the second prediction result and the third prediction result respectively; Step 4: Extract feature values corresponding to key feature types from the original analysis data as target analysis data; Step 5: Input the first prediction result, the second prediction result, the third prediction result and the target analysis data into the fourth prediction model to obtain a fourth prediction result; the fourth prediction result is a risk prediction result for high-grade cervical cell lesions.
2. The method according to claim 1, characterized in that The patient's original test results consist of personal basic information, personal life data, TCT test data, and HPV subtype test data.
3. The method according to claim 2, characterized in that The data desensitization processing is to perform data operations according to the desensitization strategy corresponding to the desensitization level of each data field; The desensitization levels include high level, low level, and no desensitization; the corresponding desensitization strategies are data deletion, data replacement, and no operation.
4. The method according to claim 1, characterized in that: The first analysis model, the second analysis model and the third analysis model are all prediction models trained based on machine learning algorithms, and the three are trained using different machine learning algorithms; The machine learning algorithm selects any three of decision tree, random forest, support vector machine, and gradient boosting algorithm.
5. The method according to claim 1, characterized in that The fourth prediction model is a prediction model trained by an integrated learning method based on the first analysis model, the second analysis model and the third analysis model.
6. The method according to claim 1, characterized in that The first prediction model, the second prediction model, the third prediction model and the fourth prediction model are obtained in the following manner: S1. Obtain original training data for training from the database module and divide it into a training set and a validation set; The original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values; S2. Based on the training set, a machine learning algorithm is used to train the first prediction model, the second prediction model, and the third prediction model respectively, and performance evaluation and optimization of each model pair are performed based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements; S3, using the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combining the prediction results to obtain a basic training prediction result; S4. Calculate the degree of association between each feature type in the original training data and the original data label according to the first degree of association calculation rule; Filter and obtain key feature types according to the first data screening rule; Extract the key feature type and the corresponding feature value from the original training data as target feature data; S5. Create a meta-classifier, and use the basic prediction result and the target feature data as input, and the original data label corresponding to the target feature data as output to perform model training; S6, fine-tuning the meta-classifier trained in S5 by adjusting the prediction threshold to the preset threshold; S7, using the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combining the prediction results to obtain a basic validation prediction result; S8. Using the basic verification prediction results, the meta-classifier trained in S6 is subjected to performance evaluation and optimization to obtain a fourth prediction model.
7. The method according to claim 6, characterized in that The first correlation calculation rule can be implemented by a combination of one or more of the following methods: Pearson correlation coefficient, Spearman rank, Euclidean distance, cosine similarity, and Jaccard correlation coefficient.
8. The method according to claim 6, characterized in that The first data screening rule is: Sort all feature types by relevance, and select the top N feature types as Key feature type.
9. The method according to claim 6, characterized in that The meta-classifier is implemented using any one of the algorithms selected from logistic regression, random forest, support vector machine, neural network, gradient boosting algorithm, and K-nearest neighbor algorithm.
10. A risk prediction system for high-grade cervical cell lesions, characterized in that , The system consists of a data collection module, a database module, a data processing module, a model training module, a prediction analysis module and a user interface module.
Citation Information
Patent Citations
Plasma sample cancer early screening method based on ensemble learning
CN113611404A
Method for identifying ground glass pulmonary nodules
CN114692748A
Cervical cell high-level lesion risk prediction method and system
CN117727452A
Cervical cancer diagnosis method and apparatus using artificial intelligence-based medical image analysis and software program therefor
US20210090248A1
Cervical cancer screening support system, cervical cancer screening support method, recording medium carrying cervical cancer screening support program, and smartphone built with smartphone application carrying cervical cancer screening support program
US20230281815A1