Method and System for Predicting the Risk of High-Grade Cervical Intraepithelial Neoplasia
By integrating multiple cervical test data and integrated learning methods, multi-dimensional feature information is extracted, and the accuracy and specificity of early diagnosis of cervical cancer in the prior art is solved, and personalized cervical cancer risk prediction and early diagnosis are achieved.
Patent Information
- Application Number
- CN202311554606.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-11-21
AI Technical Summary
Existing early diagnosis methods of cervical cancer such as TCT and HPV detection have limited accuracy and specificity, and cannot fully explore the correlation between different characteristics, limiting the performance of the prediction model.
Using machine learning and statistical analysis technology, multi-dimensional feature information is extracted through the integration and integration of multiple cervical assay data, multiple analytical models are used to predict, and personalized risk prediction is achieved through fine-tuning of prediction thresholds.
It improves the accuracy and stability of early prediction of cervical cancer, provides a more reliable decision-making basis, and achieves a more comprehensive and accurate data view to adapt to the needs of different application scenarios.
Smart Images

Figure CN117727452B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and system for predicting the risk of high-grade cervical cell lesions. Background Art
[0002] The pathogenesis of cervical cancer usually undergoes a long precancerous lesion stage, that is, high-grade lesions of cervical cells. The lesions at this stage can be controlled and treated through effective screening and early diagnosis, thereby reducing the incidence and mortality of cervical cancer. In current clinical practice, the screening of cervical cancer mainly relies on cervical cytology examination (TCT) and high-risk human papillomavirus (HPV) detection. However, single TCT or HPV detection has certain limitations in the early diagnosis of cervical cancer, and its accuracy and specificity are limited.
[0003] In order to improve the accuracy of early diagnosis of cervical cancer, many studies focus on using machine learning techniques to comprehensively analyze and predict various feature information. These feature information include cervical cell morphological features, HPV virus subtype detection results, and clinical information of patients, etc. However, there are still certain limitations in the existing feature extraction and selection, and the correlation between different features cannot be fully explored, thus limiting the performance of the prediction model. Summary of the Invention
[0004] In view of the above-mentioned limitations, the present invention proposes a method and system for predicting the risk of high-grade cervical cell lesions, which analyzes a large amount of patient data by means of machine learning and statistical analysis techniques, extracts feature information from multiple dimensions, uses an analysis and prediction model for integrated prediction, and realizes personalized risk prediction through fine-tuning of the prediction threshold. This system can improve the accuracy of early prediction of cervical cancer and provide more reliable decision-making basis for medical professionals.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for predicting the risk of high-grade cervical cell lesions, the method comprising the following steps:
[0007] Step 1, collect the original test results of patients and perform data desensitization processing to obtain test result data;
[0008] Step 2, perform data preprocessing on the test result data to obtain original analysis data;
[0009] Step 3, input the original analysis data into a first analysis model, a second analysis model, and a third analysis model respectively to obtain a first prediction result, a second prediction result, and a third prediction result;
[0010] Step 4: Extract the feature values corresponding to the key feature types from the original analysis data as the target analysis data;
[0011] Step 5: Input the first prediction result, the second prediction result, the third prediction result, and the target analysis data into the fourth prediction model to obtain the fourth prediction result; the fourth prediction result is the risk prediction result of high-grade cervical cell lesions;
[0012] The fourth prediction model is a prediction model trained by means of ensemble learning based on the first analysis model, the second analysis model, and the third analysis model.
[0013] Furthermore, the first prediction model, the second prediction model, the third prediction model, and the fourth prediction model are obtained in the following manner:
[0014] S1: Obtain the original training data for training from the database module and divide it into a training set and a validation set;
[0015] The original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values;
[0016] S2: Based on the training set, use machine learning algorithms to train the first prediction model, the second prediction model, and the third prediction model respectively, and perform performance evaluation and optimization of each model pair based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements;
[0017] The types of machine learning algorithms selected when training the first prediction model, the second prediction model, and the third prediction model are different from each other;
[0018] S3: Use the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combine the prediction results to obtain the basic training prediction result;
[0019] S4: Calculate the correlation degree between each feature type in the original training data and the original data label according to the first correlation degree calculation rule; screen and obtain the key feature types according to the first data screening rule; extract the key feature types and corresponding feature values from the original training data as the target feature data; the target feature data and the corresponding original data labels form the target original data;
[0020] S5: Create a meta-classifier, and use the basic prediction result and the target feature data as input and the original data label corresponding to the target feature data as output for model training;
[0021] S6: Fine-tune the meta-classifier trained in S5 by adjusting the prediction threshold to the preset threshold;
[0022] S7. Use the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combine the prediction results to obtain the basic validation prediction result;
[0023] S8. Evaluate and optimize the performance of the meta-classifier trained in S6 with the help of the basic validation prediction result to obtain the fourth prediction model.
[0024] Compared with the prior art, the present invention has the following advantages:
[0025] (1) By integrating various cervical test data, including patient age, TCT test results, test results of 13 types of HPV viruses, and viral load information, etc., comprehensively evaluate the risk of high-grade cervical lesions and improve the accuracy of risk prediction;
[0026] (2) Adopt the ensemble learning method, effectively combine the prediction results of multiple basic models, and improve the accuracy and stability of the prediction of high-grade cervical lesions;
[0027] (3) When performing data processing and analysis, fuse the output data of the basic model and the original data highly relevant to the target; thereby integrating data from different sources and providing a more comprehensive and accurate data view;
[0028] (4) By retaining the valid data in the original data, it can ensure that the model can learn and train in a larger data space, thereby improving the performance of the model; by increasing the data dimension and improving the overall quality of the data, the overall prediction effect of the model is improved;
[0029] (4) By fine-tuning the prediction threshold, it is possible to balance the accuracy and sensitivity of the prediction according to the needs of specific application scenarios, and achieve personalized risk prediction and intervention.
[0030] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart of a method for predicting the risk of high-grade cervical lesions provided by an embodiment of the present invention.
[0032] Figure 2 It is a structural diagram of a system for predicting the risk of high-grade cervical lesions provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention. To further understand the present invention, the following provides a detailed description of the present invention in combination with the best embodiments.
[0034] The invention point of the present invention is to provide a method and system for predicting the risk of high-grade cervical cell lesions, which analyzes a large amount of patient data by means of machine learning and statistical analysis techniques, extracts feature information from multiple dimensions, uses an analysis prediction model for integrated prediction, and realizes personalized risk prediction through fine-tuning of the prediction threshold, so as to improve the early prediction accuracy of cervical cancer.
[0035] One aspect of the present invention lies in a method for predicting the risk of high-grade cervical cell lesions, with reference to Figure 1 , the method includes the following steps:
[0036] Step 1: Collect the original test results of patients and perform data desensitization processing to obtain test result data;
[0037] The data desensitization processing refers to: performing data operations according to the desensitization strategies corresponding to the desensitization levels of each data field;
[0038] Step 2: Perform data preprocessing on the test result data to obtain original analysis data;
[0039] The data preprocessing includes data cleaning, data filling, and data conversion;
[0040] Step 3: Input the original analysis data into the first analysis model, the second analysis model, and the third analysis model respectively to obtain the first prediction result, the second prediction result, and the third prediction result;
[0041] The first analysis model, the second analysis model, and the third analysis model are all prediction models trained based on machine learning algorithms;
[0042] Step 4: Extract the feature values corresponding to the key feature types from the original analysis data as the target analysis data;
[0043] The key feature types are determined during the model training process;
[0044] Step 5: Input the first prediction result, the second prediction result, the third prediction result, and the target analysis data into the fourth prediction model to obtain the fourth prediction result; the fourth prediction result is the risk prediction result of high-grade cervical cell lesions;
[0045] The fourth prediction model is a prediction model trained by means of an ensemble learning method based on the first analysis model, the second analysis model, and the third analysis model.
[0046] Another aspect of the present invention lies in a high-grade cervical cell lesion risk prediction system, with reference to Figure 2 , the system consists of the following modules:
[0047] A data collection module, configured to collect the original test results of patients and perform data desensitization processing on the collected data;
[0048] A database module, configured to store the test result data and provide data operation functions;
[0049] A data processing module, configured to perform data preprocessing on the test result data;
[0050] A model training module, configured to train and adjust and optimize the prediction model;
[0051] A prediction and analysis module, configured to predict the high-grade cervical cell lesion risk by means of the prediction model to obtain a prediction result;
[0052] A user interface module, configured to present the analysis result.
[0053] As an embodiment, the original test results of the patients consist of personal basic information, personal life data, TCT test data, and HPV subtype test data.
[0054] The personal basic information includes: name, age, date of birth, ID number, ethnicity, occupation, nationality, user ID, current place of residence, telephone number, email, height, weight, driver's license information, social security card information, and residence permit information.
[0055] The personal life data includes: age at first sexual intercourse, number of years of sexual life, number of previous sexual partners, reproductive system diseases of the spouse, HPV infection status of the spouse, family history of tumors, exercise and fitness status, menstrual cycle, age at menarche, last menstrual period, and whether menopause has occurred.
[0056] The TCT test data are the results of TCT examinations, specifically including the following index values: NILM (Negative for Intraepithelial Lesion or Malignancy), ASC-US (Atypical Squamous Cells of Undetermined Significance), ASC-H (Atypical Squamous Cells-High-grade), LSIL (Low-grade Squamous Intraepithelial Lesion), HSIL (High-grade Squamous Intraepithelial Lesion), SCC (Squamous Cell Carcinoma), AGC-NOS (Atypical Glandular Cells-Not Otherwise Specified), CGIN (Cervical Glandular Intraepithelial Neoplasia), AIS (Adenocarcinoma In Situ).
[0057] The HPV subtype test data are the results of HPV detections. The HPV detection results are the viral load data of 14 high-risk HPV types, specifically including the viral load data of the following HPV types: HPV16, HPV18, HPV 31, HPV 33, HPV35, HPV 39, HPV 45, HPV 51, HPV 52, HPV 56, HPV 58, HPV 59, HPV 66, and HPV 68.
[0058] As an example, the TCT test data are collected by the following method:
[0059] (1) Obtain pictures of TCT test reports;
[0060] (2) Perform optical character recognition on the pictures of TCT test reports through an optical character recognition algorithm to obtain the recognition results of TCT test reports;
[0061] (3) Use regular matching to perform matching recognition on the recognition results of TCT test reports to obtain TCT test data.
[0062] The collection steps of the HPV subtype test data are the same as those of the TCT test data, and will not be elaborated here.
[0063] As an embodiment, the desensitization levels include high level, low level, and no desensitization; the corresponding desensitization strategies are data deletion, data replacement, and no operation respectively.
[0064] Specifically, the data fields with a high desensitization level include: name, ID number, current place of residence, phone number, email, driver's license information, social security card information, and residence permit information;
[0065] The data fields with a low desensitization level include: date of birth, ethnicity, occupation, and nationality;
[0066] The desensitization levels of the remaining data fields are all no desensitization.
[0067] For the date of birth, replace the month and day numbers; for ethnicity, occupation, and nationality, replace them with corresponding numerical codes.
[0068] It can be understood that adopting the above desensitization strategy can achieve the protection of patient privacy, and at the same time ensure that the data with analytical value is not ignored.
[0069] As an embodiment, the fourth prediction result consists of prediction classification result data and result confidence data;
[0070] The prediction classification result data is of integer type and is used to indicate whether cervical biopsy sampling is required for the patient;
[0071] The result confidence data is of decimal type; the closer the value is to zero, the smaller the reference significance of the classification result, and the closer the value is to 1, the greater the reference significance of the classification result.
[0072] It can be understood that the confidence level is the degree of certainty of the model for the given result. When the degree of certainty given by the model is small, the reference significance of the result given by the model is small. On the contrary, more attention should be paid to the judgment result given by the model. When the model result is inconsistent with the human judgment, more medical evidence is needed to prove the reliability of the human judgment.
[0073] As an embodiment, the model training module performs prediction model training, specifically including the training of the first prediction model, the second prediction model, the third prediction model, and the fourth prediction model.
[0074] The first prediction model, the second prediction model, the third prediction model, and the fourth prediction model are obtained in the following manner:
[0075] S1. Obtain the original training data for training from the database module and divide it into a training set and a validation set;
[0076] The original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values;
[0077] S2. Use machine learning algorithms to train the first prediction model, the second prediction model, and the third prediction model respectively based on the training set, and perform performance evaluation and optimization on each model pair based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements;
[0078] The types of machine learning algorithms selected when training the first prediction model, the second prediction model, and the third prediction model are different from each other;
[0079] S3. Use the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combine the prediction results to obtain the basic training prediction results;
[0080] S4. Calculate the correlation degree between each feature type in the original training data and the original data label according to the first correlation degree calculation rule; screen and obtain the key feature types according to the first data screening rule; extract the key feature types and the corresponding feature values from the original training data as the target feature data; the target feature data and the corresponding original data label form the target original data;
[0081] S5. Create a meta-classifier, and use the basic prediction results and the target feature data as inputs and the original data label corresponding to the target feature data as the output for model training;
[0082] S6. Fine-tune the meta-classifier trained in S5 by adjusting the prediction threshold to a preset threshold;
[0083] S7. Use the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combine the prediction results to obtain the basic validation prediction results;
[0084] S8. Evaluate and optimize the performance of the meta-classifier trained in S6 with the help of the basic validation prediction results to obtain the fourth prediction model.
[0085] It should be noted that the purpose of adjusting the prediction threshold to fine-tune the model in S6 is to enable the trained model to balance the prediction accuracy and sensitivity according to the requirements of specific application scenarios, so as to achieve personalized risk prediction and intervention.
[0086] As an embodiment, in S2, the machine learning algorithm can be any one of decision tree, random forest, support vector machine, and gradient boosting algorithm.
[0087] As an embodiment, the first correlation degree calculation rule can be implemented by a combination of one or more of Pearson correlation coefficient, Spearman rank, Euclidean distance, cosine similarity, and Jaccard correlation coefficient.
[0088] As an embodiment, the first data screening rule is as follows:
[0089] Sort all feature types according to the degree of association, and select the top N feature types with the highest degree of association as the key feature types.
[0090] As an embodiment, the first data screening rule can also be as follows:
[0091] Screen the feature types with an association degree greater than the preset association degree threshold as the key feature types.
[0092] As an embodiment, the meta-classifier can be implemented using any one of the algorithms such as logistic regression, random forest, support vector machine, neural network, gradient boosting algorithm, and K-nearest neighbor algorithm.
[0093] As an embodiment, the method of the present invention can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device.
[0094] The method of the present invention can be implemented in the form of a software program, and the software program can be executed by a processor to implement the above-mentioned steps or functions. Similarly, the software program (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices.
[0095] In addition, some steps or functions of the method of the present invention can be implemented using hardware. For example, as a circuit that cooperates with the processor to execute each step or function.
[0096] In addition, a part of the method of the present invention can be applied as a computer program product. For example, computer program instructions, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for calling the method of the present invention can be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in the working memory of a computer device that runs according to the program instructions.
[0097] As an embodiment, the present invention also provides a device, which includes a memory for storing computer program instructions and a processor for executing the program instructions. When the computer program instructions are executed by the processor, the device is triggered to run the method and / or technical solution based on the foregoing multiple embodiments.
[0098] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0099] Finally, it should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the said element.
[0100] In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of the invention claimed.
[0101] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the technical content disclosed above without departing from the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for predicting the risk of high-grade cervical cell lesions, characterized in that the method comprises the following steps: Step 1: Collect the original test results of the patient and perform data desensitization processing, wherein data deletion is performed on the ID number, current residence, and telephone number fields, and the date of birth field is replaced to obtain the test result data; the original test results of the patient consist of personal basic information, personal life data, TCT test data, and HPV subtype test data; Step 2: Perform data preprocessing on the test result data to obtain the original analysis data; Step 3: Input the original analysis data into the first analysis model, the second analysis model, and the third analysis model respectively to obtain the first prediction result, the second prediction result, and the third prediction result; Step 4: Extract the feature values corresponding to the key feature types from the original analysis data as the target feature data; Step 5: Input the first prediction result, the second prediction result, the third prediction result, and the target feature data into the fourth prediction model to obtain the fourth prediction result; the fourth prediction result consists of prediction classification result data and result confidence data, which is the risk prediction result of high-grade cervical cell lesions. When the model prediction result is inconsistent with the human judgment, more medical evidence is required to prove the reliability of the human judgment; The fourth prediction model is obtained in the following manner: S3: Use the first prediction model, the second prediction model, and the third prediction model to predict the training set respectively, and combine the prediction results to obtain the basic training prediction result; S4: Calculate the correlation degree between each feature type in the original training data and the original data label according to the first correlation degree calculation rule; the original training data consists of the original feature data and the corresponding original data label; the original feature data includes several feature types and corresponding feature values; the first correlation degree calculation rule includes one or a combination of methods such as Pearson correlation coefficient, Spearman rank, Euclidean distance, cosine similarity, and Jaccard correlation coefficient; select the key feature types according to the first data screening rule; extract the key feature types and corresponding feature values from the original training data as the target feature data; the first data screening rule is: sort all feature types according to the correlation degree, and select the feature types ranked in the top N in terms of correlation degree as the key feature types; S5: Create a meta-classifier, and use the basic training prediction result and the target feature data as the input and the original data label corresponding to the target feature data as the output for model training; S6: Fine-tune the meta-classifier trained in S5 by adjusting the prediction threshold to a preset threshold; S7: Use the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combine the prediction results to obtain the basic validation prediction result; S8: Evaluate and optimize the performance of the meta-classifier trained in S6 with the help of the basic validation prediction result to obtain the fourth prediction model.
2. The method according to claim 1, characterized in that the data desensitization processing is to perform data operations according to the desensitization strategies corresponding to the desensitization levels of each data field; The desensitization levels include high level, low level, and no desensitization; the corresponding desensitization strategies are data deletion, data replacement, and no operation respectively.
3. The method according to claim 1, wherein the first analysis model, the second analysis model, and the third analysis model are all prediction models trained based on machine learning algorithms, and the three are trained using different machine learning algorithms; the machine learning algorithms are selected from any three of decision tree, random forest, support vector machine, and gradient boosting algorithm.
4. The method according to claim 1, wherein the first prediction model, the second prediction model, and the third prediction model are obtained in the following manner: S1. Obtain the original training data for training from the database module and divide it into a training set and a validation set; S2. Respectively train the first prediction model, the second prediction model, and the third prediction model based on the training set using machine learning algorithms, and perform performance evaluation and optimization on each model pair based on the validation set to obtain the first prediction model, the second prediction model, and the third prediction model that meet the preset performance requirements.
5. The method according to claim 1, wherein the meta-classifier is implemented using any one of the algorithms of logistic regression, random forest, support vector machine, neural network, gradient boosting algorithm, and K-nearest neighbor algorithm.
6. A high-grade cervical cell lesion risk prediction system, characterized in that , The system consists of a data collection module, a data processing module, a model training module, a prediction analysis module, and a user interface module; the data collection module is used to collect the original test results of patients and perform data desensitization processing, wherein data deletion is performed on the ID number, current residence, and phone number fields, and replacement is performed on the date of birth field to obtain the test result data; the original test results of patients consist of personal basic information, personal life data, TCT test data, and HPV subtype test data; the data processing module is used to perform data preprocessing on the test result data to obtain the original analysis data; the model training module includes a first prediction model, a second prediction model, a third prediction model, and a fourth prediction model, and is used for the training and adjustment and optimization of the prediction model; specifically, input the original analysis data into the first analysis model, the second analysis model, and the third analysis model respectively to obtain the first prediction result, the second prediction result, and the third prediction result; extract the feature values corresponding to the key feature types from the original analysis data as the target feature data; The fourth prediction model is obtained in the following manner: S3. Use the first prediction model, the second prediction model, and the third prediction model to respectively predict the training set, and combine the prediction results to obtain the basic training prediction result; S4. Calculate the correlation degree between each feature type in the original training data and the original data label according to the first correlation degree calculation rule; the original training data consists of original feature data and corresponding original data labels; the original feature data includes several feature types and corresponding feature values; the first correlation degree calculation rule includes one or a combination of methods such as Pearson correlation coefficient, Spearman rank, Euclidean distance, cosine similarity, and Jaccard correlation coefficient; screen and obtain key feature types according to the first data screening rule; Extract the key feature types and corresponding feature values from the original training data as target feature data; the first data screening rule is: sort all feature types according to the correlation degree, and select the feature types ranked top N in terms of correlation degree as key feature types; S5. Create a meta-classifier, and use the basic training prediction result and the target feature data as input, and the original data label corresponding to the target feature data as output for model training; S6. Fine-tune the meta-classifier trained in S5 by adjusting the prediction threshold to a preset threshold; S7. Use the first prediction model, the second prediction model, and the third prediction model to predict the validation set respectively, and combine the prediction results to obtain the basic validation prediction result; S8. Evaluate and optimize the performance of the meta-classifier trained in S6 with the help of the basic validation prediction result to obtain the fourth prediction model; The prediction analysis module is used to predict the risk of high-grade cervical cell lesions with the help of a prediction model; specifically, input the first prediction result, the second prediction result, the third prediction result, and the target feature data into the fourth prediction model to obtain the fourth prediction result; the fourth prediction result consists of prediction classification result data and result confidence data, which is the prediction result of the risk of high-grade cervical cell lesions. When the model prediction result is inconsistent with the human judgment, more medical evidence is needed to prove the reliability of the human judgment; The user interface module is used to present the analysis result.