Cell segmentation and recognition system and method based on HE stained pathological sections
By evaluating the usability of the pre-selected model and retraining it, and combining feature extraction and stitching architecture, the problem of insufficient accuracy in cell segmentation and recognition in panoramic digital slices of gastric cancer was solved, achieving higher accuracy in cell segmentation and recognition.
Patent Information
- Application Number
- CN202310819839.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-05
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-07-05
AI Technical Summary
Existing pan-cell segmentation and classification models are not accurate enough in panoramic digital slices of gastric cancer, making it difficult to effectively distinguish tumor cells with high cellular heterogeneity and morphological differences, resulting in inaccurate cell segmentation and further affecting the accuracy of cell identification.
The usability of the pre-trained pre-selected model is evaluated, and a decision is made on whether to retrain the model based on the evaluation results. By setting up a target feature extractor and a feature splicing architecture, the target features of HE-stained pathological sections are extracted to achieve cell segmentation and classification.
It improves the accuracy of cell segmentation and recognition, enhances the quality and training efficiency of the model, and ensures the accuracy of segmentation and classification results.
Smart Images

Figure CN116798033B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical pathology analysis technology, and in particular to a cell segmentation and identification system and method based on HE-stained pathological sections. Background Technology
[0002] Currently, the latest pan-cell segmentation and classification models cannot achieve high accuracy in panoramic digital slices of gastric cancer (average accuracy is less than 74%). This is because gastric cancer is a tumor with very high cellular heterogeneity; cells of the same type can exhibit significant morphological differences in different regions, and cell nuclei within the tumor are often clustered together, making them difficult to distinguish.
[0003] The invention patent with application number CN202211528536.1 discloses a method, apparatus, and device for image segmentation and model training. This invention inputs the candidate category names of the image segmentation scene / task along with the image into an image segmentation model. The image segmentation model automatically maps the input category names to text embedding vectors in a unified category representation space, extracts image features, and performs image segmentation based on the image features and text embedding vectors to obtain the image's location mask and the corresponding category information. This eliminates the need to manually establish unified category names and category representation spaces, making it applicable to various image segmentation scenes / tasks using different category systems. This improves the generalization ability and robustness of the image segmentation model, and enhances the accuracy of image segmentation.
[0004] However, existing technologies do not evaluate the training quality of image segmentation models. When the training quality of image segmentation models is not high, cell segmentation is not accurate enough, and furthermore, subsequent cell recognition is also not accurate enough.
[0005] In view of this, a solution is urgently needed to address the above problems. Summary of the Invention
[0006] One of the objectives of this invention is to provide a cell segmentation and recognition system based on HE-stained pathological sections. The system evaluates the usability of pre-trained pre-selected models and retrains pre-selected models that are deemed unusable, thereby improving the quality and training efficiency of the models and further enhancing the accuracy of cell segmentation and recognition.
[0007] The cell segmentation and identification system based on HE-stained pathological sections provided in this embodiment of the invention includes:
[0008] The evaluation module is used to evaluate the usability of pre-trained pre-selected models and obtain the evaluation results;
[0009] The retraining module is used to use the corresponding pre-selected model as the cell segmentation and classification model if the evaluation result is usable; otherwise, it obtains the training record set of the corresponding pre-selected model and retrains the cell segmentation and classification model based on the training record set.
[0010] The settings module is used to configure the target feature extractor;
[0011] The feature determination module is used to determine the target features in HE-stained pathological slide images based on the target feature extractor and the acquired target task.
[0012] The segmentation and recognition module is used to obtain segmentation and classification results based on the cell segmentation and classification model and the target features corresponding to the target task.
[0013] Preferably, the settings module includes:
[0014] The base model determination submodule is used to determine the base model of the preset base feature extractor;
[0015] The stride setting submodule is used to set the stride of the convolutional layers in the base model to 1.
[0016] The reconnection submodule is used to remove the max-pooling layer in the base model and reconnect the unconnected layer ports after the max-pooling layer is removed to obtain the target feature extractor.
[0017] Preferably, the feature determination module includes:
[0018] The task type acquisition submodule is used to acquire the task type of the target task, which includes: cell nucleus segmentation and cell nucleus classification.
[0019] The first parameter determination submodule is used to determine the first constraint parameters of the target feature extractor when the task type is cell nucleus segmentation.
[0020] The second parameter determination submodule is used to determine the second constraint parameters of the target feature extractor when the task type is cell nucleus classification.
[0021] The layer feature determination submodule is used to determine the layer features of different feature layers extracted by the target feature extractor after being constrained by the first constraint parameter or the second constraint parameter.
[0022] The layer feature splicing submodule is used to splice the layer features of each feature layer according to the preset splicing architecture to obtain the target feature.
[0023] Preferably, the evaluation module includes:
[0024] The pathology image set acquisition submodule is used to acquire manually annotated pathology image sets;
[0025] The parsing submodule is used to parse each pathological image in the pathological image set to obtain the annotation result set;
[0026] The first calculation submodule is used to obtain the prediction result set output after each pathological image is input into the pre-selected model, and to calculate the first number of prediction results in the prediction result set.
[0027] The second calculation submodule is used to determine the second number of prediction results in the prediction result set that meet the accuracy standard based on the labeled result set, the prediction result set, and the preset accuracy standard.
[0028] The third calculation submodule is used to calculate the cell classification accuracy based on the first number and the second number;
[0029] The first accuracy determination submodule is used to evaluate the result as usable if the cell classification accuracy is greater than or equal to a preset first threshold.
[0030] The second accuracy determination submodule is used to determine the evaluation result as unusable if the cell classification accuracy is less than a preset first threshold.
[0031] Preferred, the retraining module includes:
[0032] The splitting submodule is used to split the training record set based on the difference in the first labeler of the first training record in the training record set, and obtain multiple training record subsets.
[0033] The traversal submodule is used to sequentially traverse the first training record in the training record subset.
[0034] The associated value retrieval submodule is used to retrieve the associated value between the first record type and the target record type of the first training record during each iteration.
[0035] The second marker determination submodule is used to determine the first marker corresponding to the subset of training records being traversed, and to use it as the second marker;
[0036] The first annotation experience value acquisition submodule is used to acquire the first annotation experience value of the second annotator corresponding to the first training record of the first record type;
[0037] The target training sample determination submodule is used to determine the target training samples based on the association value and the first annotation empirical value.
[0038] The sample input submodule is used to input the target training samples into a preset neural network model to obtain a cell segmentation and classification model.
[0039] Preferably, the associated value retrieval submodule includes:
[0040] The analysis target determination unit is used to obtain a template based on preset analysis targets and determine the first and second analysis targets for correlation analysis.
[0041] An analysis item acquisition unit is used to acquire a first analysis item for a first analysis target and a second analysis item for a second analysis target.
[0042] The cross-analysis pairing determination unit is used to determine multiple first cross-analysis pairings based on the first analysis item and the second analysis item;
[0043] The pairing value acquisition unit is used to acquire the pairing value of the first cross-analysis pairing item based on the first element information of the first analysis item and the second element information of the second analysis item corresponding to the first cross-analysis pairing item.
[0044] The cross-analysis pairing selection unit is used to select the corresponding first cross-analysis pairing as the second cross-analysis pairing if the pairing value is greater than or equal to a preset third threshold.
[0045] The cross-analysis result acquisition unit is used to acquire the cross-analysis results of the second cross-analysis pairings;
[0046] The correlation value determination unit is used to determine the correlation value based on the cross-analysis results.
[0047] Preferably, the first annotation experience value acquisition submodule includes:
[0048] The historical annotation record acquisition unit is used to acquire at least one second record type of the historical annotation records of the second marker.
[0049] The record count calculation unit is used to calculate the number of historical annotation records for each second record type;
[0050] The experience level determination unit is used to query a preset record number-experience level comparison database to determine the experience level.
[0051] The marking seniority acquisition unit is used to acquire the marking seniority of the second marker.
[0052] The second annotation experience value determination unit is used to determine the second annotation experience value of the second marker corresponding to the historical annotation record of the second record type based on experience level and marking years;
[0053] The first annotation experience value determination unit is used to take the corresponding second annotation experience value as the first annotation experience value if the second record type is consistent with the first record type.
[0054] Preferably, the second annotation experience value determination unit includes:
[0055] The first seniority record determination sub-unit is used to determine the first seniority record of multiple employees based on the acquired data tagging industry employment information database.
[0056] The second seniority record calculation subunit is used to analyze and obtain reliable second seniority records from the first seniority record according to a preset reliability analysis template.
[0057] The correction factor acquisition sub-unit is used to calculate the average of the recorded years of service corresponding to the second years of service record, and to obtain the correction factor by dividing the marked years of service by the average.
[0058] The second annotation experience value calculation subunit is used to calculate the second annotation experience value of the historical annotation record corresponding to the second record type by the second annotator, based on the experience level and the correction factor.
[0059] Preferably, the target training sample determination submodule includes:
[0060] The reference degree determination unit is used to query the preset association value-reference degree library and determine the reference degree of the association value;
[0061] The judgment value acquisition unit is used to fuse the reference degree and the first annotation experience value to obtain the judgment value;
[0062] The second training record acquisition unit is used to take the corresponding first training record as the second training record if the judgment value is greater than or equal to the preset fourth threshold.
[0063] The third number calculation unit is used to calculate the third number of the second training record;
[0064] Model parameter acquisition module, used to obtain the model parameters of the pre-selected model;
[0065] The sample number calculation unit is used to estimate the model based on the model parameters and the preset sample number, and to determine the target sample number.
[0066] The target training sample determination unit is used to determine the target training sample if the third number reaches the target number of samples.
[0067] The supplementary unit is used to supplement the third training record with the corresponding number of third training records according to the difference between the third number and the target number of samples if the third number does not reach the target number of samples. The third training record and the corresponding second training record are used together as the target training sample.
[0068] The cell segmentation and identification method based on HE-stained pathological sections provided in this invention includes:
[0069] Step 1: Evaluate the usability of the pre-trained pre-selected model and obtain the evaluation results;
[0070] Step 2: If the evaluation result is usable, use the corresponding pre-selected model as the cell segmentation and classification model; otherwise, obtain the training record set of the corresponding pre-selected model and retrain the cell segmentation and classification model based on the training record set.
[0071] Step 3: Set up the target feature extractor;
[0072] Step 4: Based on the target feature extractor, determine the target features in the HE-stained pathological slide images according to the acquired target task;
[0073] Step 5: Based on the cell segmentation and classification model, obtain the segmentation and classification results according to the target features corresponding to the target task.
[0074] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings.
[0075] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0076] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0077] Figure 1 This is a schematic diagram of a cell segmentation and identification system based on HE-stained pathological sections in an embodiment of the present invention;
[0078] Figure 2 This is a schematic diagram of a cell segmentation and identification method based on HE-stained pathological sections in an embodiment of the present invention. Detailed Implementation
[0079] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0080] This invention provides a cell segmentation and identification system based on HE-stained pathological sections, such as... Figure 1 As shown, it includes:
[0081] Evaluation module 1 is used to evaluate the usability of pre-trained pre-selected models and obtain evaluation results. The pre-selected models are neural network models trained using multiple manual records of cell segmentation and classification until convergence. The usability evaluation of the pre-trained pre-selected models is as follows: the prediction accuracy of the pre-selected models is evaluated. If the accuracy is lower than a preset value, it is determined to be unusable; otherwise, it is usable.
[0082] Retraining module 2 is used to use the corresponding pre-selected model as the cell segmentation and classification model if the evaluation result is usable; otherwise, it obtains the training record set of the corresponding pre-selected model and retrains the cell segmentation and classification model based on the training record set. The training record set is a collection of model training records of the pre-selected model that needs to be retrained.
[0083] Module 3 is used to set the target feature extractor; the target feature extractor is used to extract features from cell images, which is achieved through feature extraction technology.
[0084] Feature determination module 4 is used to determine target features in HE-stained pathological slide images based on the target feature extractor and the acquired target task; the target task is, for example, cell segmentation and cell identification; the target features are the features required to perform the target task.
[0085] The segmentation and recognition module 5 is used to obtain segmentation and classification results based on the cell segmentation and classification model and the target features corresponding to the target task. The segmentation results include, for example, identifying where the cell nucleus is and where tissue fluid is; the classification results include identifying the cell type of the identified cell nucleus, such as tumor cells, lymphocytes, and stromal cells.
[0086] The working principle and beneficial effects of the above technical solution are as follows:
[0087] The pre-trained models are evaluated for usability. If the evaluation result indicates that the pre-selected model is usable, it is directly used as the cell segmentation and classification model. Otherwise, the training record set of the corresponding pre-selected model is obtained. Generally, poor model training quality does not mean that all training data is unusable. Therefore, the cell segmentation and classification model is retrained using the training record set (e.g., removing training data that does not meet the training requirements from the training record set and retraining). This greatly improves the quality of the model and the training efficiency. Then, based on the target feature extractor and the target task, the target features are determined and input into the cell segmentation and classification model to obtain the segmentation and classification results.
[0088] This application performs usability assessment on pre-trained pre-selected models, and retrains pre-selected models that are deemed unusable by the assessment results, thereby improving the quality and training efficiency of the models and further enhancing the accuracy of cell segmentation and recognition.
[0089] In one embodiment, the setting module includes:
[0090] The base model determination submodule is used to determine the base model of the preset base feature extractor; the preset base feature extractor is: a feature extraction device based on Inception v3 as the base network; the base model is: the network model of Inception v3.
[0091] The stride setting submodule is used to set the stride of the convolutional layers in the base model to 1. A convolutional layer consists of several convolutional submodules. The stride of the convolutional layer represents the extraction accuracy; the longer the stride, the higher the extraction accuracy.
[0092] The reconnection submodule removes the max-pooling layers from the base model and reconnects the unconnected layer ports after the removal of the max-pooling layers to obtain the target feature extractor. Max-pooling layers are used to reduce the dimensionality of information extracted by convolutional layers and also enhance the invariance of image features. Removing the max-pooling layers is beneficial for extracting image features from the cell nucleus layer.
[0093] The working principle and beneficial effects of the above technical solution are as follows:
[0094] This application resets the stride of the convolutional layers in the basic feature extractor and removes the max pooling layer, increasing the size of the feature image while preserving important low-level information of cell nucleus segmentation, thus improving the accuracy of subsequent recognition.
[0095] In one embodiment, the feature determination module includes:
[0096] The task type acquisition submodule is used to acquire the task type of the target task. The task type includes: cell nucleus segmentation and cell nucleus classification. Cell nucleus segmentation is: segmenting cell nuclei in histological images; cell nucleus classification is: classifying cell nuclei in histological images.
[0097] The first parameter determination submodule is used to determine the first constraint parameter of the target feature extractor when the task type is cell nucleus segmentation; the first constraint parameter is: the parameter that constrains the target feature extractor to extract only the features used for cell nucleus segmentation;
[0098] The second parameter determination submodule is used to determine the second constraint parameters of the target feature extractor when the task type is cell nucleus classification; the second constraint parameter is: the parameter that constrains the target feature extractor to extract only the features used for cell nucleus classification;
[0099] The layer feature determination submodule is used to determine the layer features of different feature layers extracted by the target feature extractor after being constrained by the first constraint parameter or the second constraint parameter; the feature layer is: the feature depth extracted by each convolutional layer; the layer feature is: the feature extracted by each convolutional layer.
[0100] The layer feature stitching submodule is used to stitch together layer features from various feature layers according to a preset stitching architecture to obtain the target feature. The preset stitching architecture is the DenseNet architecture; the target feature is a stitched feature composed of low-level features, mid-level features, and high-level features.
[0101] The working principle and beneficial effects of the above technical solution are as follows:
[0102] Based on the different task types of the target task, this application determines the first constraint parameter and the second constraint parameter of the target feature extractor, and determines the layer features of different feature levels extracted by the target feature extractor after being constrained by the first constraint parameter or the second constraint parameter; introduces a stitching architecture to obtain the target features stitched together from the layer features of each feature level, so as to obtain the original target features of the tissue cell image as much as possible, thereby improving the accuracy of subsequent predictions.
[0103] In one embodiment, the evaluation module includes:
[0104] The pathological image set acquisition submodule is used to acquire manually annotated pathological image sets; the pathological image set includes: multiple pathological images with manual cell segmentation and cell classification annotations;
[0105] The parsing submodule is used to parse each pathological image in the pathological image set to obtain the annotation result set; the annotation result set is: the collection of manual annotation results corresponding to the pathological images;
[0106] The first calculation submodule is used to obtain the prediction result set output by each pathological image after inputting it into the pre-selected model, and to calculate the first number of prediction results in the prediction result set; the prediction result set is: the set of prediction results corresponding to each pathological image obtained by inputting the pathological image into the pre-trained pre-selected model; the first number is: the number of prediction results in the prediction result set.
[0107] The second calculation submodule is used to determine the second number of prediction results in the prediction result set that meet the accuracy standard based on the labeled result set, the prediction result set, and the preset accuracy standard. The accuracy standard is IoU (Intersection over Union), which is a standard for measuring the accuracy of detecting corresponding objects in a specific dataset. The second number is the number of prediction results with an IoU greater than 0.5.
[0108] The third calculation submodule is used to calculate the cell classification accuracy based on the first number and the second number; the cell classification accuracy is: the second number divided by the first number;
[0109] The first accuracy determination submodule is used to determine whether the evaluation result is usable if the cell classification accuracy is greater than or equal to a preset first threshold; the first threshold is preset manually.
[0110] The second accuracy determination submodule is used to determine the evaluation result as unusable if the cell classification accuracy is less than a preset first threshold.
[0111] The working principle and beneficial effects of the above technical solution are as follows:
[0112] This application introduces a set of manually annotated pathological images. Based on the annotation result set, prediction result set, and preset accuracy standard of each pathological image in the set, the cell classification accuracy is determined, which improves the suitability of determining the cell classification accuracy. The introduction of a first threshold to determine the usability of the pre-selected model is more appropriate.
[0113] In one embodiment, the retraining module includes:
[0114] The splitting submodule is used to split the training record set based on the different first labelers of the first training record in the training record set, to obtain multiple training record subsets; the first labeler is: the person who incorporates the training records in the training record set into the training data of the pre-selected model; the training record subset is: the set of the first training records labeled by each first labeler;
[0115] The traversal submodule is used to sequentially traverse the first training record in the training record subset.
[0116] The association value acquisition submodule is used to obtain the association value between the first record type and the target record type of the first training record during each traversal; the first record type is: the type of the first training record; the target record type is, for example: records of manual cell segmentation and cell classification;
[0117] The second marker determination submodule is used to determine the first marker corresponding to the subset of training records being traversed, and to use it as the second marker;
[0118] The first annotation experience value acquisition submodule is used to acquire the first annotation experience value of the second annotator corresponding to the first training record of the first record type; the larger the first annotation experience value, the more likely the corresponding first training record is to be used as the target training sample.
[0119] The target training sample determination submodule is used to determine the target training sample based on the association value and the first annotation empirical value; the target training sample is the training sample used for retraining.
[0120] The sample input submodule is used to input the target training samples into a preset neural network model to obtain a cell segmentation and classification model.
[0121] The working principle and beneficial effects of the above technical solution are as follows:
[0122] When it is found that the quality of the pre-trained pre-selected model is not high, it is necessary to retrain the corresponding pre-selected model. Generally, training samples are obtained again. However, not all of the previous training samples are without training value. Therefore, a solution is urgently needed.
[0123] This application determines the target training samples for retraining based on the empirical value of the first labeler corresponding to the first record type and the correlation value between the first record type and the target record type, thereby improving the rationality of obtaining the target training samples.
[0124] In one embodiment, the associated value retrieval submodule includes:
[0125] The analysis target determination unit is used to determine the first analysis target and the second analysis target for correlation analysis based on a preset analysis target acquisition template. The analysis target acquisition template is: constraining the acquisition behavior to only acquire the analysis target; the first analysis target is, for example: a first record type; the second analysis target is: a target record type.
[0126] The analysis item acquisition unit is used to acquire a first analysis item of a first analysis target and a second analysis item of a second analysis target; the first analysis item is, for example, the type semantics corresponding to a first record type; the second analysis item is, for example, the type semantics corresponding to a target record type; the type semantics are extracted based on semantic analysis technology, which can be achieved.
[0127] The cross-analysis pairing item determination unit is used to determine multiple first cross-analysis pairing items based on the first analysis item and the second analysis item; the first cross-analysis pairing item is a pairing combination item in which the first analysis item and the second analysis item are randomly paired;
[0128] The pairing value acquisition unit is used to acquire the pairing value of the first cross-analysis pairing item based on the first element information of the first analysis item and the second element information of the second analysis item corresponding to the first cross-analysis pairing item. The first element information is the data node information of the first analysis item; the second element information is the data node information of the second analysis item; the pairing value is, for example, 90. The higher the pairing value, the more likely the corresponding first cross-analysis pairing items are to be related to each other. When acquiring the pairing value, the descriptive features between the first element information and the second element information are extracted, and the pairing value corresponding to the descriptive features is determined according to the manually preset descriptive feature-pairing value library.
[0129] The cross-analysis pairing selection unit is used to select the corresponding first cross-analysis pairing as the second cross-analysis pairing if the pairing value is greater than or equal to a preset third threshold; the third threshold is preset manually.
[0130] The cross-analysis result acquisition unit is used to acquire the cross-analysis results of the second cross-analysis pair; the cross-analysis result is the difference value between the first analysis item and the second analysis item corresponding to the second cross-analysis pair;
[0131] The correlation value determination unit is used to determine the correlation value based on the cross-analysis results. The smaller the difference in the cross-analysis results, the larger the correlation value.
[0132] The working principle and beneficial effects of the above technical solution are as follows:
[0133] This application introduces an analysis target acquisition template to determine the first analysis target and the second analysis target, thereby improving the accuracy of analysis target determination; based on the pairing value of the first cross-analysis pairing item composed of the first analysis item of the first analysis target and the second analysis item of the second analysis target, the second cross-analysis pairing item is determined, thereby improving pairing efficiency; based on the cross-analysis results of the second cross-analysis pairing item, the correlation value is determined, thereby improving the suitability of correlation value acquisition.
[0134] In one embodiment, the first annotation experience value acquisition submodule includes:
[0135] The historical annotation record acquisition unit is used to acquire at least one second record type of the historical annotation records of the second marker; the second record type is: the record category of the historical annotation records of the second marker;
[0136] The record count calculation unit is used to calculate the number of historical annotation records for each second record type; the record count is the number of historical annotation records of the same second record type.
[0137] The experience level determination unit is used to query a preset record number-experience level comparison database to determine the experience level; the record number-experience level comparison database is a database that stores the correspondence between multiple record numbers and experience levels;
[0138] The marking seniority acquisition unit is used to acquire the marking seniority of the second marker; the marking seniority is: the number of years the second marker has been engaged in data marking work;
[0139] The second annotation experience value determination unit is used to determine the second annotation experience value of the second marker corresponding to the historical annotation record of the second record type based on experience level and marking years;
[0140] The first annotation experience value determination unit is used to take the corresponding second annotation experience value as the first annotation experience value if the second record type is consistent with the first record type.
[0141] The working principle and beneficial effects of the above technical solution are as follows:
[0142] This application determines the experience level of the second marker for the second record type based on the number of historical annotation records corresponding to the second record type and an introduced record number-experience level comparison library, thereby improving the accuracy of experience level acquisition. By introducing marker seniority, an adjusted second annotation experience value is determined, further improving the accuracy of second annotation experience value acquisition.
[0143] In one embodiment, the second annotation experience value determination unit includes:
[0144] The first seniority record determination subunit is used to determine the first seniority records of multiple employees based on the acquired data-tagged industry employment information database; the data-tagged industry employment information database is a database of information on employees in multiple data-tagged industries; the first seniority record is a record that records the number of years an employee has worked in the industry;
[0145] The second seniority record determination subunit is used to analyze and obtain reliable second seniority records from the first seniority record according to a preset reliability analysis template. The reliability analysis template is: the constraint analysis behavior only performs reliability analysis, for example: the year of employment cannot be earlier than the year the industry appeared; the second seniority record is: a reliable first seniority record.
[0146] The correction factor acquisition sub-unit is used to calculate the average length of service of the corresponding record for the second length of service record. The marked length of service is divided by the average length of service to obtain the correction factor. The average length of service is, for example, 2.5. The larger the correction factor, the more the corresponding experience level is adjusted upward.
[0147] The second annotation experience value calculation subunit is used to calculate the second annotation experience value of the historical annotation record corresponding to the second record type by the second marker, based on the experience level and the correction factor. During calculation, the experience level and the correction factor are multiplied accordingly to obtain the second annotation experience value.
[0148] The working principle and beneficial effects of the above technical solution are as follows:
[0149] This application introduces a credibility analysis template to screen the first seniority records in the data labeling industry employment information database, determine the credible second seniority records to calculate the average seniority, and improve the accuracy of obtaining the average seniority; the correction factor obtained by dividing the labeled seniority by the average value is used to correct the experience degree, which is more reasonable and further improves the accuracy of the second labeled experience value.
[0150] In one embodiment, the target training sample determination submodule includes:
[0151] The reference degree determination unit is used to query a preset association value-reference degree database to determine the reference degree of an association value; the association value-reference degree database is a database that stores the correspondence between multiple association values and reference degrees; the larger the association value, the higher the corresponding reference degree;
[0152] The decision value acquisition unit is used to fuse the reference degree and the first annotation empirical value to obtain the decision value. The data fusion is as follows: the first annotation empirical value and the reference degree are multiplied accordingly. The higher the decision value, the more suitable it is to be used as a target training sample.
[0153] The second training record acquisition unit is used to take the corresponding first training record as the second training record if the judgment value is greater than or equal to the preset fourth threshold; the fourth threshold is preset manually.
[0154] The third number calculation unit is used to calculate the third number of the second training record;
[0155] The model parameter acquisition module is used to obtain the model parameters of the pre-selected model; the model parameters are the input parameters of the pre-selected model.
[0156] The sample number calculation unit is used to estimate the target sample number based on the model parameters and the preset sample number estimation model. The sample number estimation model is to train the neural network model using multiple manual records of the number of samples trained based on the model parameters until the neural network model converges. The target sample number is the ideal number for training the cell segmentation and classification model.
[0157] The target training sample determination unit is used to determine the target training sample if the third number reaches the target number of samples.
[0158] The supplementary unit is used to supplement the target number of training records if the third number does not reach the target number of samples. The third training records, along with the corresponding second training records, are used as the target training samples. The third training records are: manually recorded cell segmentation and classification data obtained from large-scale data acquisition for supplementary training.
[0159] The working principle and beneficial effects of the above technical solution are as follows:
[0160] This application introduces an association value-reference value library. Based on the determined reference value and a first standard empirical value, a judgment value is jointly determined, improving the rationality and accuracy of the judgment value. A third number of training records is calculated after removing unqualified training samples. Model parameters and sample number are used to estimate the model, determining the ideal target sample number for model training. The third number is compared with the target sample number; if training data is insufficient, it is supplemented accordingly. This fully utilizes the already obtained second training records, supplementing only when the target training samples are insufficient. This measure improves model training quality while ensuring training efficiency.
[0161] This invention provides a method for cell segmentation and identification based on HE-stained pathological sections, such as... Figure 2 As shown, it includes:
[0162] Step 1: Evaluate the usability of the pre-trained pre-selected model and obtain the evaluation results;
[0163] Step 2: If the evaluation result is usable, use the corresponding pre-selected model as the cell segmentation and classification model; otherwise, obtain the training record set of the corresponding pre-selected model and retrain the cell segmentation and classification model based on the training record set.
[0164] Step 3: Set up the target feature extractor;
[0165] Step 4: Based on the target feature extractor, determine the target features in the HE-stained pathological slide images according to the acquired target task;
[0166] Step 5: Based on the cell segmentation and classification model, obtain the segmentation and classification results according to the target features corresponding to the target task.
[0167] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A cell segmentation and recognition system based on HE-stained pathological sections, characterized in that, The method comprises the following steps: An evaluation module is configured to evaluate the availability of a pre-trained preselected model and obtain an evaluation result; A retraining module is configured to use the preselected model as a cell segmentation and classification model if the evaluation result is available, or to obtain a training record set of the preselected model, retrain the cell segmentation and classification model based on the training record set, and set a target feature extractor; A feature determination module is configured to determine target features in a HE-stained pathological section image based on the target feature extractor and a target task; A segmentation and identification module is configured to obtain segmentation and classification results based on the cell segmentation and classification model and target features corresponding to the target task; The retraining module performs the following operations: Split the training record set based on different first markers of first training records in the training record set to obtain a plurality of training record subsets; Iterate through the first training records in the training record subsets one by one; Determine a first analysis target and a second analysis target for correlation analysis based on a preset analysis target acquisition template during each iteration; Determine a plurality of first cross-analysis pairing items based on a first analysis item of the first analysis target and a second analysis item of the second analysis target; Obtain a pairing value of the first cross-analysis pairing item based on first element information of the first analysis item and second element information of the second analysis item corresponding to the first cross-analysis pairing item; If the pairing value is greater than or equal to a preset third threshold value, use the first cross-analysis pairing item as a second cross-analysis pairing item; Determine an association value between a first record type of the first training record and a target record type based on a cross-analysis result of the second cross-analysis pairing item; Determine a first marker corresponding to the training record subset being iterated through as a second marker; Calculate the number of historical annotation records of each second record type of the second marker; Query a preset record number-experience degree reference library to determine an experience degree; Obtain the marker's length of service; Determine a plurality of first length of service records of the employees based on the obtained data marker industry employee information; Calculate the average value of the record length of the reliable second length of service records in the first length of service records, divide the marker's length of service by the average value to obtain a correction factor; Calculate a second annotation experience value of the second marker for the historical annotation records of the second record type based on the experience degree and the correction factor; If the second record type is consistent with the first record type, use the second annotation experience value as a first annotation experience value; Determine a target training sample based on the association value and the first annotation experience value; Input the target training sample into a preset neural network model to obtain the cell segmentation and classification model. The setting module comprises:
2. The HE staining pathological section based cell segmentation and recognition system of claim 1, wherein, A base model determination submodule is configured to determine a base model of a preset base feature extractor; A step length setting submodule is configured to set the step length of the convolution layer in the base model to 1; A reconnection submodule is configured to delete the maximum pooling layer in the base model and reconnect the unconnected hierarchical ports after deleting the maximum pooling layer to obtain the target feature extractor. The feature determination module comprises:
3. The HE staining pathological section based cell segmentation and recognition system of claim 1, wherein, The task type acquisition submodule is configured to acquire a task type of the target task, and the task type includes: nucleus segmentation and nucleus classification. The first parameter determination submodule is configured to determine a first constraint parameter of the target feature extractor when the task type is nucleus segmentation. The second parameter determination submodule is configured to determine a second constraint parameter of the target feature extractor when the task type is nucleus classification. The layer feature determination submodule is configured to determine layer features of different feature layers extracted by the target feature extractor after being constrained by the first constraint parameter or the second constraint parameter. The layer feature splicing submodule is configured to splice the layer features of the different feature layers according to a preset splicing architecture to obtain a target feature.
4. The HE staining pathological section based cell segmentation and recognition system of claim 1, wherein, The evaluation module includes: The pathological image set acquisition submodule is configured to acquire a set of manually annotated pathological images. The analysis submodule is configured to analyze each pathological image in the set of pathological images to obtain a set of annotation results. The first calculation submodule is configured to acquire a set of prediction results output by each pathological image when input into a preselected model, and calculate a first number of prediction results in the set of prediction results. The second calculation submodule is configured to determine a second number of prediction results in the set of prediction results that meet a preset accuracy standard according to the set of annotation results, the set of prediction results, and the preset accuracy standard. The third calculation submodule is configured to calculate a cell classification accuracy rate according to the first number and the second number. The first accuracy rate determination submodule is configured to determine that the evaluation result is available if the cell classification accuracy rate is greater than or equal to a preset first threshold. The second accuracy rate determination submodule is configured to determine that the evaluation result is unavailable if the cell classification accuracy rate is less than the preset first threshold.
5. The HE staining pathology slice based cell segmentation and recognition system of claim 1, wherein, The retraining module determines the target training sample according to the correlation value and the first annotation experience value, and includes: The reference degree determination unit is configured to query a preset correlation value-reference degree library to determine a reference degree of the correlation value. The determination value acquisition unit is configured to perform data fusion on the reference degree and the first annotation experience value to obtain a determination value. The second training record acquisition unit is configured to take the corresponding first training record as a second training record if the determination value is greater than or equal to a preset fourth threshold. The third number calculation unit is configured to calculate a third number of the second training records. The model parameter acquisition unit is configured to acquire model parameters of the preselected model. The sample number calculation unit is configured to determine a target sample number according to the model parameters and a preset sample number estimation model. The target training sample determination unit is configured to take the corresponding second training record as the target training sample if the third number reaches the target sample number. The supplement unit is configured to supplement a corresponding number of third training records according to a difference between the third number and the target sample number if the third number does not reach the target sample number, and take the third training records and the corresponding second training records together as the target training sample.
6. A method for cell segmentation and recognition based on HE-stained pathological sections, characterized in that, The method includes: Step 1: performing availability evaluation on a pre-trained preselected model to acquire an evaluation result; Step 2: taking the corresponding preselected model as a cell segmentation and classification model if the evaluation result is available, or otherwise, acquiring a set of training records of the corresponding preselected model and retraining the cell segmentation and classification model according to the set of training records; Step 3: setting a target feature extractor; Step 4: based on the target feature extractor, according to the obtained target task, determining the target feature in the HE-stained pathological section image; Step 5: based on the cell segmentation and classification model, according to the target feature corresponding to the target task, obtaining the segmentation result and the classification result; The step 2: if the evaluation result is available, the corresponding preselected model is used as the cell segmentation and classification model, otherwise, the training record set corresponding to the preselected model is obtained, and the cell segmentation and classification model is retrained according to the training record set, including: Based on the different first markers of the first training records in the training record set, the training record set is split to obtain a plurality of training record subsets; The first training records in the training record subsets are traversed in turn; Each time the first analysis target and the second analysis target are determined based on the preset analysis target acquisition template for correlation analysis; According to the first analysis item of the first analysis target and the second analysis item of the second analysis target, a plurality of first cross-analysis pairing items are determined; According to the first element information of the first analysis item and the second element information of the second analysis item corresponding to the first cross-analysis pairing item, the pairing value of the first cross-analysis pairing item is obtained; If the pairing value is greater than or equal to the preset third threshold value, the corresponding first cross-analysis pairing item is used as the second cross-analysis pairing item; According to the cross-analysis result of the second cross-analysis pairing item, the correlation value of the first record type and the target record type of the first training record is determined; The first marker corresponding to the training record subset being traversed is determined as the second marker; The record number of the historical annotation record of each second record type of the second marker is calculated; The experience degree is determined by querying the preset record number-experience degree comparison library; The marker seniority of the second marker is obtained; The first seniority record of a plurality of employees is determined based on the obtained data marker industry employee information database; The average value of the record seniority corresponding to the reliable second seniority record in the first seniority record is calculated, and the marker seniority and the average value are divided to obtain a correction factor; According to the experience degree and the correction factor, the second annotation experience value of the second marker corresponding to the second record type of the historical annotation record is calculated; If the second record type and the first record type are consistent, the corresponding second annotation experience value is used as the first annotation experience value; According to the correlation value and the first annotation experience value, the target training sample is determined; The target training sample is input into the preset neural network model to obtain the cell segmentation and classification model.
Citation Information
Patent Citations
Image segmentation and model training method, device and equipment
CN115631205A
Artificial intelligence assisted early gastric cancer screening system
CN111710394A