Cell grading system based on small-sample recognition

Through the CNN model combined with Triplet Loss training and GCN grading and combined with doctors' prior knowledge, the problem of insufficient feature extraction in the grading of small sample abnormal cells was solved, and higher grading accuracy and discernment ability were achieved.

CN114399666BActive Publication Date: 2025-08-05SHENZHEN DONGHUI PRECISION MECHANICAL & ELECTRICAL CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210058535.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-08-05
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

In the grading of small samples abnormal cells, representative features cannot be effectively extracted, resulting in low grading accuracy.

Method used

The CNN model is used to train the embedded feature similarity function with Triplet Loss formula, build graph structures through weighted models, use GCN model to perform cell grading, and fuse doctor prior knowledge to improve grading accuracy.

Benefits of technology

With a small number of training cell samples, the grading accuracy and discernment of abnormal cells are significantly improved, and the grading results are more accurate and objective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399666B_ABST
    Figure CN114399666B_ABST
Patent Text Reader

Abstract

The present application relates to a cell grading system based on small sample identification, which includes: a first acquisition module for acquiring a training set, the training set including known labeled cell samples; a training module for inputting the training set into a CNN model and training based on the Triplet Loss formula; a second acquisition module for acquiring unknown labeled cell samples; a first calculation module for calculating the similarity score between the unknown labeled cell sample and the known labeled cell sample using an embedded feature similarity function; a second calculation module for calculating the level type score of the unknown labeled cell sample as the known labeled cell sample based on the similarity score; a construction module for establishing a weighted model based on the level type score of the unknown labeled sample as the known labeled sample, and constructing a graph structure through the weighted model; a cell grading module for inputting the graph structure into a GCN model for cell grading. The present application has the effect of improving the accuracy of abnormal cell grading with a small number of training cell samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence and medicine, and in particular to a cell grading system based on small sample identification. Background Art

[0002] Cervical cancer is a malignant tumor that seriously endangers women's health and life. With the development of image processing and artificial intelligence technologies, automated pathology diagnosis technology has emerged. The key to this technology is the detection and grading of abnormal cervical cells in microscopic images.

[0003] The grading of abnormal cervical cells is based on the TBS diagnostic criteria, which further divides abnormal cervical cells into categories such as ASC-US, ASC-H, LSIL, HSIL, SCC keratinization, and SCC non-keratinization.

[0004] Among related technologies, abnormal cytology screening based on image processing is currently the most widely used screening method. Abnormal cell features are usually obtained from the color, shape, texture, etc. of the cell image to screen for abnormal cells.

[0005] However, for small samples with a small number of abnormal cells and high similarity, due to the limited number of abnormal cells themselves, the data of each type of sample is even more scarce when divided into several levels of abnormal cells.

[0006] Furthermore, abnormal cells in small samples have many common features, and the similarity between different types of abnormal cells is high. When extracting features from abnormal cells, the most representative features cannot be effectively extracted, thereby reducing the accuracy of abnormal cell grading. Summary of the Invention

[0007] In order to improve the accuracy of abnormal cell classification with a small number of training cell samples, the present application provides a cell classification system based on small sample identification.

[0008] This application provides a cell classification system based on small sample identification, which adopts the following technical solutions:

[0009] A cell classification system based on small sample identification includes a first acquisition module for acquiring a training set, wherein the training set includes cell samples with known labels;

[0010] A training module, configured to input the training set into a CNN model, perform training based on a Triplet Loss formula, and obtain an embedding feature similarity function;

[0011] The second acquisition module is used to obtain unknown label cell samples;

[0012] A first calculation module is used to calculate the similarity score between the unknown label cell sample and the known label cell sample using the embedded feature similarity function;

[0013] A second calculation module is used to calculate the level type score of the unknown labeled cell sample as a known labeled cell sample based on the similarity score;

[0014] A construction module, configured to establish a weighted model based on the grade type score of the unknown-labeled cell sample as a known-labeled cell sample, and construct a graph structure using the weighted model;

[0015] The cell grading module is used to input the graph structure into the GCN model to perform cell grading and determine the cell level of the unknown label cell sample.

[0016] By adopting the above technical solution, the training set is trained using the CNN model combined with the Triplet Loss formula, so that the distance between samples of the same type in the training set is closer and the distance between samples of different types is farther, so that the CNN model can extract more representative features of each cell sample in the training set, and then calculate the embedded feature similarity function; the embedded feature similarity function is used to calculate the similarity function of the unknown label cell sample to the known label cell sample, and the known label cell sample most similar to the unknown label cell sample can be obtained. After constructing the graph structure through the weighted model, it can be input into the GCN model to confirm the cell level of the unknown label cell sample, which can effectively extract the most representative features and has a significant effect on classifying abnormal cells with high similarity, thereby improving the accuracy of abnormal cell classification with a small number of training cell samples.

[0017] Optionally, the Triplet Loss formula is:

[0018]

[0019] in, is a sample randomly selected from the training set, For Similar cell samples, For Different types of cell samples, f(x) is the characteristic expression of element x, α is and The distance between and The minimum distance between them.

[0020] By adopting the above technical solution, Triplet Loss is used to train samples with less differences, which is more suitable for abnormal cell samples. It makes the distance between features of the same type in the training set closer and the distance between features of different types farther, so that more representative features can be extracted, thereby improving the recognition, accuracy and reliability of model training.

[0021] Optionally, the similarity score is calculated as follows:

[0022]

[0023] Among them, x i and y i Represents the value of the i-th dimension of the two normalized eigenvectors.

[0024] Optionally, the second calculation module includes: a determination submodule, configured to determine a maximum similarity score from the similarity scores of the unknown labeled cell sample and the known labeled cell sample;

[0025] A mapping submodule is used to map the maximum similarity score into a grade type score of the unknown label cell sample to a known label cell sample based on the abnormal cell diagnosis index.

[0026] Optionally, the construction module includes: establishing a submodule for establishing a weighted model using the level type score of the unknown label sample as the known label sample and the embedded feature similarity function combined with the doctor's prior knowledge to obtain the diagnostic indicator weight;

[0027] A construction submodule is used to calculate the distance between the known labeled cell sample and the unknown labeled cell sample using the diagnostic indicator weight to construct the graph structure.

[0028] By adopting the above technical solution, integrating TBS (The Bethesda System, classification and reporting rules of vaginal cytology), combining the doctor's prior knowledge, and constructing a graph structure, the grading results can be made more accurate and objective.

[0029] Optionally, the abnormal cell diagnostic indicators include: nuclear-cytoplasmic ratio, nuclear division rate, nuclear polarity, nuclear eccentricity, nuclear atypia, cell circular fit, nuclear area coefficient, vacuoles to cytoplasm area ratio, nucleolus to nucleus area ratio, nucleolus to nucleus area ratio, keratinization degree, nuclear groove concave area, nuclear staining depth, nuclear staining uniformity, cytoplasm richness, cell outline clarity, glandular cell disorder, cell cluster crowding, and cell cluster size distribution. At least one of the following.

[0030] Optionally, the calculation formula for establishing the weighted model is:

[0031]

[0032] Among them, S j is the level type score, θ is the weighting coefficient of the level type score in the weight calculation, is the feature embedding function, and σ is the sample length parameter generated by the graph structure.

[0033] Optionally, the cell classification module includes: an input submodule for inputting the graph structure into a GCN model, and using an aggregation function formula to make the unknown label cell sample features in the graph structure obtain the adjacent known label cell sample features to obtain aggregate features;

[0034] A determination submodule is configured to determine the cell level of the unknown-label cell sample based on the aggregated features.

[0035] Optionally, the aggregation function formula is:

[0036]

[0037] in, is the k-th layer feature of the unknown label sample v, mean is the average function, w is the diagnostic indicator weight, and σ is the nonlinear function.

[0038] 1. By training the training set using the CNN model combined with the Triplet Loss formula, the distance between samples of the same type in the training set is made closer, and the distance between samples of different types is made farther, so that the CNN model can extract more representative features of each cell sample in the training set, and then calculate the embedded feature similarity function; the embedded feature similarity function is used to calculate the similarity function of the unknown label cell sample to the known label cell sample, and the known label cell sample most similar to the unknown label cell sample can be obtained. After constructing the graph structure through the weighted model, it can be input into the GCN model to confirm the cell level of the unknown label cell sample, which can effectively extract the most representative features and has a significant effect on classifying abnormal cells with high similarity, thereby improving the accuracy of abnormal cell classification with a small number of training cell samples;

[0039] 2. The Triplet Loss function is used to train samples with less variability and is more suitable for abnormal cell samples. It makes the distances between features of the same type in the training set closer and the distances between features of different types farther apart, allowing the extraction of more representative features, thereby improving the discrimination, accuracy, and reliability of model training.

[0040] 3. By incorporating TBS (The Bethesda System, classification and reporting guidelines for vaginal cytology) and combining it with doctors' prior knowledge, a graph structure is constructed, making the grading results more accurate and objective. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a structural block diagram of a cell classification system based on small sample identification in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The present application is further described in detail below with reference to the accompanying drawings.

[0043] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document, unless otherwise specified, generally indicates an "or" relationship between the related objects.

[0045] The embodiment of the present application provides a cell grading system 100 based on small sample identification, which can be stored and / or executed by electronic devices, including but not limited to mobile terminal devices, PCs, workstations, and dedicated medical devices for cancer-related analysis.

[0046] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings. Figure 1 As shown, the system includes:

[0047] A first acquisition module 101 is used to acquire a training set;

[0048] In this embodiment, the training set consists of a small number of abnormal cervical cells with known labels. The abnormal cervical cells are taken from people of different ages and different conditions. The range of abnormal cervical cells includes all lesion levels in TBS (The Bethesda System, classification and reporting details of vaginal cytology), which are divided into categories such as ASC-US, ASC-H, LSIL, HSIL, SCC keratinization, and SCC non-keratinization.

[0049] After obtaining the training set, it needs to be preprocessed to make the cell image features of the training set more prominent.

[0050] The specific method for preprocessing the training set is to horizontally flip, rotate, shift, scale, align, and crop all cell feature images in the training set to obtain normalized images, so that the key points of the cell images in the training set are more prominent in each image. It should be noted that the training set can be preprocessed using image grayscale, histogram equalization, or other preprocessing methods, which are not specifically limited in this embodiment.

[0051] The preprocessed training set was divided into triplets, where the triplets were standard cell samples, similar cell samples, and non-similar cell samples. Using different features to train the training set improved the accuracy of feature extraction.

[0052] A training module 102 is used to input the training set into the CNN model and perform training based on the Triplet Loss formula to obtain an embedding feature similarity function;

[0053] The training process of the CNN model for a small number of cell samples with known labels is as follows:

[0054] Load the CNN model and input the training set into the CNN model, and use the parameters of the CNN model as the initial parameters to extract the image features of the preprocessed training set;

[0055] Based on Triplet Loss loss function learning, the CNN model parameters are continuously optimized through the loss function to obtain the best embedding feature similarity function.

[0056] Optionally, the Triplet Loss formula is:

[0057]

[0058] in, is a sample randomly selected from the training set, For Similar cell samples, For Different types of cell samples, f(x) is the characteristic expression of element x, α is and The distance between and The minimum distance between them.

[0059] By using the Triplet Loss function to train a CNN model for a small number of cell samples with known labels and small differences, the distances between features of the same type in the training set are brought closer, while the distances between features of different types are made farther, allowing more representative features to be extracted, thereby improving the recognition, accuracy, and reliability of model training.

[0060] After the CNN model is trained by the first acquisition module 101 and the training module 102, labels need to be constructed for cell samples with known labels and cell samples with unknown labels to be predicted. Therefore, the cell classification system 100 based on small sample recognition also includes:

[0061] The second acquisition module 103 is used to acquire an unknown label cell sample;

[0062] In this embodiment, the unknown label cell sample is preprocessed, and the preprocessing method is the same as the preprocessing method in the first acquisition module 101, which will not be described in detail here.

[0063] A first calculation module 104 is configured to calculate a similarity score between an unknown labeled cell sample and a known labeled cell sample using an embedded feature similarity function;

[0064] A second calculation module 105 is configured to calculate a grade type score of the unknown-label cell sample as a known-label cell sample based on the similarity score;

[0065] Optionally, the similarity score is calculated as:

[0066]

[0067] Among them, x i and y i Represents the value of the i-th dimension of the two normalized eigenvectors.

[0068] In this embodiment, the first calculation module 104 combines the unknown labeled cell sample with the embedded feature similarity function to obtain a similarity score between the unknown labeled cell sample and the known labeled cell sample, and then the second calculation module 105 further determines the features of the unknown labeled cell sample.

[0069] Specifically, the second calculation module 105 includes:

[0070] The determination submodule is used to determine the maximum similarity score from the similarity scores of the unknown label cell sample and the known label cell sample, and select the most similar features of the unknown label sample and the known label cell sample so that the features of the known label cell sample referenced by it are the most representative.

[0071] It should be noted that the maximum similarity score may be determined using a maximum function algorithm, or other statistical algorithms may be used to determine the maximum similarity, which is not specifically limited in this embodiment.

[0072] A mapping submodule, for mapping the maximum similarity score into a level type score of the unknown label cell sample to the known label cell sample based on the abnormal cell diagnosis index;

[0073] Among them, the diagnostic indicators of abnormal cells include: nuclear-cytoplasmic ratio, nuclear division rate, nuclear polarity, nuclear eccentricity, nuclear atypia, cell circular fit, nuclear area coefficient, vacuoles to cytoplasm area ratio, nucleolus to nucleus area ratio, nucleolus to nucleus area ratio, keratinization degree, nuclear groove concave area, nuclear staining depth, nuclear staining uniformity, cytoplasm richness, cell outline clarity, glandular cell disorder, cell cluster crowding, and cell cluster size distribution.

[0074] It should be noted that this embodiment can use cell morphology priors and basic mathematical formulas to define diagnostic indicators and formulate the diagnostic indicators. The diagnostic indicators can also be determined based on the doctor's prior knowledge, without specific limitation.

[0075] The following is a detailed description of the diagnostic indicators through the diagnostic indicator formula. The specific formula for the diagnostic indicator of abnormal cells is as follows:

[0076] (1) The specific formula for the nucleus-cytoplasm ratio is:

[0077]

[0078] Among them, R is the nucleus-cytoplasm ratio, A C is the cell area, A n is the cell nucleus area.

[0079] (2) The formula for calculating the nuclear division rate is:

[0080]

[0081] Among them, K is the nuclear division rate, x p 、x pp They represent the first and second derivatives of the x component of the curve, y p 、y pp They represent the first and second derivatives of the y component of the curve respectively.

[0082] (3) The specific formula for nuclear polarity is:

[0083]

[0084] Among them, P n is the nuclear polarity value, N is the total number of cells in the sample, θ i is the eccentricity angle of the i-th cell, is the average eccentricity angle, x1 is the horizontal coordinate of the cell center, y1 is the vertical coordinate of the cell center, x2 is the horizontal coordinate of the cell nucleus center, and y2 is the vertical coordinate of the cell nucleus center.

[0085] (4) The specific formula for nuclear eccentricity is:

[0086]

[0087] Where d is the nuclear eccentricity, (x1, y1) is the cell center, and (x2, y2) is the cell nuclear center.

[0088] (5) The specific formula for nuclear atypia is:

[0089]

[0090] Among them, C is the degree of nuclear atypia, D is the short axis of the rectangle surrounding the nucleus, L is the long axis of the rectangle surrounding the nucleus, and A is the n is the cell nucleus area.

[0091] (6) The specific formula for the cell circular fitting degree is:

[0092]

[0093] Among them, A n is the area of the cell nucleus, d1 and d2 are the long axis and short axis of the cell nucleus respectively, P is the perimeter of the cell nucleus, (x i ,y i )、(x i+1 ,y i+1 ) are two adjacent points on the boundary of the cell nucleus.

[0094] (7) The specific formula for the cell nucleus area coefficient is:

[0095]

[0096] Among them, A index is the cell nucleus area coefficient, A n is the cell nucleus area, A m The area of the nuclei of cells in the same layer is the mean.

[0097] (8) The specific formula for the ratio of vacuolar to cytoplasmic area is:

[0098]

[0099] Among them, K vacuoles is the ratio of vacuolar to cytoplasmic area, HSV vacuoles For whiteness, A vacuoles is the cavitation area, A cy is the cytoplasmic area.

[0100] (9) The specific formula for the ratio of nucleolus to nucleus area is:

[0101]

[0102] Among them, K nucleolus is the ratio of nucleolus to nucleus area, HSV nucleolus A is the degree of blackness,nucleolus is the nucleolus area output by the model, A n is the cell nucleus area.

[0103] (10) The calculation formula of angularity is:

[0104]

[0105] Among them, K orange is the keratinization degree, HSV orange Orange degree, A orange is the angular area, A cy is the cytoplasmic area.

[0106] (11) The specific formula for the concave area of the nuclear groove is:

[0107]

[0108] Among them, A groove is the concave area of the nuclear groove, n is the number of lines detected, x i Indicates the width of each line.

[0109] (12) The specific formula for nuclear staining intensity is:

[0110]

[0111] in, is the depth of nuclear staining, f(x,y) is the grayscale value of the nuclear image (x,y), l is the length of the nuclear image, and d is the width of the nuclear image.

[0112] (13) The specific formula for nuclear staining uniformity is:

[0113]

[0114] Where S is the uniformity of nuclear staining, f(x,y) is the grayscale value of the nuclear image at position (x,y), l is the length of the nuclear image, and d is the width of the nuclear image.

[0115] (14) The specific formula for cytoplasm richness is:

[0116]

[0117] Among them, T cy is the cytoplasm richness, S(x,y) is the gradient value of the Sobel convolution kernel at the point (x,y) in the cytoplasm of the image, and n is the number of cytoplasm pixels.

[0118] (15) The specific formula for cell outline clarity is:

[0119] F=∑x ∑ y G 2 (x,y);

[0120] Where F represents the cell contour clarity, and G(x,y) is the gradient value after the Laplace operator is convolved with the image at the (x,y) point in the cytoplasm nuclear contour.

[0121] (16) The specific formula for glandular cell disorder is:

[0122]

[0123] Among them, H gc is the disorder degree of glandular cells, p(x i ) is the probability of glandular cells in the unit rectangle, N g is the number of unit rectangles.

[0124] (17) The specific formula for cell cluster crowding is:

[0125]

[0126] Among them, O cm is the cell cluster crowding degree, A overlap is the overlapping cell area, A ncm is the total area of cell nuclei in the cell cluster, N ncm is the number of overlapping cells.

[0127] (18) The specific formula for the cell cluster size distribution is:

[0128]

[0129] Among them, B cm is the size distribution of cell clusters, A i is the area of the ith cell cluster, is the average area of cell clusters, N cm is the number of cell clusters.

[0130] After the first calculation module 104 and the second calculation module 105 determine that the unknown labeled cell sample has the most similar characteristics to the known labeled cell sample, the unknown labeled cell sample needs to be graded. Therefore, the cell grading system 100 based on small sample identification further includes:

[0131] A construction module 106 is configured to establish a weighted model based on the level type scores of the unknown-labeled cell samples as known-labeled cell samples, and to construct a graph structure using the weighted model;

[0132] Specifically, the building block 106 includes:

[0133] Establish a submodule to use the level type score of the unknown label sample to the known label sample and the embedded feature similarity function combined with the doctor's prior knowledge to establish a weighted model to obtain the diagnostic indicator weight;

[0134] A submodule is constructed to calculate the distance between known labeled cell samples and unknown labeled cell samples through the diagnostic indicator weights and construct a graph structure.

[0135] Optionally, the calculation formula for establishing a weighted model is:

[0136]

[0137] Among them, B cm is the size distribution of cell clusters, A i is the area of the ith cell cluster, is the average area of cell clusters, N cm is the number of cell clusters.

[0138] It is worth noting that the level of each abnormal cell needs to refer to different diagnostic indicators. The weights of different diagnostic indicators are also different. They need to be calibrated by experienced doctors and are a reflection of the doctor's prior knowledge.

[0139] The cell classification module 107 is used to input the graph structure into the GCN model for cell classification and determine the cell level of the unknown label cell sample.

[0140] Specifically, the cell classification module 107 includes:

[0141] The input submodule is used to input the graph structure into the GCN model and use the aggregation function formula to make the unknown label cell sample features in the graph structure obtain the adjacent known label cell sample features to obtain the aggregated features;

[0142] The determination submodule is used to determine the cell level of unknown label cell samples based on aggregated features.

[0143] Optionally, the aggregation function formula is:

[0144]

[0145] in, is the k-th layer feature of the unknown label sample v, mean is the average function, w is the diagnostic indicator weight, and σ is the nonlinear function.

[0146] In this embodiment, construction module 106 integrates TBS (The Bethesda System, a classification and reporting system for vaginal cytology) and physicians' prior knowledge to construct a graph structure. After constructing the graph structure, input cell grading module 107 uses an aggregate function formula to map the features of known labeled cell samples adjacent to the features of unknown labeled cell samples to features of unknown abnormal labeled cell samples. This allows each unknown labeled cell sample to obtain an aggregated feature, which is then used for grading, resulting in more accurate and objective grading results.

[0147] The cell grading system 100 based on small sample identification can effectively extract the most representative features of abnormal cells in a small sample. The first acquisition module 101 acquires a training set and divides the training set into standard cell samples, similar cell samples, and non-similar cell samples. The CNN model of the training module 102 is trained in combination with the Triplet Loss formula to make the distance between similar samples in the training set closer and the distance between non-similar samples farther. As a result, the CNN model extracts more representative features of each sample in the training set and calculates an embedded feature similarity function. After the second acquisition module 103 acquires an unknown label cell sample, the first calculation module 104 and the second calculation module 105 use the embedded feature similarity function to calculate a similarity function between the unknown label sample and the known label sample. The known label cell sample that is most similar to the unknown label cell sample can be obtained. Then, the weighted model of the construction module 106 is used to construct a graph structure. The cell grading module 107 can input the graph structure into the GCN model to confirm the cell level of the unknown label cell sample. The system has a significant effect on classifying abnormal cells with high similarity, and improves the accuracy of abnormal cell grading with a small number of training cell samples.

[0148] Various objects such as various messages / information / equipment / network elements / systems / devices / actions / operations / processes / concepts that may appear in this application are named. It can be understood that these specific names do not constitute a limitation on the relevant objects. The names assigned may change with factors such as scenarios, contexts or usage habits. The understanding of the technical meaning of the technical terms in this application should be mainly determined from the functions and technical effects embodied / executed in the technical solutions.

[0149] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0150] This specific embodiment is merely an explanation of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed. However, as long as such modifications are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. A cell classification system based on small sample identification, characterized in that: include: A first acquisition module is used to acquire a training set, wherein the training set includes cell samples with known labels; A training module, configured to input the training set into a CNN model, perform training based on a Triplet Loss formula, and obtain an embedding feature similarity function; The second acquisition module is used to obtain unknown label cell samples; A first calculation module is used to calculate the similarity score between the unknown label cell sample and the known label cell sample using the embedded feature similarity function; A second calculation module is used to calculate the level type score of the unknown labeled cell sample as a known labeled cell sample based on the similarity score; A construction module, configured to establish a weighted model based on the grade type score of the unknown-labeled cell sample as a known-labeled cell sample, and construct a graph structure using the weighted model; A cell grading module is used to input the graph structure into the GCN model for cell grading and determine the cell level of the unknown label cell sample; The second calculation module includes: a determination submodule, configured to determine a maximum similarity score from the similarity scores of the unknown labeled cell sample and the known labeled cell sample; A mapping submodule is used to map the maximum similarity score into a level type score of the unknown label cell sample to the known label cell sample based on the abnormal cell diagnostic index, wherein the level type score represents the similarity between the unknown label cell sample and the known label cell sample at the cellular level.

2. A cell classification system based on small sample identification according to claim 1, characterized in that: The Triplet Loss formula is: in, is a sample randomly selected from the training set, For Similar cell samples, For Different types of cell samples, f(x) is the characteristic expression of element x, α is and The distance between and The minimum distance between them.

3. A cell classification system based on small sample identification according to claim 1 or 2, characterized in that: The calculation formula of the similarity score is: Among them, x i and y i Represents the value of the i-th dimension of the two normalized eigenvectors.

4. A cell classification system based on small sample identification according to claim 1, characterized in that: The building blocks include: Establishing a submodule for establishing a weighted model using the level type score of the unknown label sample as the known label sample and the embedded feature similarity function combined with the doctor's prior knowledge to obtain the diagnostic indicator weight; A construction submodule is used to calculate the distance between the known labeled cell sample and the unknown labeled cell sample using the diagnostic indicator weight to construct the graph structure.

5. A cell classification system based on small sample identification according to claim 1 or 4, characterized in that: The abnormal cell diagnostic indicators include: nuclear-cytoplasmic ratio, nuclear division rate, nuclear polarity, nuclear eccentricity, nuclear atypia, cell circular fit, nuclear area coefficient, vacuoles to cytoplasm area ratio, nucleolus to nucleus area ratio, nucleolus to nucleus area ratio, keratinization degree, nuclear groove concave area, nuclear staining depth, nuclear staining uniformity, cytoplasm richness, cell outline clarity, glandular cell disorder, cell cluster crowding, and cell cluster size distribution.

6. A cell classification system based on small sample identification according to claim 4, characterized in that: The calculation formula for establishing the weighted model is: Among them, S j is the level type score, θ is the weighting coefficient of the level type score in the weight calculation, is the feature embedding function, and σ is the sample length parameter generated by the graph structure.

7. A cell classification system based on small sample identification according to claim 1 or 6, characterized in that: The cell classification module includes: An input submodule, configured to input the graph structure into the GCN model, and use an aggregation function formula to make the unknown label cell sample features in the graph structure obtain the adjacent known label cell sample features to obtain aggregated features; A determination submodule is configured to determine the cell level of the unknown-label cell sample based on the aggregated features.

8. A cell classification system based on small sample identification according to claim 7, characterized in that: The aggregation function formula is: in, is the k-th layer feature of the unknown label cell sample v, mean is the average function, w is the diagnostic index weight, and σ is the nonlinear function.

Citation Information

Patent Citations

  • Cell classification method based on cell fluorescence image

    CN116310531A