Incremental learning-based cell morphology test system iteration method and device
Through an incremental learning method, using the out-of-distributed sample detection module and attribute learning memory library, efficient training and learning of the intelligent blood cell morphology test system is achieved, solving the problem of inefficient training in the face of new cell categories and attributes.
Patent Information
- Application Number
- CN202510054215.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing intelligent blood cell morphology test system is weak in the face of the ever-increasing cell categories and cell attribute characteristics analysis and learning, which leads to high training time cost and low training efficiency.
Using the iterative method of the cell morphology test system based on incremental learning, the image data of the target cells is obtained through the preset out-of-distributed sample detection module and attribute learning memory library, and the data set is constructed, and incremental learning is updated until all new cell data are marked as learned, and the final round of test system is output.
It realizes that the new cell data is continuously learned and memorized without forgetting old cell data, reduces the consumption of time and computing power, improves training efficiency, and solves the shortcomings of existing systems in training and learning capabilities.
Smart Images

Figure CN120014634A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image analysis technology, and in particular, to an iterative method and device for a cell morphology inspection system based on incremental learning. Background Art
[0002] In the field of cytological morphological analysis, due to the large number of cell categories and the emergence of various different morphologies, the existing intelligent blood cell morphology inspection system faces the problems of constantly adding new cell categories and weak ability to analyze and learn cell attribute characteristics during application. Therefore, the blood cell morphology inspection system needs to continuously and repeatedly learn a large amount of original cell data during training and learning, which leads to the high time cost and low training efficiency of training the blood cell morphology inspection system. Summary of the invention
[0003] The present application provides an iterative method and device for a cell morphology inspection system based on incremental learning to solve one or more technical problems existing in the prior art and at least provide a beneficial choice or create conditions.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0005] According to one aspect of an embodiment of the present application, an iterative method for a cell morphology inspection system based on incremental learning is provided, which is applied to a cell morphology inspection system, wherein the cell morphology inspection system includes a preset out-of-distribution sample detection module and an attribute learning memory library, wherein the attribute learning memory library stores a plurality of learned and memorized cell data, and the method includes:
[0006] S1, acquiring image data of target cells, and constructing a first data set according to the image data;
[0007] S2, inputting the first data set into the out-of-distribution sample detection module to obtain a new category cell data set in the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to represent cell data that has not been learned and memorized;
[0008] S3, performing iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module, until each of the new cell data is marked as learned and memorized cell data, and outputting the final round of target cell morphology inspection system;
[0009] The incremental learning of the current round of the cell morphology inspection system includes the following steps:
[0010] Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell;
[0011] The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
[0012] In one embodiment of the present application, based on the aforementioned scheme, the incremental learning includes attribute incremental learning and category incremental learning, wherein the attribute incremental learning is used to characterize the incremental learning of attribute characteristics of each cell attribute of the target new cell data, and the category incremental learning is used to characterize the incremental learning of category characteristics of the target new cell data.
[0013] In one embodiment of the present application, based on the above scheme, the cell morphology inspection system outputted in the previous round is incrementally learned according to each cell attribute of the target new cell data, including:
[0014] Obtain the previous attribute learning memory library of the cell morphology inspection system output in the previous round;
[0015] Determine, according to each cell attribute of the target new cell data, a plurality of similarity weight data between each cell attribute that has been learned and memorized in the previous attribute learning memory library, to obtain a similarity weight data set;
[0016] For each cell attribute of the target new cell data, a target cell attribute corresponding to the cell attribute is obtained in the previous attribute learning memory library, and attribute feature migration and fusion are performed on each of the cell attributes and the cell attributes according to the similarity weight data set to obtain a cell attribute fusion vector, wherein the target cell attribute is a cell attribute learned and memorized in historical rounds;
[0017] Determine the attribute feature vector of the target new cell data according to each of the cell attribute fusion vectors;
[0018] A feature fusion vector is determined according to the attribute feature vector and the attribute feature vectors of each cell data in historical rounds, and the previous attribute learning memory library is updated according to the feature fusion vector to complete the attribute incremental learning of the cell morphology inspection system.
[0019] In one embodiment of the present application, based on the above scheme, the category incremental learning of the cell morphology inspection system is performed by the following steps:
[0020] Binding the feature fusion vector to the target new cell data to determine a mapping relationship between the feature fusion vector and the target new cell data;
[0021] The category features of the target new cell data are determined according to the mapping relationship, and the out-of-distribution sample detection module is updated according to the category features to complete the category incremental learning of the cell morphology inspection system.
[0022] In one embodiment of the present application, based on the above scheme, the method of determining multiple similarity weight data between each cell attribute of the target new cell data and each learned and memorized cell data in the previous attribute learning memory library according to each cell attribute of the target new cell data, and obtaining a similarity weight data set includes:
[0023] For each cell attribute of the target new cell data, calculating the Hamming distance between the cell attribute and the attribute feature of each cell data in the attribute learning memory library;
[0024] The similarity weight data corresponding to each of the cell attributes of the target new cell data are determined according to each of the Hamming distances, and the similarity weight data set is determined according to each of the similarity weight data.
[0025] In one embodiment of the present application, based on the above scheme, the new category cell dataset is obtained by the following steps:
[0026] Obtaining the category probability and distance score of each of the image data in the first data set through an out-of-distribution sample detection module;
[0027] For each image data, if the product value of the category probability and the distance score of the image data is greater than a preset threshold, then determining that the cell data corresponding to the image data is the new cell data;
[0028] Wherein, each of the image data corresponds to single cell data.
[0029] In one embodiment of the present application, based on the aforementioned scheme, the cell morphology inspection system is trained according to a preset class balance loss function, a preset cosine distillation loss function, a preset feature enhancement loss function and a preset attribute space separation loss function, and the class balance loss function, the cosine distillation loss function, the feature enhancement loss function and the attribute space separation loss function are used for attribute incremental learning.
[0030] In one embodiment of the present application, based on the aforementioned scheme, the cell morphology inspection system is also trained according to a preset cross entropy loss function and a preset distillation loss function, and the cross entropy loss function and the distillation loss function are used for category incremental learning.
[0031] According to one aspect of an embodiment of the present application, a cell morphology inspection system iteration device based on incremental learning is provided, which is applied to a cell morphology inspection system, wherein the cell morphology inspection system comprises a preset out-of-distribution sample detection module and an attribute learning memory library, wherein the attribute learning memory library stores a plurality of learned and memorized cell data, and the device comprises:
[0032] an acquisition unit, configured to acquire image data of target cells and construct a first data set according to the image data;
[0033] An input unit, used for inputting the first data set into the out-of-distribution sample detection module to obtain a new category cell data set in the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to represent cell data that has not been learned and memorized;
[0034] An incremental learning unit, used for performing iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module, until each of the new cell data is marked as learned and memorized cell data, and outputting a final round of target cell morphology inspection system;
[0035] The incremental learning of the current round of the cell morphology inspection system includes the following steps:
[0036] Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell;
[0037] The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
[0038] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in the above embodiment is implemented.
[0039] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions of the processors, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the methods described in the above embodiments.
[0040] Beneficial effects of the present application: The present application obtains a new category of cell data set in the first data set through a preset out-of-distribution sample detection module, that is, an out-of-distribution class detection method, that is, screens out new cell data that have not been learned and memorized, and performs iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data, wherein one round of incremental learning only needs to use one new cell data, at which time it can be determined that the new cell data used in the current round of incremental learning is the target new cell data, then the cell morphology inspection system output in the previous round can be incrementally learned according to each cell attribute of the target new cell data, until each cell attribute is incrementally learned, thereby obtaining a feature fusion vector of the target new cell data, so that the entire cell morphology inspection system can complete incremental learning through the feature fusion vector, that is, the attribute learning memory library and the out-of-distribution sample detection module can be updated, so that the next round of incremental learning can be performed on the cell morphology inspection system output in the current round. Incremental learning, while not forgetting the old cell data, continuously realizes incremental learning of new cell data, reduces catastrophic forgetting, and thereby reduces the consumption of time and computing power, thereby improving the training efficiency of the cell morphology inspection system.
[0041] A target attribute discriminator and a target category classifier are obtained, and the cell attributes and cell categories of each new cell are obtained. The cell attributes and cell categories obtained through learning can be used to iteratively evolve the cytological morphology inspection system, that is, each new cell data is marked as old cell data that has been learned and remembered by the cytological morphology inspection system, and stored in a local memory database. New cell data of new categories are continuously learned without forgetting old cell data of old categories, and the new cell data obtained through continuous learning and memory are stored in the attribute learning memory library and converted into old cell data, thereby completing the iterative update of the entire cell morphology inspection system and solving the problem of weak analysis and learning capabilities existing in the prior art.
[0042] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0044] Figure 1 It is a flow chart of an iterative method of a cell morphology inspection system based on incremental learning according to an embodiment of the present application;
[0045] Figure 2 The overall logic diagram of incremental learning of the cell morphology inspection system shown in the embodiment of the present application;
[0046] Figure 3 It is a logic diagram for calculating the product value of category probability and distance score according to an embodiment of the present application;
[0047] Figure 4 It is a flow chart of judging new and old category cell data according to an embodiment of the present application;
[0048] Figure 5 is an architecture diagram of an attribute learning memory library according to an embodiment of the present application;
[0049] Figure 6 Detailed schematic diagram of ILA incremental learning phase 2 and phase 3 according to an embodiment of the present application;
[0050] Figure 7 Detailed schematic diagram of ILA incremental learning phases one and two according to an embodiment of the present application;
[0051] Figure 8 A schematic diagram of target mapping parameter generation for ILA incremental learning phases one and two according to an embodiment of the present application;
[0052] Fig. 9 A simple flowchart of incremental learning according to an embodiment of the present application is shown;
[0053] Fig.10 A block diagram of an iterative cell morphology inspection system based on incremental learning according to an embodiment of the present application;
[0054] Fig.11 4 is a diagram showing the relationship between cell attributes and cell categories according to an embodiment of the present application. DETAILED DESCRIPTION
[0055] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0056] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.
[0057] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or micro-control node devices.
[0058] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0059] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0060] The implementation details of the technical solution of the embodiment of the present application are described in detail below:
[0061] According to one aspect of the present application, a cell morphology inspection system iteration method based on incremental learning is provided. Figure 1 The flowchart of the iterative method of the cell morphology inspection system based on incremental learning according to the embodiment of the present application is shown. The iterative method of the cell morphology inspection system based on incremental learning includes at least steps S1 to S3, which are described in detail as follows:
[0062] In step S1, image data of target cells are acquired, and a first data set is constructed according to the image data.
[0063] Specifically, during the inspection process, image data is obtained from a blood cell smear by using a camera device, and a first data set D1 is constructed based on the obtained image data. The newly acquired first data set D1 is input into the cell morphology inspection system, and the first data set D1 contains both old cell category data (old cell data is the cell data that has been learned and memorized in the attribute learning memory bank in the cell morphology inspection system output in the previous round), and new cell category data (new cell data), wherein one new cell data represents a new category of cells. It should be noted that the new cell data includes multiple cell attributes, that is, it has multiple attribute features. The attribute features corresponding to some cell attributes in the new cell data may have been learned in the incremental learning of the historical rounds, and some cell attributes have not been learned. Among them, if some cell attributes have been learned in the previous round, the learned target cell attributes are migrated and fused with the cell attributes of this round. If they have not been learned, a high-dimensional extractor is dynamically expanded to extract the attribute features of these cell attributes that have not been learned and memorized.
[0064] It should be distinguished that the new cell data represents a new category of cell data, which includes multiple cell attribute features, some of which may have been learned and some may not have been learned.
[0065] Therefore, the incremental learning of the present application is actually divided into two categories. The first category is attribute incremental learning, which performs incremental learning on cell attributes, and the second category is category incremental learning, which performs incremental learning on category features for new cell data. The new cell data can be represented by t, wherein the incremental learning of cell attributes can be specifically as follows: if the cell attribute has not been learned and memorized in the incremental learning of historical rounds, the cell attribute is used as a new cell attribute. Traverse each cell attribute in the new cell data to obtain one or more new cell attributes, then construct a corresponding high-dimensional extractor based on the set of these new cell attributes to perform feature extraction of the new cell attribute. If the cell attribute has been learned and memorized in the incremental learning of historical rounds, then find the target cell attribute that has been learned and memorized for the cell attribute in the historical rounds, and perform weighted fusion of these target cell attributes with the cell attributes of the current round to obtain a cell attribute fusion vector corresponding to the current round.
[0066] In step S2, the first data set is input into the out-of-distribution sample detection module to obtain a new category cell data set within the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to characterize cell data that has not been learned and memorized.
[0067] Specifically, the out-of-distribution sample detection module, that is, the model network that uses the preset out-of-distribution class detection method for detection, is an out-of-distribution class detection (OCD) based on attribute features (i.e., the cell attributes described in this application) and the KNN (K-NearestNeighbor) algorithm, hereinafter referred to as the OCD method, which can be Figure 2 As shown, the first data set D1 is processed by the OCD method, the old category cell data and the new category cell data are screened out, and a data set D2 containing all the new cell category data is obtained. The cell category of data set D2 has been labeled by the inspector during the morphological inspection process. The learning value of data set D2 is sorted by the active learning method, and the samples with the most learning value are screened out to obtain a number of new category data with annotations after classification, and they are added to the new category data set d. The new category data set d is the new category cell data set represented in this application. The number of samples in the new category cell data set is greater than the threshold T, where T can be set to an arbitrary value. Among them, since the out-of-distribution sample detection module is updating in iterative rounds, each time a new cell data is incrementally learned, the cell category corresponding to the new cell data will be memorized, and the judgment accuracy of the next round of new and old category cell data can be improved.
[0068] In one embodiment of the present application, the new category cell dataset is obtained by the following steps:
[0069] Obtaining the category probability and distance score of each of the image data in the first data set through an out-of-distribution sample detection module;
[0070] For each image data, if the product value of the category probability and the distance score of the image data is greater than a preset threshold, then determining that the cell data corresponding to the image data is the new cell data;
[0071] Wherein, each of the image data corresponds to single cell data.
[0072] Specifically, combined Figure 3 and Figure 4 As shown, the OCD method is described in detail:
[0073] Stage 1: Feature extraction and normalization
[0074] Use a convolutional neural network model that has been pre-trained on a dataset of x blood cells. For each sample in the training set, validation set, and test set, the model will output the predicted probability of x categories, extracting the feature vectors of different network outputs during the prediction process.
[0075] Feature extraction is divided into several steps:
[0076] Extract the output features of the last block of layer2 and layer3 in the model. These features represent the intermediate features learned after the input data has been processed by convolution at different depths. Extract the output of the global average pooling layer. The global average pooling layer can capture the global information of the entire image and summarize it into a single vector.
[0077] The feature vectors are normalized using the L2 norm, and the components of each feature vector are adjusted to the same scale range, so that the extracted feature vectors have the same scale and the accuracy of distance calculation is improved.
[0078] The OCD method stores the feature vectors of all samples (samples are image data described in this application) together with their corresponding sample indexes in a feature array. This helps to quickly retrieve features and identify the source of samples in subsequent steps. The feature array is divided into three, including feature vectors of the training set, as well as feature vectors of the validation set and the test set, in preparation for subsequent distance calculation and OOD (Out-of-Distribution Detection) detection.
[0079] Phase 2: Distance calculation and scoring
[0080] The Faiss library is used to build an index. For each test sample, the Euclidean distance between the test sample and all training set samples is calculated. The Euclidean distance is used to measure the similarity between two feature vectors. The smaller the distance, the more similar the two samples are.
[0081] In order to reduce the impact of extreme values on the distance score, after calculating the distance, all distances are arranged from small to large. Then, the middle k distances among these distances are selected (that is, the outliers at the front end and back end are excluded), and these distances are inverted and averaged to be the score of the test sample. The higher the score, the lower the similarity between the sample and the training set, and the more likely it is an OOD sample (representing sample data that has not been learned and remembered).
[0082] Phase 3: OOD Detection
[0083] Before performing OOD detection, a threshold is preset to determine whether a sample is OOD. The last layer of the ResNet34 model is modified to output x+1 categories, where the x+1th category is used as the OOD category.
[0084] The samples in the validation set are predicted to obtain the probability that each sample belongs to the OOD category. These probabilities represent the confidence of the model that the sample is OOD. At the same time, for each sample in the validation set, the OCD method also calculates its distance score. Then, the OOD category probability of each sample is multiplied by its distance score to obtain a product value. Finally, these product values are sorted, and the product values corresponding to the last 5% of samples are selected as the threshold, which is used to distinguish between ID samples (representing samples that have been learned and memorized) and OOD samples.
[0085] For each test sample, the OCD method calculates the category probability and distance score of its OOD category according to the processing method of the validation set, and multiplies the two. If the product value of the test sample is greater than the set threshold, the sample is determined to be an OOD sample (new cell data); otherwise, it is determined to be an ID sample (old cell data).
[0086] In step S3, the cell morphology inspection system is subjected to iterative rounds of incremental learning according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module, until each of the new cell data is marked as learned and memorized cell data, and the target cell morphology inspection system of the final round is output;
[0087] The incremental learning of the current round of the cell morphology inspection system includes the following steps:
[0088] Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell;
[0089] The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
[0090] The incremental learning includes attribute incremental learning and category incremental learning. The attribute incremental learning is used to represent the incremental learning of attribute characteristics of each cell attribute of the target new cell data, and the category incremental learning is used to represent the incremental learning of category characteristics of the target new cell data.
[0091] The incremental learning of the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data includes:
[0092] Obtain the previous attribute learning memory library of the cell morphology inspection system output in the previous round;
[0093] Determine, according to each cell attribute of the target new cell data, a plurality of similarity weight data between each cell attribute that has been learned and memorized in the previous attribute learning memory library, to obtain a similarity weight data set;
[0094] For each cell attribute of the target new cell data, a target cell attribute corresponding to the cell attribute is obtained in the previous attribute learning memory library, and attribute feature migration and fusion are performed on each of the cell attributes and the cell attributes according to the similarity weight data set to obtain a cell attribute fusion vector, wherein the target cell attribute is a cell attribute learned and memorized in historical rounds;
[0095] The feature fusion vector is determined according to each of the cell attribute fusion vectors, and the previous attribute learning memory library is updated according to the feature fusion vector to complete the attribute incremental learning of the cell morphology inspection system.
[0096] Specifically, the following combination Figure 5 , 6 , 7, 8 for detailed explanation:
[0097] First, this application adopts the ILA (Incremental learning based on cell Attributes) method to construct a progressive attribute-driven dynamic expansion network, designs an attention-weighted dynamic feature fusion network, and designs a cosine distillation loss to achieve efficient learning and classification of newly added category cells, while maintaining the model's discrimination performance for the attributes and categories of new and old category cells, reducing catastrophic forgetting, and reducing time and computing power consumption, thereby improving training efficiency.
[0098] The following is a detailed explanation of the ILA method:
[0099] like Figure 5 The initialization phase shown is: using the existing cell category data (i.e., the cell data that has been learned and memorized in the cell morphology inspection system output in the previous round) to initialize the low-dimensional attribute extractor L, the high-dimensional attribute extractor H1, the attribute discriminator F1 and the category classifier C1, and constructing / updating the attribute learning memory library.
[0100] The network contains a shared L, which is composed of the first two residual blocks of a pre-trained model, responsible for extracting basic features from the input data. It contains a high-dimensional attribute feature extractor H1, which is composed of the last two residual blocks of the pre-trained model. Therefore, the low-dimensional attribute extractor L and the high-dimensional attribute extractor H are different parts of the same model network, and the output of the low-dimensional attribute extractor L is the input of the high-dimensional attribute extractor. Before processing the new cell category task, the low-dimensional attribute extractor L and the high-dimensional attribute extractor H1 are trained based on the existing cell data to provide model support for subsequent incremental tasks.
[0101] The attribute discriminator F1 maps the output h1 of the high-dimensional attribute extractor H1 through a fully connected layer to obtain A1. The final output of the mapping network is an n-dimensional attribute feature layer that can be divided into Q attribute combinations, that is, an attribute feature vector. In the initialization stage, since the model has only one high-dimensional attribute extractor, the attribute discriminator F1 does not need to perform attention-weighted dynamic fusion on the attribute feature vector A1.
[0102] The cell category classifier C1, based on the attribute discriminator F1, finally adds a layer of k-class classification layer as the output of cell category classification. The training of the attribute classification network enables the network to classify cells directly by attribute features. It should be noted that after the attribute discriminator obtains the output (the output is the feature fusion vector described in this application), the output is used for training to achieve incremental learning at the attribute level (i.e., the attribute incremental learning described in this application). This is an independent training process, and the training of attribute incremental learning is carried out through the attribute incremental network. After the training is completed, the feature fusion vector output by the attribute incremental network is imported, and then the category discriminator is used for training to achieve category incremental learning, and category incremental learning is carried out through the category incremental network.
[0103] Specifically, Fig.11 As shown in Figure 1, each cell category corresponds to multiple cell attributes, and each image data contains a cell category label (i.e. Fig.11 At the same time, according to the knowledge of blood cell morphology, several attributes are defined for each type of cell (each new cell data in this application represents a type of cell), and each attribute contains several attribute features (each attribute represents each cell attribute described in this application). Some data can be seen in the figure, which shows some attribute features of 15 types of cells. In the table, except for the first row (cell category) and the first column (cell attribute), the content of each cell represents the semantic feature value of its corresponding attribute.
[0104] Build an attribute learning memory library to record the attribute features (i.e., cell attributes) corresponding to each learned cell category. The present invention uses the form of DataFrame, where each row represents a cell category and each column corresponds to an attribute feature. This structure allows for easy access, query, and manipulation of data. When inputting cell data, first read the CSV file containing the cell data attributes and load it into the Pandas DataFrame. Set the cell category as the index of the DataFrame for subsequent queries. Save the constructed attribute learning memory library as a new CSV file format for loading and use in subsequent operations.
[0105] Stage 1: Import the attribute features (attribute features, i.e., the cell attributes described in this application) of the new category cell t (new category cell data set d) into the attribute learning memory bank, obtain the record of cell attribute learning by searching the attribute learning memory bank, calculate the Hamming distance between the new category cell t and each old category cell (the old category cell is the cell data that has been learned and memorized) at the attribute feature level, and obtain the attribute similarity weight between the new category cell (new cell data) and each old category cell (old cell data).
[0106] The method of determining multiple similarity weight data between each cell attribute of the target new cell data and each cell data that has been learned and memorized in the previous attribute learning memory library according to each cell attribute of the target new cell data to obtain a similarity weight data set includes:
[0107] For each cell attribute of the target new cell data, calculating the Hamming distance between the cell attribute and the attribute feature of each cell data in the attribute learning memory library;
[0108] The similarity weight data corresponding to each of the cell attributes of the target new cell data are determined according to each of the Hamming distances, and the similarity weight data set is determined according to each of the similarity weight data.
[0109] When a new category cell t is input, the morphological attribute features of the new category cell t are first imported and added to the attribute learning memory library. Then, the attribute features of each old category cell in the attribute memory library are traversed, that is, each old category cell is read row by row, and the attribute corresponding to each cell is read column by column. The present invention uses the Hamming distance to calculate the distance between the attribute of the new category cell and the attribute of each old category cell, and performs a weighted summation on the distance of each attribute to obtain the attribute difference distance D between the new category cell and the i-th old category cell. i .
[0110] Specifically, for a new category cell t, its attribute features are extracted from the attribute learning memory bank and represented as M t, then traverse each old category cell c i Attribute characteristics And calculate the Hamming distance D, the formula is as follows:
[0111]
[0112] Where Q is the number of cell attributes defined in cell morphology. t [j] and Represent the feature labels of the new and old category cells on the jth attribute respectively.
[0113] According to the Hamming distance D(t,c i ), the similarity between the new category cell and each old category cell at the attribute level is represented as the weight w(t,c i ), that is, the similarity weight data set described in this application, the weight is inversely proportional to the distance, the greater the distance, the smaller the weight.
[0114]
[0115] The weight w(t,c i ), used to characterize the c i The attribute correlation between the old category cell and the new category cell t. The larger the weight, the higher the correlation between the two cells. In the subsequent fusion step, the high-dimensional extractor H corresponding to the old category cell t The larger the proportion of output, the more fully existing knowledge can be utilized, providing a basis for subsequent feature fusion and model updating.
[0116] Phase 2: When the tth new cell data is input, the model dynamically expands a high-dimensional attribute extractor H according to the attribute learning situation t , the low-dimensional attribute extractor L and the high-dimensional attribute extractor {H 1, H 2, …,H t-1} parameters, and the output of all high-dimensional attribute extractors {h 1, h 2, …,h t}(output update parameters described in the embodiment of the present application).
[0117] The incremental learning of the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data includes:
[0118] Obtain the previous attribute learning memory library of the cell morphology inspection system output in the previous round;
[0119] Determine, according to each cell attribute of the target new cell data, a plurality of similarity weight data between each cell attribute that has been learned and memorized in the previous attribute learning memory library, to obtain a similarity weight data set;
[0120] For each cell attribute of the target new cell data, a target cell attribute corresponding to the cell attribute is obtained in the previous attribute learning memory library, and attribute feature migration and fusion are performed on each of the cell attributes and the cell attributes according to the similarity weight data set to obtain a cell attribute fusion vector, wherein the target cell attribute is a cell attribute learned and memorized in historical rounds;
[0121] Determine the attribute feature vector of the target new cell data according to each of the cell attribute fusion vectors;
[0122] A feature fusion vector is determined according to the attribute feature vector and the attribute feature vectors of each cell data in historical rounds, and the previous attribute learning memory library is updated according to the feature fusion vector to complete the attribute incremental learning of the cell morphology inspection system.
[0123] When the image data of the newly added cell of the tth category (i.e., the target new cell data of the tth category) is input, the ILA method first freezes the attribute extractor {L,H 1, H 2, …,H t-1} (i.e., the corresponding feature output in the cell morphology inspection system output by the historical round), that is, close its gradient to prevent the model weight of the old task from being affected during the training of the new task (i.e., the incremental learning of the current round). At the same time, by searching the attribute learning memory library, the attribute features of the new category cell t that have been learned by the extractor in the previous stage (the previous stage represents the historical round) (i.e., the target cell attributes) and the attribute features that have not been learned by the extractor in the previous stage (i.e., the cell attributes that have not been learned in the historical rounds) are obtained, and the number of unlearned attributes is recorded as N. At this time, the model dynamically expands N residual blocks as a new high-dimensional attribute extractor H t This approach will H t It is used to learn the attribute features of new cell data t, so that the model can effectively learn the features of new cell data t; on the other hand, according to the learning situation of the attribute features, a high-dimensional attribute extractor is dynamically introduced to avoid the increased complexity caused by the static introduction of a fixed extractor, thereby reducing the loss of computing power.
[0124] At this time, each high-dimensional attribute extractor {H 1, H 2, …,H t}Parallel independent output model prediction feature vector h i , where hi is the output of the i-th high-dimensional attribute extractor, with dimension d. In general, the high-dimensional attribute extractor H t Compared with all high-dimensional attribute extractors {H 1, H 2, …,H t-1} parallel output feature vector {h 1, h 2, …,h t}.
[0125] Stage 3: All high-dimensional attribute extractor outputs {h1,h2,…,h t} Input to the attribute feature discriminator F t , attribute discriminator F t It contains the following three functions: (1) transform the feature vector {h1,h2,…,h t}After mapping in the fully connected layer, we get the attribute feature vector {A1,A2,…,A t}; (2) Obtain the learned attributes according to the attribute learning memory library, and transfer the learned attribute features in the attribute feature vector to A t (3) for the attribute feature vector {A1, A2, …, A t} Perform weighted dynamic feature fusion based on the attention mechanism. Finally, the attribute increment network is trained using the designed loss function.
[0126] The following is a detailed description of the contents of Phase 3:
[0127] like Figure 6 As shown, h t is the feature currently output by the new cell data t, which includes multiple cell attributes (i.e., attribute features). Then, an attribute incremental learning is performed on each cell attribute, such as cell attribute feature j. Figure 7 As shown, in the incremental learning of historical rounds (i.e. h1,h2,…h t-1 ) to find out which cell attribute features j have been learned. Figure 7 In the above example, h1 and h3 have already learned the target cell attribute j. At this time, the similarity weight data in stage 1 is used to fuse with Aj in the current round of incremental learning At to obtain Aj-fusion, which is the cell attribute fusion vector of the jth cell attribute. Then, by fusing these cell attribute fusion vectors, the attribute feature vector At of the current round (i.e., the attribute feature vector described in stage 3) can be obtained. At corresponds to the new category cell t.
[0128] Updating the previous attribute learning memory base according to the feature fusion vector may be specifically as follows:
[0129] The attribute feature vector At of the current round is combined with the attribute feature vector set {A1, A2, …, A t-1} to merge, that is, Figure 8 As shown, for the current round of feature fusion vector Afusion, the attribute incremental learning is completed.
[0130] The output of the high-dimensional attribute extractor {h1,h2,…,h t}, and input it into the attribute feature discriminator. i Map it to n dimensions through a fully connected layer to obtain the attribute feature vector A i At this time, through the feedback of the attribute learning memory bank, the attribute features of the new category cell t that have been learned by the previous stage extractor are obtained. For each learned attribute feature, the index and number M corresponding to the learned attribute and the corresponding high-dimensional attribute extractor number are recorded. According to the index range corresponding to each attribute, each attribute feature A can be intercepted i In the output A of one of the learned attribute features ij , which represents the output of the i-th high-dimensional attribute extractor that has learned the attribute feature j.
[0131] According to the number M of extractors corresponding to the learned attribute features, the cell attribute similarity weights obtained in stage 1 are imported to filter out the weights of the old category cells corresponding to the high-dimensional attribute extractors that have learned the attribute, and these weights are normalized to facilitate subsequent feature fusion.
[0132]
[0133] By fusing the attribute feature A corresponding to the high-dimensional extractor that has learned the attribute feature ij , and obtain the fusion feature vector Aj-fusion of attribute j. The calculation formula is as follows:
[0134]
[0135] Replace the newly expanded high-dimensional attribute extractor H with Aj-fusion (feature fusion vector) t The corresponding attribute feature vector A t This is equivalent to fusing the output of the model that has learned the attribute features, and using this fused output as the newly expanded high-dimensional attribute extractor H t The output of attribute j. This operation traverses all the learned attributes of the newly added category cell t and finally completes A t Transfer of learned attribute feature outputs in .
[0136] The attribute incremental learning and category incremental learning of the present application are performed using two different models, namely the attribute incremental network and the category incremental network. The model of the attribute incremental network uses the attention weighted dynamic fusion method to process the attribute feature vector {A 1, A 2, …,A t}, t attribute feature vectors from different high-dimensional attribute extractors are merged according to their importance. The present invention uses a feedforward neural network to generate an attention score β for each attribute feature vector. The structure of the feedforward network consists of two fully connected layers. For each attribute feature vector A i , the feed-forward network first maps the feature vector to the hidden layer through the first fully connected layer:
[0137]
[0138] Among them, W1 is the weight matrix of the first fully connected layer, with a size of [d, e], is the attention score of the i-th attribute feature vector after the first layer of attention network, b1 is the bias vector of this layer, with a size of [e], where e is the size of the hidden layer. Next, the activation function ReLu is introduced to increase the nonlinear expression ability of the model:
[0139]
[0140] After ReLU activation, the feedforward network further maps the output of the hidden layer to a scalar through a second fully connected layer:
[0141]
[0142] Among them, W2 is the weight matrix of the second fully connected layer, with a size of [e,1], and b2 is the bias term, with a size of [1]. i Is a scalar, representing the attribute feature vector A i Attention score. When the network is initialized, the weight matrices W1 and W2 are randomly generated by He initialization to ensure that the initial values of the weights have an appropriate scale, which is conducive to the training stability of the deep network. Throughout the process, the weights W1, W2 and biases b1, b2 of the feedforward neural network are all trainable parameters optimized by the back-propagation algorithm.
[0143] To assign these attention scores β i Converted to an attention weight α that can be directly applied i , it is necessary to i The present invention introduces the Softmax function to normalize all scores to probability values between 0 and 1, and the sum is 1.
[0144]
[0145] Among them, α i is the attribute feature vector A i The final attention weight of indicates the importance of the attribute feature in weighted fusion. All attention weights α i Normalization ensures that their sum is 1, so that they can be used in the weighted sum operation.
[0146] Then, each attribute feature vector {A1,A2,…,A t}(attribute feature vector set) is weighted and summed by attention weights to obtain a fused feature representation Afusion. The specific formula is as follows:
[0147]
[0148] Among them, Afusion (target mapping parameter) is the fused attribute feature vector, which represents the comprehensive representation of all attribute feature vector outputs from different high-dimensional attribute extractors. This fusion feature contains all the attribute feature vectors from {H1, H2, …, H t}The key information of the high-level feature extractor is obtained, and the weight of each feature extractor is dynamically adjusted according to the different input situations through the attention mechanism, thereby achieving adaptive feature fusion and capturing the key information from all related extractors.
[0149] Finally, by applying the Softmax activation function, these raw prediction values are converted into distributions representing the probability of each attribute. Specifically, according to the index range corresponding to each attribute, each attribute is traversed. For attribute k, the calculation formula of the Softmax function is:
[0150]
[0151] Among them, A fusion-k is the output of the fusion attribute feature vector Afusion for attribute k, and C is the number of features corresponding to an attribute. Through the above transformation, for each attribute k, the probability p(k) of its attribute feature is between 0 and 1, and for this attribute, the sum of the probabilities of its attribute features is 1. Finally, all attribute prediction probabilities are saved in a two-dimensional array P, such as: P = [p(1), p(2), ..., p(Q)], where Q is the total number of attributes.
[0152] The training set S for each incremental step t t , which consists of the dataset U of the newly introduced category t and the old category sample set V stored in memory t-1 Composition, namely S t =U t ∪Vt-1 The old category samples in the memory are obtained by random sampling, and each old category cell stores L samples. By combining new data with a small amount of old data, the model can learn the characteristics of the new category cell attributes in each incremental step while maintaining the ability to recognize the characteristics of the old category cell attributes.
[0153] The loss function of the training process is L old With L new The two loss functions are used alternately in each training step.
[0154] In one embodiment of the present application, the cell morphology inspection system is trained according to a preset class balance loss function, a preset cosine distillation loss function, a preset feature enhancement loss function and a preset attribute space separation loss function, and the class balance loss function, the cosine distillation loss function, the feature enhancement loss function and the attribute space separation loss function are used for attribute incremental learning.
[0155] L new It focuses on learning the features of new categories of blood cell data and consists of four parts:
[0156] L new =L class +λ1·L enhance +λ2·L att-sep +λ3·L cos-dist
[0157] L old It focuses on helping the model remember old categories of blood cell data. At the same time, it focuses on old categories with fewer samples by assigning higher penalties to incorrectly predicted samples. It consists of three parts:
[0158] L old =L class +λ4·L cos-dist +λ5·L att-sep
[0159] The following are the design ideas of the four loss functions:
[0160] Class balance loss L class :
[0161] A class-balanced focal loss is adopted to avoid severe forgetting due to the imbalance of data amount between new and old classes.
[0162]
[0163] Where n is the data set S tThe number of samples in , p(k) represents the predicted probability of the kth attribute output by the model, β is a hyperparameter for controlling category imbalance, which is used to impose additional penalties on scarce old category data to ensure that they are not ignored during training. γ is a hyperparameter for controlling task difficulty, which strengthens the learning of old category samples, allowing the model to better learn those challenging features.
[0164] The present invention assigns a higher penalty to the old class, L class It takes into account the class imbalance when learning features to better distinguish old and new classes. It avoids discarding samples that are valuable for learning accurate decision boundaries. old The hyperparameters β, γ and L in new The hyperparameters used in are different, so the former penalizes misclassification of old class samples more.
[0165] Cosine distillation loss L cos-dist :
[0166] The old knowledge of the model is preserved by minimizing the difference in attribute prediction between the new model being trained and the old model imported from the previous stage. The loss function consists of two parts: attribute difference loss L att And cosine similarity loss L cos , the calculation formula is as follows:
[0167] L cos-dist =L att +L cos
[0168] Among them, the attribute difference loss L att Used to measure the component-by-component difference between the new and old models on each attribute; cosine similarity loss L cos , which is used to evaluate the similarity between the new and old models in the overall direction of the attribute prediction vector.
[0169] Attribute difference loss L att The design is as follows: First, the attribute feature output A of the new model new And the attribute feature output A of the old model old Temperature softening to obtain and The calculation formula is as follows:
[0170]
[0171] Where T is the temperature parameter, which is used to control the smoothness of attribute prediction. old (k) represents the kth attribute feature output of the old model in the previous stage, Represents the kth attribute feature output of the new model. att The calculation of is as follows:
[0172]
[0173] Attribute difference loss L att , quantified the component-by-component difference between the new and old models in the characteristic values of each attribute, reflecting the accuracy of the two models in predicting specific attributes, emphasizing that the model needs to maintain the accuracy of each attribute during the updating process, ensuring that the model can not only adapt to new tasks during incremental learning, but also effectively retain and utilize the knowledge of old tasks.
[0174] Cosine similarity loss L cos It measures the directional similarity between the attribute prediction vectors of the new and old models, emphasizing the consistency of the model in the overall prediction trend. It evaluates the relative position of two vectors in space by calculating the angle between them, thus reflecting the model's overall understanding of the attribute features. When the model maintains a similar prediction direction during the update process, although the specific feature values may be different, the consistency and integration of new and old knowledge can still be ensured. This enables the model to transfer existing knowledge more smoothly when facing new tasks, reducing uncertainty in the learning process.
[0175]
[0176] Finally, the attribute difference loss L att And cosine similarity loss L cos By adding the ratio ε, we get the cosine distillation loss L cos-dist .
[0177]
[0178] Among them, the combination ratio ε is used to adjust the relative importance of attribute difference loss and cosine similarity loss in the final loss, and the value range is [0,1].
[0179] Feature enhancement loss L enhance
[0180] In order to learn the attribute characteristics of new categories of cells in a targeted manner, the present invention introduces the feature enhancement loss L enhance .L enhance With L class Both use the class balance loss calculation method, but L enhance Only focus on the newly expanded high-dimensional attribute extractor H t Output h t The attribute feature A obtained after mapping t .
[0181]
[0182] Among them, p t(k) represents attribute feature A t The predicted probability of the kth attribute in .
[0183] The loss L of attribute space separation att-sep
[0184] The purpose of the attribute space separation loss is to ensure that the new class and the old class remain separated in the attribute space by measuring the L2 distance between the attribute prediction values of the new model for the new and old category data. In this way, when the model learns a new category, by constraining the attribute predictions of the new and old categories not to be too close, the interference of the learning of the new category cells on the knowledge of the old category cells is reduced, thereby alleviating the "forgetting" problem in incremental learning.
[0185]
[0186] in, and They are the newly expanded high-dimensional attribute extractor H t For the dataset S t and dataset V t-1 Output attribute feature vector of the kth attribute. min It is the preset minimum distance, ensuring that the new and old category cells are sufficiently separated in the attribute space.
[0187] During model training, by alternating between L new and L old Calculate and optimize the loss, and finally train the high-dimensional attribute extractor H t , attribute discriminator F t Through this training mode, the network can continuously learn the attribute features of newly added categories, while reducing the forgetting of the attribute features of old categories, thereby realizing the incremental learning based on cell attributes of the present invention.
[0188] Phase 4: Load the model trained in Phase 3 (i.e., obtain Afusion), update the category classifier Ct based on the attribute incremental network structure, and use Ct to map the feature fusion vector Afusion to the current cell category B t (i.e., the mapping relationship between the feature fusion vector and the target new cell data described in this application), using the preset cross entropy loss function and the preset distillation loss function, train and optimize the category incremental network to achieve incremental learning at the cell category level (i.e., the category incremental learning described in this application).
[0189] Through the training in stage three, an attribute discriminator Ft is obtained that can classify the attribute characteristics of new and old categories of cells. This application loads the model obtained in stage three, and updates the category classifier Ft at the same time, and modifies its fully connected layer output to t (because t-1 types of cells have been learned before, it needs to be modified to t, and each cell category data can be identified in parallel). The parameter t changes dynamically with the number of new categories. The fused attribute feature vector Afusion output by the attribute discriminator Ft in the attribute increment network of the third stage is mapped to the category through the updated classifier Ct, realizing the mapping of the model from attributes to categories, and obtaining the category feature vector B t . Use the Softmax function to transform the category feature vector B t Converted into a probability distribution p t , the processing is as follows, where t represents the total number of categories:
[0190]
[0191] Use the loss function to train the category feature vector p i , we obtain the category incremental network and realize the incremental learning of cell categories.
[0192] The loss function of the training process is L cls , which consists of two parts: L cls =L dist +L ce , which helps the model learn new categories of cells without forgetting old categories of cells. The dataset is S t , using cell category labels. The following are the specific designs of the two functions:
[0193] Cross entropy loss L ce :
[0194]
[0195] Among them, q i is the true label of the cell category, p i The probability distribution predicted by the model. The cross entropy loss function is used to measure the difference between the predicted probability distribution and the true label probability distribution, guiding model training to minimize the difference, thereby improving classification accuracy.
[0196] Distillation loss L dist :
[0197]
[0198] Among them, T is the temperature parameter, which is used to adjust the smoothness of the probability distribution in the Softmax function. and The old model and the new model are respectivelyt The Softmax probability distribution of is calculated as follows:
[0199]
[0200] Distillation loss L dist Measuring the difference between the output of the new model and the output of the old model, through this measurement, the new model can learn and inherit the behavior of the old model, especially in terms of probability distribution. While keeping the model size small, the new model can absorb the deep knowledge and probability relationship of the old model when dealing with different categories, thereby maintaining or improving the performance and generalization ability of the model as much as possible while reducing the computational complexity.
[0201] Specifically, by storing the new cell data obtained through continuous learning and memory in the attribute learning memory bank and turning it into old cell data, the new cell data of new categories can be continuously learned without forgetting the old cell data of old categories, and multiple rounds of iterative updates of the cell morphology inspection system can be performed, thereby solving the problems of weak analysis and learning capabilities existing in the prior art.
[0202] Can be as Fig. 9 As shown, Fig. 9 This is a flowchart of incremental learning of an embodiment of the present application, which screens out a new category cell data set d by inputting a cell sample data stream (a first data set D1), and labels the new category cell samples in the new category cell data set d. The labeled samples are integrated into the model of the cell morphology inspection system through the ILA method to achieve system evolution.
[0203] Fig.10 This is a logic diagram of an iterative device of a cell morphology inspection system based on incremental learning in an embodiment of the present application. Fig.10 This is a block diagram of an iterative device 300 for a cell morphology inspection system based on incremental learning according to an embodiment of the present application. According to an iterative device 300 for a cell morphology inspection system based on incremental learning according to an embodiment of the present application, the system 300 includes: an acquisition unit 301, an input unit 302, and an incremental learning unit 303.
[0204] An acquisition unit 301 is used to acquire image data of target cells and construct a first data set according to the image data;
[0205] An input unit 302 is used to input the first data set into the out-of-distribution sample detection module to obtain a new category cell data set in the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to represent cell data that has not been learned and memorized;
[0206] The incremental learning unit 303 is used to perform iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module until each of the new cell data is marked as learned and memorized cell data, and output a final round of target cell morphology inspection system;
[0207] The incremental learning of the current round of the cell morphology inspection system includes the following steps:
[0208] Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell;
[0209] The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
[0210] As another aspect, the present application further provides a computer-readable storage medium on which a program product capable of implementing the method provided above in this specification is stored. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary implementations of the present application described in the above "Embodiment Method" section of this specification.
[0211] According to the embodiment of the present application, the program product for implementing the above method can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0212] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0213] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0214] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0215] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0216] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0217] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be performed without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An iterative method for a cell morphology inspection system based on incremental learning, characterized in that: Applied to a cell morphology inspection system, the cell morphology inspection system includes a preset out-of-distribution sample detection module and an attribute learning memory library, the attribute learning memory library stores a plurality of learned and memorized cell data, the method includes: S1, acquiring image data of target cells, and constructing a first data set according to the image data; S2, inputting the first data set into the out-of-distribution sample detection module to obtain a new category cell data set in the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to represent cell data that has not been learned and memorized; S3, performing iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module, until each of the new cell data is marked as learned and memorized cell data, and outputting the final round of target cell morphology inspection system; The incremental learning of the current round of the cell morphology inspection system includes the following steps: Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell; The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
2. The iterative method for cell morphology inspection system based on incremental learning according to claim 1, characterized in that: The incremental learning includes attribute incremental learning and category incremental learning. The attribute incremental learning is used to represent the incremental learning of attribute characteristics of each cell attribute of the target new cell data, and the category incremental learning is used to represent the incremental learning of category characteristics of the target new cell data.
3. The iterative method for cell morphology inspection system based on incremental learning according to claim 2 is characterized in that: The incremental learning of the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data includes: Obtain the previous attribute learning memory library of the cell morphology inspection system output in the previous round; Determine, according to each cell attribute of the target new cell data, a plurality of similarity weight data between each cell attribute that has been learned and memorized in the previous attribute learning memory library, to obtain a similarity weight data set; For each cell attribute of the target new cell data, a target cell attribute corresponding to the cell attribute is obtained in the previous attribute learning memory library, and attribute feature migration and fusion are performed on each of the cell attributes and the cell attributes according to the similarity weight data set to obtain a cell attribute fusion vector, wherein the target cell attribute is a cell attribute learned and memorized in historical rounds; Determine the attribute feature vector of the target new cell data according to each of the cell attribute fusion vectors; A feature fusion vector is determined according to the attribute feature vector and the attribute feature vectors of each cell data in historical rounds, and the previous attribute learning memory library is updated according to the feature fusion vector to complete the attribute incremental learning of the cell morphology inspection system.
4. The iterative method for cell morphology inspection system based on incremental learning according to claim 3 is characterized in that: The incremental learning of the cell morphology inspection system is performed by the following steps: Binding the feature fusion vector to the target new cell data to determine a mapping relationship between the feature fusion vector and the target new cell data; The category features of the target new cell data are determined according to the mapping relationship, and the out-of-distribution sample detection module is updated according to the category features to complete the category incremental learning of the cell morphology inspection system.
5. The iterative method for cell morphology inspection system based on incremental learning according to claim 4 is characterized in that: The method of determining multiple similarity weight data between each cell attribute of the target new cell data and each cell data that has been learned and memorized in the previous attribute learning memory library according to each cell attribute of the target new cell data to obtain a similarity weight data set includes: For each cell attribute of the target new cell data, calculating the Hamming distance between the cell attribute and the attribute feature of each cell data in the attribute learning memory library; The similarity weight data corresponding to each of the cell attributes of the target new cell data are determined according to each of the Hamming distances, and the similarity weight data set is determined according to each of the similarity weight data.
6. The iterative method for cell morphology inspection system based on incremental learning according to claim 5, characterized in that: The new category cell dataset is obtained by the following steps: Obtaining the category probability and distance score of each of the image data in the first data set through an out-of-distribution sample detection module; For each image data, if the product value of the category probability and the distance score of the image data is greater than a preset threshold, then determining that the cell data corresponding to the image data is the new cell data; Wherein, each of the image data corresponds to single cell data.
7. The iterative method for cell morphology inspection system based on incremental learning according to claim 6, characterized in that: The cell morphology inspection system is trained according to a preset class balance loss function, a preset cosine distillation loss function, a preset feature enhancement loss function and a preset attribute space separation loss function, and the class balance loss function, the cosine distillation loss function, the feature enhancement loss function and the attribute space separation loss function are used for attribute incremental learning.
8. The iterative method for cell morphology inspection system based on incremental learning according to claim 7, characterized in that: The cell morphology inspection system is also trained according to a preset cross entropy loss function and a preset distillation loss function, and the cross entropy loss function and the distillation loss function are used for category incremental learning.
9. An iterative device for a cell morphology inspection system based on incremental learning, characterized in that: Applied to a cell morphology inspection system, the cell morphology inspection system includes a preset out-of-distribution sample detection module and an attribute learning memory library, the attribute learning memory library stores a plurality of learned and memorized cell data, and the device includes: an acquisition unit, configured to acquire image data of target cells and construct a first data set according to the image data; An input unit, used for inputting the first data set into the out-of-distribution sample detection module to obtain a new category cell data set in the first data set, wherein the new category cell data set includes a plurality of new cell data, and the new cell data is used to represent cell data that has not been learned and memorized; An incremental learning unit, used for performing iterative rounds of incremental learning on the cell morphology inspection system according to each of the new cell data to update the attribute learning memory library and the out-of-distribution sample detection module, until each of the new cell data is marked as learned and memorized cell data, and outputting a final round of target cell morphology inspection system; The incremental learning of the current round of the cell morphology inspection system includes the following steps: Acquire the target new cell data of the current round, and perform incremental learning on the cell morphology inspection system outputted in the previous round according to each cell attribute of the target new cell data, until each cell attribute of the target new cell data is incrementally learned, and obtain the feature fusion vector of the target new cell; The attribute learning memory library and the out-of-distribution sample detection module are updated according to the feature fusion vector to output the cell morphology inspection system of the current round.
Citation Information
Patent Citations
Method for performing class addition learning on microscope cell image detection model by using incremental learning
CN110059672A
Target counting-oriented class incremental learning network modeling method and device
CN115937597A
Deep network incremental learning method for multi-view children tumor pathological image classification
CN116363461A
Method and system for cell annotation with adaptive incremental learning
EP3478728A1
Meta few-shot class incremental learning
US20230153380A1