Electrophoresis image recognition method and device, electronic equipment and storage medium
Through multi-layer classification model and image enhancement technology, the misjudgment problem caused by the inconspicuous banding area in the electrophoretic image is solved, the accurate classification of the electrophoretic image and the correct recognition of the lane features are achieved, and the accuracy and reliability of the electrophoretic image recognition are improved.
Patent Information
- Application Number
- CN202510427188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
In the immunofixed electrophoretic image processing, the prior art causes chromatic aberration and blur of the band area due to differences between samples and unstable electrophoretic conditions, resulting in misjudgment or neglect, which affects the accuracy of electrophoretic image classification.
The multi-layer classification model is used to identify the electrophoretic images, including the first classification model for negative and positive prediction, the second classification model further analyzes the suspected positive images, and the third classification model recognizes the lane features, optimizes the model parameters through image enhancement and supervision comparison learning to improve classification accuracy.
The accurate classification of electrophoretic images, especially the accurate identification of suspected positive images and the correct verification of lane features, improve the accuracy and reliability of electrophoretic image recognition.
Smart Images

Figure CN120356047A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to an electrophoresis image recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] In the current process of processing immunofixation electrophoresis (IFE) images, due to factors such as differences between samples, inconsistent shooting angles of electrophoresis images, and unstable electrophoresis conditions, the color difference in the band region of the lanes in the electrophoresis image often weakens, the band region appears unclear, or blurring occurs. Due to the unclear morphological features and low color rendering of the band region in the electrophoresis image, it is extremely easy to cause misjudgment or be ignored, resulting in inaccurate classification of the electrophoresis image. Summary of the Invention
[0003] Embodiments of this application disclose an electrophoresis image recognition method, apparatus, electronic device, and storage medium, which can enable a computer device to accurately classify electrophoresis images, thereby improving the accuracy of the computer device in recognizing electrophoresis images.
[0004] The first aspect of the embodiments of this application discloses an electrophoresis image recognition method, including:
[0005] The computer device recognizes a target electrophoresis image through a first classification model to obtain a positive / negative prediction value corresponding to the target electrophoresis image, and determines a first classification result corresponding to the target electrophoresis image according to the positive / negative prediction value; the first classification model is trained according to a first data set, and the first data set includes multiple first sample electrophoresis images and the overall positive / negative labels corresponding to each of the first sample electrophoresis images;
[0006] If the first classification result is suspected positive, the computer device analyzes the target electrophoresis image through a second classification model to obtain a second classification result corresponding to the target electrophoresis image; the second classification model is trained according to a suspected data set, and the suspected data set includes multiple second sample electrophoresis images marked as suspected data and the overall positive / negative labels corresponding to each of the second sample electrophoresis images;
[0007] If the first classification result or the second classification result is positive, the computer device recognizes each lane included in the target electrophoresis image through a third classification model to obtain a third classification result corresponding to each lane; the third classification model is trained according to a second data set, and the second data set includes multiple third sample electrophoresis images and the positive / negative labels corresponding to each lane in each of the third sample electrophoresis images;
[0008] The computer device determines the target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result, or the third classification result.
[0009] In some possible embodiments, the determining the first classification result corresponding to the target electrophoresis image according to the positive-negative prediction value includes:
[0010] If the positive-negative prediction value is greater than the first prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is positive;
[0011] If the positive-negative prediction value is less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is negative;
[0012] If the positive-negative prediction value is not greater than the first prediction threshold and not less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is suspected positive.
[0013] In some possible embodiments, the method further includes:
[0014] The computer device obtains multiple second-sample electrophoresis images included in the current batch;
[0015] The computer device performs image enhancement processing on each second-sample electrophoresis image in the current batch through two different image enhancement algorithms to obtain two enhanced training images corresponding to each second-sample electrophoresis image;
[0016] The computer device generates embedding vectors corresponding to the two enhanced training images respectively according to the two enhanced training images corresponding to each second-sample electrophoresis image through a second classification model to be trained;
[0017] The computer device combines the multiple enhanced training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs;
[0018] The computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs;
[0019] The computer device calculates a loss value according to the similarities of the multiple groups of sample pairs;
[0020] The computer device adjusts the parameters of the second classification model to be trained according to the loss value until the second classification model to be trained meets the convergence condition.
[0021] In some possible embodiments, the computer device combines multiple enhanced training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs, including:
[0022] The computer device converts the embedding vectors corresponding to the multiple enhanced training images corresponding to the current batch into a first tensor;
[0023] The computer device combines and splices the first tensors corresponding to the multiple enhanced training images in pairs to obtain second tensors corresponding to the multiple groups of sample pairs.
[0024] In some possible embodiments, the computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs, including:
[0025] The computer device performs a dot product calculation on the embedding vectors corresponding to the two enhanced training images included in the target sample pair to obtain the initial similarity corresponding to the target sample pair, where the target sample pair is any one of the sample pairs;
[0026] The computer device adjusts the initial similarity according to the temperature coefficient to obtain the similarity corresponding to the target sample pair.
[0027] In some possible embodiments, before the computer device calculates the loss value according to the similarity of the positive sample pairs, the method further includes:
[0028] The computer device obtains the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs;
[0029] The computer device generates a mask according to the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs;
[0030] The computer device calculates the loss value according to the similarity of the multiple groups of sample pairs, including:
[0031] The computer device determines the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs included in the current batch according to the mask and the similarity of each group of sample pairs; the positive sample pair is a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included are the same, and the negative sample pair is a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included are different;
[0032] The computer device calculates the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs, where the first positive sample pair is any one of the positive sample pairs;
[0033] The computer device calculates the average value according to the ratio corresponding to each group of the positive samples, and calculates the loss value according to the average value.
[0034] In some possible embodiments, the second classification model to be trained includes a feature extractor and a fully connected layer; the computer device generates embedding vectors corresponding to the two enhanced training images respectively according to the two enhanced training images corresponding to each second sample electrophoresis image through the second classification model to be trained, including:
[0035] The computer device extracts the features of the two enhanced training images corresponding to each second sample electrophoresis image through the feature extractor to be trained, and obtains the embedding vectors corresponding to the two enhanced training images respectively;
[0036] The computer device adjusts the parameters of the second classification model to be trained according to the loss value until the second classification model to be trained meets the convergence condition, including:
[0037] The computer device adjusts the parameters of the feature extractor to be trained according to the loss value until the feature extractor to be trained meets the convergence condition, and the feature extractor training is completed;
[0038] The method further includes:
[0039] When the feature extractor training is completed, the computer device generates a feature vector corresponding to the input second sample electrophoresis image through the trained feature extractor, and determines the predicted classification result corresponding to the input second sample electrophoresis image through the fully connected layer to be trained according to the feature vector corresponding to the input second sample electrophoresis image;
[0040] The computer device adjusts the parameters of the fully connected layer according to the predicted classification result corresponding to the input second sample electrophoresis image and the corresponding overall positive and negative label until the fully connected layer to be trained meets the convergence condition, and then the fully connected layer training is completed, and a trained second classification model is obtained.
[0041] A second aspect of the embodiments of the present application discloses an electrophoresis image recognition device, and the device includes:
[0042] A first classification module, configured to recognize a target electrophoresis image through a first classification model, obtain a positive and negative predicted value corresponding to the target electrophoresis image, and determine a first classification result corresponding to the target electrophoresis image according to the positive and negative predicted value; the first classification model is trained according to a first data set, and the first data set includes multiple first sample electrophoresis images and overall positive and negative labels corresponding to each of the first sample electrophoresis images;
[0043] A second classification module, configured to, if the first classification result is a suspected positive, analyze the target electrophoresis image through a second classification model to obtain a second classification result corresponding to the target electrophoresis image; the second classification model is trained according to a suspected data set, and the suspected data set includes multiple second sample electrophoresis images labeled as suspected data, and the overall positive and negative labels corresponding to each of the second sample electrophoresis images;
[0044] A third classification module, configured to, if the first classification result or the second classification result is positive, identify each lane included in the target electrophoresis image through a third classification model to obtain a third classification result corresponding to each lane; the third classification model is trained according to a second data set, and the second data set includes multiple third sample electrophoresis images, and the positive and negative labels corresponding to each lane in each of the third sample electrophoresis images;
[0045] An image recognition module, configured to determine a target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result, or the third classification result.
[0046] A third aspect of the embodiments of the present application discloses an electronic device, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is enabled to implement the electrophoresis image recognition method according to any one of the foregoing embodiments.
[0047] A fourth aspect of the embodiments of the present application discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to implement the electrophoresis image recognition method according to any one of the foregoing embodiments.
[0048] A method, device, electronic device and storage medium for identifying an electrophoresis image provided by the present application. The computer device can first identify a target electrophoresis image through a first classification model to determine a first classification result of the target electrophoresis image. When the first classification result is suspected to be positive, the computer device analyzes the target electrophoresis image through a second classification model to obtain a second classification result corresponding to the target electrophoresis image. When the first classification result or the second classification result is positive, the computer device identifies each lane included in the target electrophoresis image through a third classification model to obtain a third classification result corresponding to each lane. Finally, the computer device determines a target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result or the third classification result, so that the computer device can further analyze the target electrophoresis image through the second classification model when the first classification result of the target electrophoresis image is suspected to be positive, realizing accurate classification of the target electrophoresis image suspected to be positive, and when the first classification result or the second classification result is positive, the third classification result corresponding to each lane in the target electrophoresis image is also obtained through the third classification model to further verify the accuracy of the classification of the target electrophoresis image from the classification results of the lanes, thereby improving the accuracy of the computer device in identifying electrophoresis images. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 A schematic diagram of an electrophoresis image provided by an embodiment of the present application;
[0051] Figure 2 A flowchart of a method for identifying an electrophoresis image provided by an embodiment of the present application;
[0052] Figure 3 A flowchart of training a first classification model to be trained provided by an embodiment of the present application;
[0053] Figure 4 A flowchart of training a third classification model to be trained provided by an embodiment of the present application;
[0054] Figure 5 A flowchart of training a second classification model to be trained provided by an embodiment of the present application;
[0055] Figure 6 A flowchart of obtaining multiple groups of sample pairs according to multiple enhanced training images included in the current batch provided by an embodiment of the present application;
[0056] Figure 7 Schematic diagram for splicing the first tensors corresponding to two enhanced training images respectively to obtain a second tensor provided by an embodiment of the present application;
[0057] Figure 8 Flowchart for calculating the similarity of each group of sample pairs provided by an embodiment of the present application;
[0058] Figure 9 Flowchart for calculating a loss value according to the similarities of multiple groups of sample pairs provided by an embodiment of the present application;
[0059] Figure 10 Another flowchart for training a second classification model to be trained provided by an embodiment of the present application;
[0060] Figure 11 Block diagram of a structure of an electrophoresis image recognition device provided by an embodiment of the present application;
[0061] Figure 12 Block diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0063] It should be noted that the terms "include" and "have" and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0064] In addition, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the relationship between associated objects and indicates that three relationships can exist. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0065] The electrophoresis image recognition method provided by the embodiments of the present application can be applied to a computer device, which may include, but is not limited to, a personal computer, a tablet computer, a laptop computer, a mobile phone, etc.
[0066] Exemplarily, Figure 1 is a schematic diagram of an electrophoresis image provided by the embodiments of the present application. As Figure 1 shown, the computer device can acquire an electrophoresis image generated by the IFE technology. The electrophoresis image includes a reference lane ELP (Electrophoresis, abbreviated as ELP), a heavy chain lane, and a light chain lane. The heavy chain lane includes lane G, lane A, and lane M, and the light chain lane includes lane κ and lane λ. Among them, the reference lane ELP includes multiple band regions with different color development degrees. Lane G and lane κ have band regions with higher color development degrees and clearer boundaries. Lane λ has band regions with lower color development degrees and blurrier boundaries. Lane A and lane M have band regions with extremely low color development degrees. The band regions corresponding to lane G and lane κ in the reference lane ELP have lower color development degrees in the dotted line region 101.
[0067] In the electrophoresis image, if one or more lanes in the heavy chain lane and the light chain lane have band regions corresponding to the band regions of the reference lane ELP, that is, when the same band region exists in the reference lane ELP and one or more lanes in the heavy chain lane and the light chain lane at the same time, it can be considered that the electrophoresis image shows positive, and the lanes with the corresponding band regions show positive. On the contrary, if each lane in the heavy chain lane and the light chain lane does not have a band region corresponding to each band region in the reference lane ELP, it can be considered that the electrophoresis image shows negative, and the heavy chain lane and the light chain lane corresponding to the electrophoresis image show negative.
[0068] The recognition process of electrophoresis images mainly includes two aspects. One is the recognition of the overall positive and negative of the electrophoresis image, and the other is the recognition of the positive and negative corresponding to each lane in the electrophoresis image. In the recognition process of electrophoresis images, it is necessary to first recognize the overall positive and negative of the electrophoresis image. When it is determined that the electrophoresis image is positive, then recognize the positive and negative of each lane of the electrophoresis image to determine the positive and negative classification results of each lane of the electrophoresis image; when the electrophoresis image is recognized as negative, then each lane of the electrophoresis image must also be negative, so there is no need to recognize the positive and negative of each lane.
[0069] In some embodiments, the overall positive and negative of the electrophoresis image include positive and negative, but the overall positive and negative labels of the electrophoresis image include positive, negative, and suspected positive. Suspected positive can refer to the ambiguous area between positive and negative, that is, it cannot be clearly determined whether the electrophoresis image shows positive or negative.
[0070] For an electrophoresis image with a suspected positive overall positive and negative label, it is usually because the band area in the lane of the electrophoresis image has unclear morphological changes (such as unclear boundaries of the band area, blurred band area) or low color development, resulting in misjudgment or neglect of the above band area, making the overall positive and negative classification of the electrophoresis image inaccurate. For example, in Figure 1 in the dotted area 101 of the shown electrophoresis image, since the color development of the band area of the reference lane ELP in the dotted area 101 is low and the boundary is also relatively blurred, by manually identifying this electrophoresis image, especially in the process of batch identifying electrophoresis images, it is very easy to misjudge or neglect whether there is a band area showing positive in the reference lane ELP in the dotted area 101, making the overall positive and negative classification of the electrophoresis image inaccurate.
[0071] The embodiment of the present application provides an electrophoresis image recognition method, which can enable a computer device to accurately classify a suspected positive electrophoresis image to improve the accuracy of the computer device in recognizing electrophoresis images.
[0072] As Figure 2 shown, in one embodiment, an electrophoresis image recognition method is provided. The method may include the following steps:
[0073] Step 202, the computer device recognizes the target electrophoresis image through a first classification model to obtain the positive and negative prediction value corresponding to the target electrophoresis image, and determines the first classification result corresponding to the target electrophoresis image according to the positive and negative prediction value.
[0074] The first classification model is trained based on a first dataset, which includes multiple first sample electrophoresis images and the overall positive / negative labels corresponding to each first sample electrophoresis image. The overall positive / negative labels can be used to identify the positive / negative nature presented by the first sample electrophoresis images, and the overall positive / negative labels can include positive, negative, and suspected positive.
[0075] The target electrophoresis image can include the electrophoresis image obtained by subjecting a target sample to IFE processing. The target sample can include one or more of a serum sample, a urine sample, or other biological samples containing immunoglobulins.
[0076] The positive / negative prediction value can refer to the confidence level used to quantify that the first classification result of the target electrophoresis image is positive. The first classification result can include negative, positive, and suspected positive. The larger the positive / negative prediction value, the higher the degree to which the target electrophoresis image appears positive, that is, the greater the probability that the first classification result of the target electrophoresis image is positive; conversely, the smaller the positive / negative prediction value, the higher the degree to which the target electrophoresis image appears negative, that is, the greater the probability that the first classification result of the target electrophoresis image is negative; when the positive / negative prediction value is within a specific intermediate range, the first classification result of the target electrophoresis image is suspected positive, indicating that it may be positive or negative, and further analysis is required to determine its positive / negative nature.
[0077] In some embodiments, the first classification model includes a classification prediction value generation module and a classification probability calculation module. The computer device can construct a first classification model to be trained for binary positive / negative classification through a deep learning algorithm, and train the first classification model to be trained based on the first dataset. After the first classification model is trained, the computer device can retain the classification prediction value generation module in the trained first classification model that is used to output the classification prediction value, and remove the classification probability calculation module that is used to convert the classification prediction value into a classification probability, so as to output the classification prediction value as the positive / negative prediction value.
[0078] In some embodiments, the first classification model can include a neural network model. The classification prediction value generation module includes a fully connected layer, and the classification probability calculation module includes a sigmoid function (Softmax) layer. Among them, the fully connected layer can output the original prediction value for classifying the electrophoresis image, that is, the Logits value. The Softmax layer can be used to convert the Logits value output by the fully connected layer into a probability distribution and convert it into the probability value of each category, so as to output the predicted classification result of the electrophoresis image. The predicted classification result includes positive and negative. After the first classification model is trained, the computer device can remove the Softmax layer in the first classification model, only retain the fully connected layer, and use the Logits value output by the fully connected layer as the positive / negative prediction value.
[0079] It can be understood that since the Logits value output by the fully connected layer of the neural network model retains more complete numerical information and can more sensitively reflect the confidence of the first classification model in the electrophoretic image belonging to a certain category, the first classification model can classify the negative, positive, and suspected positive of the target electrophoretic image according to the Logits value.
[0080] During the training process of the first classification model based on the first data set, the first classification model can perform binary classification of the electrophoretic image as negative or positive according to the Softmax layer, so that the model parameters of the first classification model can be adjusted based on the predicted classification results and the actual overall negative / positive labels to improve the classification accuracy of the first classification model. After the first classification model is trained, the Softmax layer of the first classification model can be removed, and the Logits value of the fully connected layer can be directly used as the output, so that the first classification model can classify the negative, positive, and suspected positive of the target electrophoretic image according to the Logits value.
[0081] By removing the Softmax layer in the first classification model after the first classification model is trained and directly using the Logits value output by the fully connected layer as the negative / positive prediction value, not only can the complexity of the first classification model be effectively reduced, but also the negative, positive, and suspected positive can be classified according to the Logits value, so as to make full use of the continuity and confidence information of the Logits value to improve the accuracy and reliability of the first classification result of the first classification model.
[0082] In some embodiments, the computer device can compare the negative / positive prediction value with a preset negative / positive prediction threshold interval to determine the first classification result corresponding to the negative / positive prediction threshold interval where the negative / positive prediction value is located.
[0083] If the negative / positive prediction value is greater than the first prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoretic image is positive; if the negative / positive prediction value is less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoretic image is negative; if the negative / positive prediction value is not greater than the first prediction threshold and not less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoretic image is suspected positive. Among them, the first prediction threshold and the second prediction threshold can be set as needed according to the actual situation.
[0084] Exemplarily, the first prediction threshold is set to 0.9 and the second prediction threshold is set to 0.1. When the computer device detects that the positive-negative prediction value is greater than 0.9, it can determine that the first classification result corresponding to the target electrophoresis image is positive; when the computer device detects that the positive-negative prediction value is less than 0.1, it can determine that the first classification result corresponding to the target electrophoresis image is negative; when the computer device detects that the positive-negative prediction value is not greater than 0.9 and not less than 0.1, it determines that the first classification result corresponding to the target electrophoresis image is suspected positive.
[0085] By using the positive-negative prediction value output by the first classification model to determine the corresponding first classification result, the confidence information of the positive-negative prediction value can be fully utilized to achieve accurate classification of the first classification result corresponding to the target electrophoresis image into negative, positive, and suspected positive, so as to improve the accuracy and reliability of the first classification result.
[0086] Step 204, the computer device determines whether the first classification result is suspected positive. If the first classification result is suspected positive, step 206 is executed; if the first classification result is not suspected positive, step 208 is executed.
[0087] The computer device can compare the positive-negative prediction value with the interval composed of the second prediction threshold and the first prediction threshold to determine whether the positive-negative prediction value is within this interval. If the positive-negative prediction value is within this interval, it is determined that the first classification result is suspected positive; if the positive-negative prediction value is not within this interval, it is determined that the first classification result is not suspected positive.
[0088] When the computer device determines that the first classification result corresponding to the target electrophoresis image is suspected positive, the computer device can further analyze the target electrophoresis image according to the second classification model to determine the positive-negative of the target electrophoresis image; when the computer device determines that the first classification result corresponding to the target electrophoresis image is not suspected positive, the computer device can directly determine the positive-negative of the target electrophoresis image according to the positive-negative prediction value.
[0089] Step 206, the computer device analyzes the target electrophoresis image through the second classification model to obtain the second classification result corresponding to the target electrophoresis image.
[0090] The second classification model is trained according to a suspected data set, which includes multiple second sample electrophoresis images labeled as suspected data and the overall positive-negative labels corresponding to each second sample electrophoresis image.
[0091] Suspected data may include electrophoresis images classified as suspected positive. The suspected data set may include second sample electrophoresis images with a first classification result of suspected positive and second sample electrophoresis images of suspected positive screened by expert experience.
[0092] In some embodiments, the computer device may construct a second classification model through a supervised contrastive learning algorithm and train the second classification model with a suspected data set, so that the second classification model can obtain the features of the target electrophoresis image with the first classification result being suspected positive, and thus determine the positive or negative of the target electrophoresis image according to the features.
[0093] The supervised contrastive learning algorithm may refer to an algorithm that uses the label information of electrophoresis images to construct positive and negative sample pairs and optimizes the feature learning ability of the second classification model by comparing the similarities between positive and negative sample pairs. Through the second classification model constructed based on the supervised contrastive learning algorithm, the computer device can use the clear overall positive or negative labels of each second sample electrophoresis image in the second data set, making two second sample electrophoresis images with the same overall positive or negative label (i.e., positive sample pair) closer in the feature space, while two second sample electrophoresis images with different overall positive or negative labels (i.e., negative sample pair) are farther away in the feature space, thereby improving the discrimination ability of the second classification model for different categories of second sample electrophoresis images.
[0094] In some embodiments, the second classification model includes a feature extractor and a fully connected layer. The feature extractor can be used to extract the features of the target electrophoresis image, and the fully connected layer can be used to map the features extracted by the feature extractor to class prediction to obtain the probability that the target electrophoresis image belongs to positive or negative, so as to determine the second classification result of the target electrophoresis image. The second classification result includes positive and negative.
[0095] The computer device can extract the features of the target electrophoresis image with the first classification result being suspected positive through the feature extractor, and then obtain the second classification result of the target electrophoresis image through the fully connected layer according to the features extracted by the feature extractor.
[0096] Step 208, the computer device determines whether the first classification result or the second classification result is positive. If the first classification result is positive or the second classification result is positive, step 210 is executed; if the first classification result is negative or the second classification result is negative, step 212 is executed.
[0097] When the first classification result or the second classification result is positive, that is, when the target electrophoresis image shows positive, the computer device can classify the positive or negative of each lane of the target electrophoresis image through a third classification model; when the first classification result or the second classification result is negative, that is, when the target electrophoresis image shows negative, the computer device can directly mark the overall positive or negative label of the target electrophoresis image as negative.
[0098] Step 210, the computer device identifies each lane included in the target electrophoresis image through the third classification model to obtain the third classification result corresponding to each lane.
[0099] The third classification model is trained based on a second data set, which includes multiple third sample electrophoresis images and the positive and negative labels corresponding to each lane in each third sample electrophoresis image. The positive and negative labels corresponding to each lane can be used to identify the positivity and negativity presented by each lane in the third sample electrophoresis image. Among them, the reference lane ELP in the third sample electrophoresis image does not have a positive and negative label, and the positive and negative labels of the lanes with the same band region as the reference lane ELP are positive, and the positive and negative labels of the lanes without the same band region or without a band region as the reference lane ELP are negative.
[0100] Optionally, the second data set further includes the overall positive and negative labels corresponding to each third sample electrophoresis image, and the overall positive and negative labels corresponding to each third sample electrophoresis image are positive.
[0101] In some embodiments, the computer device can identify each lane included in the target electrophoresis image through the third classification model and obtain the features of the band regions corresponding to each lane. Each lane included in the target electrophoresis image can include a reference lane and other lanes. The other lanes include a heavy chain lane and a light chain lane. Optionally, the heavy chain lane can include lane G, lane A, and lane M, and the light chain lane can include lane κ and lane λ; the computer device can compare the features of the band region of the reference lane with the features of the band regions corresponding to the other lanes to determine the third classification results corresponding to the other lanes.
[0102] The features of the band region can include position features and morphological features. The computer device can determine the overlapping degree between the band regions in each other lane and the band regions in the reference lane according to the position features and morphological features of each band region of the reference lane and the band regions corresponding to the other lanes in the target electrophoresis image, and determine the third classification results of each other lane according to the overlapping degree.
[0103] The overlapping degree can refer to the degree of mutual coincidence between the band regions in the other lanes and the band regions in the reference lane in the target electrophoresis image. The third classification results can include positive and negative. The higher the overlapping degree, the higher the probability that the other lanes in the target electrophoresis image show positive; the lower the overlapping degree, the higher the probability that the other lanes in the target electrophoresis image show negative.
[0104] In some embodiments, the computer device can first segment the target electrophoresis image to obtain multiple lane sub-images, and then identify and extract the features corresponding to the lanes included in each lane sub-image through the third classification model, or the computer device can directly identify and extract the features corresponding to each lane included in the target electrophoresis image through the third classification model. The computer device can determine the third classification results corresponding to each lane according to the features corresponding to each lane.
[0105] Step 212: The computer device determines the target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result, or the third classification result.
[0106] The target classification result corresponding to the target electrophoresis image includes the positive and negative of the target electrophoresis image, as well as the positive and negative corresponding to each lane included in the target electrophoresis image. Since when the first classification result of the target electrophoresis image is suspected positive, the computer device can further determine the positive and negative of the target electrophoresis image suspected positive through the second classification model, the target classification result does not include suspected positive.
[0107] In some embodiments, the computer device can mark the overall positive and negative label of the target electrophoresis image and the positive and negative labels corresponding to each lane included in the target electrophoresis image according to the first classification result, the second classification result, or the third classification result, and complete the recognition of the target electrophoresis image.
[0108] If the first classification result is suspected positive, the computer device obtains the second classification result of the target electrophoresis image according to the second classification model. If the first classification result or the second classification result is negative, the computer device can mark the overall positive and negative label of the target electrophoresis image as negative, and the positive and negative labels corresponding to each lane included in the target electrophoresis image as negative, and determine the target classification result corresponding to the target electrophoresis image; if the first classification result or the second classification result is positive, the computer device can mark the overall positive and negative label of the target electrophoresis image as positive, and obtain the third classification result corresponding to the target electrophoresis image according to the third classification model, so as to mark the positive and negative corresponding to each lane in the target electrophoresis image according to the third classification result, thereby determining the target classification result corresponding to the target electrophoresis image.
[0109] Through the joint prediction and classification of the three models of the first classification model, the second classification model, and the third classification model, not only can in-depth analysis be performed on the target electrophoresis image suspected positive that is difficult to qualitatively due to the fuzzy boundary of positive and negative, so as to more accurately distinguish the overall positive and negative of the target electrophoresis image, improve the accuracy and credibility of the classification of the target electrophoresis image, but also the positive and negative of each lane included in the target electrophoresis image can be identified, thereby providing a more accurate and perfect basis for the classification result of the overall positive and negative of the target electrophoresis image.
[0110] In the embodiments of the present application, the computer device classifies the target electrophoresis image through the first classification model to obtain the first classification result corresponding to the target electrophoresis image, re-classifies the target electrophoresis image with the first classification result being suspected positive through the second classification model to obtain the second classification result corresponding to the target electrophoresis image, and classifies each lane included in the target electrophoresis image with the first classification result or the second classification result being positive through the third classification model to obtain the third classification result corresponding to each lane. Finally, according to the first classification result, the second classification result or the third classification result, the target classification result corresponding to the target electrophoresis image is determined, so that the computer device can further analyze the target electrophoresis image through the second classification model when the first classification result of the target electrophoresis image is suspected positive, realize the accurate classification of the target electrophoresis image suspected positive, and when the first classification result or the second classification result is positive, the third classification result corresponding to each lane in the target electrophoresis image is also obtained through the third classification model to further verify the accuracy of the classification of the target electrophoresis image from the classification results of the lanes, thereby improving the accuracy of the computer device in recognizing the electrophoresis image.
[0111] In some embodiments, the computer device can train the first classification model according to the first data set. Figure 3 The flowchart for training the to-be-trained first classification model provided by the embodiments of the present application is as Figure 3 shown, and may include the following steps:
[0112] Step 301, the computer device obtains the first data set.
[0113] The first data set includes multiple first sample electrophoresis images and the overall positive and negative labels corresponding to each first sample electrophoresis image. The first sample electrophoresis images can be obtained through a public electrophoresis image database or by performing IFE clinical tests. Optionally, the first sample electrophoresis images are preprocessed sample electrophoresis images. The preprocessing may include one or more of, but is not limited to, image cropping, image enhancement, and image denoising.
[0114] The overall positive and negative labels corresponding to each first sample electrophoresis image can be determined by multiple experts according to multiple-digit criteria. For example, for the same first sample electrophoresis image, expert A and expert B label its overall positive and negative label as positive, while expert C labels its overall positive and negative label as negative. Taking the labels of expert A and expert B as the standard, that is, the overall positive and negative label of this first sample electrophoresis image is positive.
[0115] Step 303, the computer device selects data from the first data set according to a certain proportion of each category of the overall positive and negative labels and divides the data into a training set and a test set.
[0116] In some embodiments, the computer device may select data from the first dataset according to a ratio of 4:1 of negative to positive for the overall negative and positive labels, and divide the selected data into a training set and a test set according to a ratio of 9:1.
[0117] It can be understood that during the collection of the first sample electrophoresis images, since the probability of the first sample electrophoresis images with positive overall negative and positive labels appearing in the actual IFE clinical detection is relatively low, and the first sample electrophoresis images with negative overall negative and positive labels are easier to obtain, it is easy to cause an extremely unbalanced ratio of overall negative and positive labels for multiple first sample electrophoresis images in the first dataset. This unbalanced data distribution may have a negative impact on the training of the first classification model, making the first classification model overly biased towards the majority class during the learning process and ignoring the features of the minority class, thereby reducing the classification accuracy of the first classification model for the minority class. Therefore, in order to ensure that the first classification model can fully learn the features of the first sample electrophoresis images with positive and negative overall negative and positive labels and improve the generalization ability and classification performance of the first classification model, it is necessary to select some data from the first dataset according to a certain ratio of each category of the overall negative and positive labels to balance the category distribution of the overall negative and positive labels in the training set.
[0118] Exemplarily, the computer device may select data from the first dataset according to a ratio of 4:1 between negative and positive for the overall negative and positive labels to construct a data subset, and then divide the training set and the test set according to a ratio of 9:1 within the data subset.
[0119] Step 305, the computer device performs K-fold cross-validation through the training set and the test set to complete the training of the first classification module.
[0120] K-fold cross-validation may refer to splitting the training set into K non-overlapping subsets, each time using a different subset as the validation set and the remaining K - 1 subsets as the training set for K times of model training and evaluation, where K is a positive integer. Optionally, K-fold cross-validation may perform K times of model training and evaluation to obtain the model parameters corresponding to K models.
[0121] In some embodiments, the computer device may use the cross-entropy loss function as the loss function during the training process of the first classification model, and perform K-fold cross-validation through the training set and the test set according to the first classification model to obtain the model parameters of K classification sub-models to complete the training of the first classification model.
[0122] The first classification model may include K classification sub-models, and each classification sub-model is constructed according to the model parameters of the first classification model that meet the convergence conditions each time in K-fold cross-validation. For example, the computer device performs five-fold cross-validation through the training set and the test set, determines the first classification sub-model in the first round of cross-validation, determines the second classification sub-model in the second round of cross-validation, and so on to determine five classification sub-models, and determines the first classification model according to the above five classification sub-models.
[0123] After determining the model parameters of the K classification sub-models according to the training set and determining the first classification model, the model performance of the first classification model is detected according to the test set to complete the training of the first classification module.
[0124] Since the computer device performs K-fold cross-validation through the training set and the test set to complete the training of the first classification module, which is similar to the prior art, it will not be elaborated here.
[0125] In some embodiments, the computer device may use the average value of the output values of the K classification sub-models as the positive and negative prediction value of the first classification model.
[0126] In the embodiments of the present application, the computer device selects data from the first data set according to a certain proportion of each category of the overall positive and negative labels, and performs K-fold cross-validation based on the selected data to complete the training of the first classification model. This can not only ensure that the first classification model fully learns the features of the positive and negative electrophoresis images during the training process, improve the recognition ability of the minority positive electrophoresis images, but also enable the model to be trained and verified multiple times on different training sets through the K-fold cross-validation method, enhance the generalization ability and stability of the first classification model, and further improve the accuracy and reliability of the first classification model for classifying target electrophoresis images.
[0127] In some embodiments, the computer device may train the third classification model according to the second data set. Figure 4 This is a flowchart for training the third classification model to be trained provided by the embodiments of the present application. As Figure 4 shown, it may include the following steps:
[0128] Step 402, the computer device obtains the second data set.
[0129] The second data set includes multiple third sample electrophoresis images, and the positive and negative labels corresponding to each lane in each third sample electrophoresis image. Among them, the above image acquisition method and the labeling of the labels are the same as those of the first sample electrophoresis image in step 301 and the labeling method of the overall positive and negative labels of the first sample electrophoresis image, and will not be elaborated here.
[0130] In some embodiments, the computer device may select a part of the electrophoresis images from the total dataset as the first dataset, and the remaining electrophoresis images as the third dataset, and use all the electrophoresis images with positive global positive and negative labels in the third dataset as the second dataset.
[0131] Since the second dataset belongs to the third dataset rather than the first dataset, using the second dataset to train the third classification model can avoid the leakage of test data or validation data into the training data, ensuring the accuracy and effectiveness when evaluating the third classification model.
[0132] In some embodiments, the computer device may select a specific number of third sample electrophoresis images from the third dataset such that the number of third sample electrophoresis images with positive positive and negative labels in different lanes is equal to balance the positive features in different lanes.
[0133] Step 404, the computer device divides the second dataset into a training set and a test set for K-fold cross-validation to complete the training of the third classification module.
[0134] The way the computer device divides the second dataset into a training set and a test set for K-fold cross-validation to complete the training of the third classification module is the same as the way the computer device completes the training of the first classification module through K-fold cross-validation using the training set and the test set in step 305, which will not be elaborated here.
[0135] In the embodiments of the present application, the computer device performs K-fold cross-validation through multiple third sample electrophoresis images with positive global positive and negative labels included in the second dataset to complete the training of the third classification model, which can enhance the generalization ability and stability of the third classification model, and further improve the accuracy and reliability of the third classification model for classifying the positive and negative of each lane included in the target electrophoresis image.
[0136] In some embodiments, the computer device may construct a second classification model based on the supervised contrastive learning algorithm and train the second classification model through the suspected dataset. Figure 5 This is a flowchart for training the second classification model to be trained provided by the embodiments of the present application. As Figure 5 shown, the following steps are further included:
[0137] Step 501, the computer device obtains multiple second sample electrophoresis images included in the current batch.
[0138] It can be understood that due to computer performance limitations, the computer device cannot process all the second sample electrophoresis images in the suspected dataset at one time, but needs to process them in batches.
[0139] Exemplarily, the computer device can set the batch size to 8 according to its own processing capacity, that is, the second classification model can process 8 images each time. Therefore, the computer device can obtain 8 second sample electrophoresis images included in the current batch.
[0140] Step 503, the computer device performs image enhancement processing on each second sample electrophoresis image in the current batch through two different image enhancement algorithms to obtain two enhanced training images corresponding to each second sample electrophoresis image.
[0141] Image enhancement can refer to transforming or adjusting the second sample electrophoresis image through various image processing techniques. The image enhancement algorithm can include one or more of, but is not limited to, size change, color jitter, adding Gaussian noise, and random cropping, etc.
[0142] Size change can refer to changing the size of the second sample electrophoresis image. Color jitter can refer to adjusting the color channels of the second sample electrophoresis image, including but not limited to adjusting one or more of the brightness, contrast, saturation, and hue of the second sample electrophoresis image. Adding Gaussian noise can refer to adding random noise that conforms to the Gaussian distribution to the second sample electrophoresis image to form noise points in the second sample electrophoresis image. Random cropping can refer to randomly selecting a sub-region in the second sample electrophoresis image for cropping to generate a corresponding enhanced training image.
[0143] Exemplarily, the computer device can first scale each second sample electrophoresis image in the current batch to a preset resolution, then add Gaussian noise within a specified variance range to the edge region of the image to obtain the first enhanced training image, and then adjust the color channels of each second sample electrophoresis image in the same batch to obtain the second enhanced training image.
[0144] In some embodiments, the two enhanced training images corresponding to one second sample electrophoresis image have the same overall positive and negative label as the second sample electrophoresis image.
[0145] In some embodiments, the computer device can randomly select two different image enhancement algorithms from multiple image enhancement algorithms to perform image enhancement processing on each second sample electrophoresis image in the current batch to obtain two enhanced training images corresponding to each second sample electrophoresis image.
[0146] Step 505, the computer device generates embedding vectors corresponding to the two enhanced training images respectively through the second classification model to be trained according to the two enhanced training images corresponding to each second sample electrophoresis image.
[0147] The embedded vector may refer to the representation of the enhanced training image converted by the second classification model to be trained into a low-dimensional vector space through a specific mapping method. The embedded vector can accurately and effectively represent the key features of the enhanced training image.
[0148] The computer device may input two enhanced training images corresponding to each second sample electrophoresis image into the second classification model to be trained, so that the second classification model to be trained can extract the features corresponding to the two enhanced training images to generate the embedded vectors corresponding to the two enhanced training images.
[0149] In some embodiments, the second classification model may include two identical and independent embedded vector acquisition modules. The first embedded vector acquisition module is used to acquire one of the enhanced training images in the second sample electrophoresis image, and the second embedded vector acquisition module is used to acquire the other enhanced training image in the second sample electrophoresis image.
[0150] It can be understood that through two independent embedded vector acquisition modules, accurate feature extraction can be performed on the two enhanced training images respectively, so that the second classification model can avoid mutual interference in the feature extraction process to independently learn the features of each enhanced training image, thereby providing a more accurate basis for calculating the similarity between the two enhanced training images in the subsequent steps, facilitating the distinction of positive multiple sample pairs in the contrast learning, and further improving the classification accuracy of the second classification model.
[0151] Exemplarily, the embedded vector acquisition module can be a multi-layer perceptron of a deep neural network, can also be a convolutional layer and a pooling layer of a convolutional neural network, and can also be a residual block of a residual neural network, but is not limited thereto.
[0152] Step 507, the computer device combines multiple enhanced training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs.
[0153] The sample pair may refer to a set composed of any two enhanced training images in the current batch. The sample pair may include positive sample pairs and negative sample pairs. The positive sample pair may include a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included in the sample pair are the same, and the negative sample pair may include a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included in the sample pair are different.
[0154] Since in the supervised contrast learning, it is necessary to adjust the model parameters of the second classification model by comparing the similarities and differences between different enhanced training images, it is necessary to combine multiple enhanced training images included in the current batch in pairs to obtain multiple groups of sample pairs, and adjust the model parameters of the second classification model through the similarities and differences of each group of sample pairs.
[0155] In some embodiments, the computer device may convert the corresponding embedding vectors of multiple enhanced training images included in the current batch into tensors and then perform tensor concatenation to obtain multiple sets of sample pairs that conform to the embedding vector sorting.
[0156] As Figure 6 shown, the step of pairwise combining multiple enhanced training images corresponding to the current batch to obtain multiple sets of sample pairs may include the following steps:
[0157] Step 602, the computer device converts the embedding vectors respectively corresponding to the multiple enhanced training images of the current batch into a first tensor.
[0158] The first tensor may refer to the representation form of the embedding vector corresponding to the enhanced training image in three-dimensional space.
[0159] In some embodiments, the computer device may insert a dimension of size 1 into the embedding vector corresponding to the enhanced training image through the unsqueeze instruction to convert it into a first tensor. The unsqueeze instruction can be used to insert a dimension of size 1 in the specified dimension to change the shape of the tensor.
[0160] Exemplarily, the batch size of the current batch is 8, and the dimension of the embedding vector corresponding to the enhanced training image is 512. The computer device can convert the embedding vector into a first tensor of [8, 1, 512] through the unsqueeze instruction.
[0161] In some embodiments, before converting the embedding vector into a tensor, the computer device first normalizes the embedding vectors of each enhanced training image and converts the normalized embedding vectors into a first tensor.
[0162] Step 604, the computer device pairwise combines and concatenates the first tensors respectively corresponding to multiple enhanced training images to obtain second tensors respectively corresponding to multiple sets of sample pairs.
[0163] The shape of the second tensor is a two-dimensional tensor. The dimension of the first dimension of the second tensor is twice the batch size, where the batch size may refer to the total number of second sample electrophoresis images included in the current batch. The dimension of the second dimension of the second tensor is the dimension of the embedding vector.
[0164] In some embodiments, the computer device may concatenate the first tensors respectively corresponding to two enhanced training images along the depth axis direction of the three-dimensional space to obtain a third tensor, and then split the third tensor along the depth axis direction of the three-dimensional space and re-concatenate it along the length axis direction of the three-dimensional space to obtain the second tensor.
[0165] Figure 7Schematic diagram of obtaining a second tensor by splicing the first tensors corresponding to two augmented training images provided by an embodiment of the present application. As Figure 7 shown, assuming that the batch size is 3 and the dimensionality feature of the embedding vector is 4, the computer device can use the unsqueeze instruction to convert the embedding vectors corresponding to two augmented training images into the first tensor 701 and the first tensor 702 with the shape of [3, 1, 4]. Subsequently, the first tensor 701 and the first tensor 702 are spliced along the depth axis Z of the three-dimensional space to obtain the third tensor 703, and then split along the depth axis Z of the three-dimensional space and re-spliced along the length axis X of the three-dimensional space to obtain the second tensor 704.
[0166] Since the dimension of the first dimension of the second tensor is the total number of augmented training images included in the current batch, and the dimension of the second dimension of the second tensor is the dimension of the embedding vector, that is, each row of the second tensor represents the embedding vector corresponding to an augmented training image. Therefore, the computer device can perform pairwise combination according to each row element of the second tensor and each column element of the transpose of the second tensor to obtain multiple groups of sample pairs.
[0167] By converting the embedding vectors corresponding to multiple augmented training images in the current batch into the first tensors, and then pairwise combining and splicing the first tensors corresponding to multiple augmented training images respectively to obtain the second tensors corresponding to multiple groups of sample pairs, not only can the embedding vectors of all augmented training images in the current batch be recorded through the second tensor, but also the sorting of multiple groups of sample pairs can conform to the embedding vector sorting, which is convenient for the computer device to directly perform unified feature comparison and learning based on the second tensor, improving the calculation efficiency.
[0168] Step 509, the computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two augmented training images included in each group of sample pairs.
[0169] The similarity of a sample pair can refer to the similarity of features between the two augmented training images in the sample pair. In supervised contrastive learning, the similarity of a sample pair can reflect the proximity degree of the sample pair in the feature space, and is used to guide the loss function of the second classification model to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.
[0170] In some embodiments, the computer device can calculate the similarity of each group of sample pairs according to the dot product of the embedding vectors corresponding to the two augmented training images included in each group of sample pairs and the temperature coefficient.
[0171] Figure 8 Flowchart of calculating the similarity of each group of sample pairs provided by an embodiment of the present application. As Figure 8As shown, the step of calculating the similarity of each sample pair according to the embedding vectors respectively corresponding to the two enhanced training images included in each sample pair may include the following steps:
[0172] Step 802, the computer device performs a dot product calculation according to the embedding vectors respectively corresponding to the two enhanced training images included in the target sample pair, and obtains the initial similarity corresponding to the target sample pair.
[0173] The target sample pair is any sample pair in the current batch. The initial similarity of the target sample pair may refer to the cosine similarity of the embedding vectors respectively corresponding to the two enhanced training images.
[0174] In some embodiments, the computer device normalizes the embedding vector corresponding to each enhanced training image, and performs a dot product calculation using the normalized embedding vector to obtain the initial similarity corresponding to the target sample pair.
[0175] Normalizing the embedding vector corresponding to each enhanced training image may include independently performing L2 normalization on the embedding vector of an enhanced training image. L2 normalization may refer to dividing each element in the embedding vector by the L2 norm of the embedding vector so that the norm of the embedding vector is 1.
[0176] It can be understood that in supervised contrastive learning, since only the relative similarity of the embedding vectors respectively corresponding to the two enhanced training images in the target sample pair is concerned, rather than the vector magnitude, the embedding vector can be normalized to highlight its similarity. And, after performing L2 normalization on the embedding vector corresponding to the enhanced training image, the norm of the embedding vector is 1, so that the dot product of the embedding vectors respectively corresponding to the two enhanced training images in the target sample pair can be equivalent to the cosine similarity.
[0177] Exemplarily, the target sample pair includes the enhanced training image E and the enhanced training image F, and the normalized embedding vector of the enhanced training image E is A = [a1, a2, …, a N , and the normalized embedding vector of the enhanced training image F is B = [b1, b2, …, b N , where N is the dimension of the embedding vector. Therefore, the computer device can perform a dot product calculation through the normalized embedding vector A and the normalized embedding vector B to obtain the initial similarity between the enhanced training image E and the enhanced training image F as:
[0178] C1 = A · B = a1 · b1 + a2 · b2 + … + a N · b N Equation (1);
[0179] where C1 represents the initial similarity between the enhanced training image E and the enhanced training image F.
[0180] Step 804: The computer device adjusts the initial similarity according to the temperature coefficient to obtain the similarity corresponding to the target sample pair.
[0181] The temperature coefficient can be used to adjust the smoothness of the similarity distribution, thereby affecting the second classification model's focus on the positive sample pairs and negative sample pairs in multiple groups of sample pairs. The temperature coefficient can be set as needed according to actual conditions.
[0182] In some embodiments, the computer device may determine the similarity corresponding to the target sample pair based on the ratio of the initial similarity and the temperature coefficient.
[0183] Exemplarily, the computer device may determine the similarity corresponding to the target sample pair by the ratio of the initial similarity C1 and the temperature coefficient τ:
[0184]
[0185] Among them, C2 represents the similarity of the target sample pair, that is, the similarity between the enhanced training image E and the enhanced training image F.
[0186] The initial similarity is determined by the dot product of the embedding vectors corresponding to the two enhanced training images contained in the target sample pair, and then the initial similarity is adjusted according to the temperature coefficient to determine the similarity of the target sample pair. This can better highlight the similarities between the positive sample pairs and at the same time widen the differences between the negative sample pairs, which helps the second classification model to more accurately learn the feature differences between different enhanced training images during the training process, thereby enhancing the classification ability and generalization performance of the second classification model.
[0187] In some embodiments, after the computer device calculates the similarity of each group of sample pairs based on the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs, the computer device may output a similarity matrix through a second classification model.
[0188] The similarity matrix may include the similarities of each sample pair included in the current batch. The row index of the similarity matrix may be the order of one enhanced training image in a sample pair among multiple enhanced training images included in the current batch, and the column index of the similarity matrix may be the order of another enhanced training image in the sample pair among multiple enhanced training images included in the current batch.
[0189] In some embodiments, after the computer device outputs the similarity matrix through the second classification model, the similarity matrix may be normalized to generate a target similarity matrix.
[0190] The normalization process may include normalizing the similarity matrix. The target similarity matrix may include the logarithmic probabilities of each pair of sample groups included in the current batch. The logarithmic probability may refer to the probability that each pair of sample groups belongs to the positive samples. The sum of the logarithmic probabilities in each row of the target similarity matrix is 1.
[0191] In some embodiments, the computer device may set each similarity on the main diagonal of the similarity matrix to zero to generate an initial similarity matrix, then perform an exponential operation on the similarities of the initial similarity matrix, accumulate the results, and calculate the sum of each row. By taking the logarithm of the sum of each row to generate the logarithmic sum, and using the difference between each similarity of the initial similarity matrix and the logarithmic sum of the corresponding row of the similarity, the normalization process of the similarity matrix is completed to determine the target similarity matrix.
[0192] By normalizing the similarity matrix, it can ensure that the similarity of each pair of samples corresponding to an enhanced training image obtains a reasonable logarithmic probability. It can not only avoid directly calculating the exponential sum in the Softmax function, avoid numerical overflow or underflow of the similarity caused by the exponential operation, but also improve the numerical stability of the similarity matrix.
[0193] Step 511, the computer device calculates the loss value according to the similarities of multiple groups of sample pairs.
[0194] In some embodiments, the computer device may calculate the loss value according to the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs.
[0195] Figure 9 This is a flowchart for calculating the loss value according to the similarities of multiple groups of sample pairs provided by the embodiments of the present application. As Figure 9 shown, it may include the following steps:
[0196] Step 901, the computer device obtains the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs respectively.
[0197] Step 903, the computer device generates a mask according to the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs respectively.
[0198] The mask may refer to a binary matrix used to selectively focus on whether the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs are the same. The dimension of the mask is the total number of enhanced training images included in the current batch. The row index of the elements in the mask may be the sorting of one of the enhanced training images in a group of sample pairs among the multiple enhanced training images included in the current batch, and the column index of the elements in the mask may be the sorting of the other enhanced training image in this sample pair among the multiple enhanced training images included in the current batch. That is, each element in the mask corresponds to multiple groups of sample pairs included in the current batch one by one.
[0199] Optionally, the mask can be a 0-1 matrix. The elements with a value of 1 in the mask can indicate that the overall positive and negative labels corresponding to the two enhanced training images included in the sample pair are the same, while the elements with a value of 0 can indicate that the overall positive and negative labels corresponding to the two enhanced training images included in the sample pair are different.
[0200] The computer device can detect whether the overall positive and negative labels corresponding to the two enhanced training images included in each sample pair are the same. If the overall positive and negative labels corresponding to the two enhanced training images included in each sample pair are the same, the computer device classifies the sample pair as a positive sample pair. If the overall positive and negative labels corresponding to the two enhanced training images included in each sample pair are different, the computer device classifies the sample pair as a negative sample pair.
[0201] It can be understood that the elements on the main diagonal of the mask represent the sample pairs formed by the enhanced training image and itself. Therefore, all elements on the main diagonal of the mask are positive sample pairs.
[0202] In some embodiments, the computer device can generate an initial mask based on the dimension of the first dimension of the second tensor, and mark the corresponding positions on the initial mask according to the overall positive and negative labels corresponding to the two enhanced training images included in each sample pair to generate a mask.
[0203] Exemplarily, the dimension of the first dimension of the second tensor, that is, the total number of enhanced training images included in the current batch, is 8. The computer device can generate an initial mask with a size of [8, 8] based on this dimension, and mark the corresponding positions on the initial mask according to the overall positive and negative labels corresponding to the two enhanced training images included in each sample pair to generate a mask with a size of [8, 8].
[0204] In some embodiments, the mask can include a first mask and a second mask. The first mask can include a mask that only identifies positive sample pairs, and the second mask can include a mask that only identifies negative sample pairs. The computer device can generate an initial mask based on the dimension of the first dimension of the second tensor, and mark the corresponding positions on the initial mask according to each positive sample pair to generate the first mask, and mark the corresponding positions on the initial mask according to each negative sample pair to generate the second mask.
[0205] Step 905, the computer device determines the similarity of each positive sample pair and each negative sample pair included in the current batch according to the mask and the similarity of each sample pair.
[0206] The computer device can distinguish positive sample pairs and negative sample pairs in multiple groups of sample pairs included in the current batch according to the mask, so as to extract the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs in the similarity matrix output by the second classification model.
[0207] In some embodiments, since the dimensions of the mask and the similarity matrix output by the second classification model are the same, the computer device can obtain the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs included in the current batch according to the product of the mask and the similarity matrix.
[0208] In some embodiments, the mask may include a first mask and a second mask. The computer device can obtain the similarity of each group of positive sample pairs included in the current batch according to the product of the first mask and the similarity matrix, and obtain the similarity of each group of negative sample pairs included in the current batch according to the product of the second mask and the similarity matrix.
[0209] Step 907, the computer device calculates the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs.
[0210] The first positive sample pair is any positive sample pair.
[0211] In supervised contrastive learning, through the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs, the similarity of the positive sample pairs can be maximized while the similarity of the negative sample pairs is minimized, so that the loss value of the second classification model is minimized during training, enabling the second classification model to better learn the features of the positive sample pairs, so that the second classification model can accurately identify the positive and negative features in the electrophoresis image, and further realize the accurate positive and negative classification of the suspected positive target electrophoresis image by the second classification model.
[0212] In some embodiments, the computer device can first perform exponential processing on the similarities of multiple groups of sample pairs to obtain the exponential similarity of each group of sample pairs, and then perform logarithmic processing on the ratio of the exponential similarity of the first positive sample pair to the exponential similarity of each group of negative sample pairs to calculate the loss value.
[0213] It can be understood that by exponentiating the similarities of multiple groups of sample pairs, the differences in the similarities of multiple groups of sample pairs can be amplified to better highlight the positive sample pairs; then, by performing logarithmic processing on the ratio of the exponential similarity of the first positive sample pair to the exponential similarity of each group of negative sample pairs, the second classification model can convert the similarity into a probability value without the need for a Softmax layer, and then calculate the loss value through this probability value.
[0214] Step 909, the computer device calculates the average value according to the ratio corresponding to each group of positive sample pairs, and calculates the loss value according to the average value.
[0215] In some embodiments, the computer device may sequentially determine the ratio corresponding to each group of positive sample pairs of each enhanced training image according to the total number of enhanced training images included in the current batch, calculate the average value of the ratios corresponding to each group of positive sample pairs of multiple enhanced training images included in the current batch, and take the negative of the average value to obtain a loss value.
[0216] Exemplarily, the computer device may determine the loss value as:
[0217]
[0218] Wherein, represents the loss value, N represents the total number of enhanced training images included in the current batch, and z i represents the embedding vector of the i-th enhanced training image included in the current batch, represents the embedding vector corresponding to the a-th enhanced training image and the i-th enhanced training image in the current batch when they are a positive sample pair, represents the embedding vector corresponding to the a-th enhanced training image and the i-th enhanced training image in the current batch when they are a negative sample pair.
[0219] The computer device generates a mask according to the overall positive and negative labels respectively corresponding to the two enhanced training images included in each group of sample pairs, obtains the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs included in the current batch based on the mask, then obtains the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs, and calculates the loss value according to the average value of the ratios corresponding to each group of positive sample pairs, so that the second classification model can maximize the similarity of positive sample pairs while minimizing the similarity of negative sample pairs. The second classification model can better learn the features of positive sample pairs, so that the second classification model can accurately identify positive and negative features in the electrophoresis image, and further realize accurate positive and negative classification of the suspected positive target electrophoresis image by the second classification model.
[0220] Step 513, the computer device determines whether the loss value of the second classification model satisfies the convergence condition. If it does not satisfy the convergence condition, step 515 is executed; if it satisfies the convergence condition, step 517 is executed.
[0221] In some embodiments, the computer device may compare the loss value of the second classification model in the current batch with a preset loss threshold to determine whether the loss value of the second classification model satisfies the convergence condition. Optionally, the convergence condition may further include that the number of batches of the second classification model meets the preset number of batches, and / or the number of iterations of the second classification model meets the preset number of iterations. The preset number of batches and the preset number of iterations can be set as required according to the actual situation.
[0222] Step 515: The computer device adjusts the parameters of the second classification model to be trained according to the loss value, and uses multiple second-sample electrophoresis images included in the next batch as multiple second-sample electrophoresis images included in the new current batch.
[0223] The computer device can adjust the parameters of the second classification model to be trained according to the loss value, and use multiple second-sample electrophoresis images included in the next batch as multiple second-sample electrophoresis images included in the new current batch, and then execute steps 503 to 513 again until the convergence condition is met.
[0224] Step 517: The computer device obtains the trained second classification model.
[0225] Since the process of training the second classification model in steps 513 to 517 is similar to the prior art, it will not be elaborated here again.
[0226] In the embodiments of the present application, the computer device calculates the similarity of multiple pairs of samples formed by pairwise combination of two enhanced training images through the embedding vector corresponding to each enhanced training image, and then calculates the loss value through the similarity of multiple pairs of samples, and adjusts the parameters of the second classification model to be trained according to the loss value, completing the training of the second classification model, which can enable the second classification model to more accurately learn the feature differences between multiple pairs of samples, thereby improving the accuracy and reliability of the second classification model in classifying suspected positive target electrophoresis images.
[0227] In some embodiments, the second classification model includes a feature extractor and a fully connected layer. The feature extractor can be used to extract the features of the target electrophoresis image, and the fully connected layer can be used to determine the second classification result of the target electrophoresis image. Figure 10 Another flowchart for training the second classification model to be trained provided by the embodiments of the present application. As Figure 10 shown, it further includes the following steps:
[0228] Step 1002: The computer device obtains multiple second-sample electrophoresis images included in the current batch.
[0229] Step 1004: The computer device performs image enhancement processing on each second-sample electrophoresis image in the current batch through two different image enhancement algorithms to obtain two enhanced training images corresponding to each second-sample electrophoresis image.
[0230] Step 1006: The computer device extracts the features of the two enhanced training images corresponding to each second-sample electrophoresis image through the feature extractor to be trained, and obtains the embedding vectors corresponding to the two enhanced training images respectively.
[0231] In some embodiments, the feature extractor may include two identical and independent embedding vector acquisition modules. The first embedding vector acquisition module is used to acquire the embedding vector of an enhanced training image in the second sample electrophoresis image, and the second embedding vector acquisition module is used to acquire the embedding vector of another enhanced training image in the second sample electrophoresis image.
[0232] Step 1008, the computer device combines multiple enhanced training images included in the current batch in pairs to obtain multiple groups of sample pairs.
[0233] Step 1010, the computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs.
[0234] Step 1012, the computer device calculates the loss value according to the similarities of multiple groups of sample pairs.
[0235] Step 1014, the computer device determines whether the loss value of the feature extractor meets the convergence condition. If it does not meet the convergence condition, it executes step 1016. If it meets the convergence condition, it executes step 1018.
[0236] Step 1016, the computer device adjusts the parameters of the feature extractor to be trained according to the loss value, and uses multiple second sample electrophoresis images included in the next batch as multiple second sample electrophoresis images included in the new current batch.
[0237] Step 1018, the computer device obtains the trained feature extractor.
[0238] The method of training the feature extractor in the above steps is similar to the method of training the second classification model in the above steps 501 to 517, and will not be elaborated here.
[0239] Step 1020, when the feature extractor is trained, the computer device generates the feature vector corresponding to the input second sample electrophoresis image through the trained feature extractor, and determines the predicted classification result corresponding to the input second sample electrophoresis image through the fully connected layer to be trained according to the feature vector corresponding to the input second sample electrophoresis image.
[0240] The fully connected layer can be used to map the feature vector extracted by the feature extractor into a predicted classification probability, so as to determine the predicted classification result corresponding to the input second sample electrophoresis image through the predicted classification probability.
[0241] During the training process of the feature extractor, the feature extractor needs to obtain the embedding vectors of each enhanced training image to calculate their similarity for contrastive learning. After the feature extractor is trained, it can accurately extract the features of the input second sample electrophoresis image. Therefore, it is no longer necessary to generate the embedding vectors of the enhanced training images, but can directly generate the feature vectors of the input second sample electrophoresis image.
[0242] In some embodiments, when the feature extractor is trained, the computer device can add a fully connected layer to be trained at the tail of the feature extractor, generate the feature vectors corresponding to the second sample electrophoresis images input in the current batch through the trained feature extractor, and determine the predicted classification results corresponding to the second sample electrophoresis images input in the current batch according to the feature vectors corresponding to the second sample electrophoresis images input in the current batch through the fully connected layer to be trained.
[0243] In some embodiments, after the computer device trains the feature extractor through the suspected data set, it trains the fully connected layer through the same suspected data set.
[0244] Step 1022, the computer device determines whether the loss value of the fully connected layer meets the convergence condition. If it does not meet the convergence condition, step 1024 is executed. If it meets the convergence condition, step 1026 is executed.
[0245] In some embodiments, the computer device can calculate the loss value of the fully connected layer according to the predicted classification results corresponding to the second sample electrophoresis images included in the current batch and the corresponding overall positive and negative labels, and the output value of the second classification model to be trained.
[0246] Step 1024, the computer device adjusts the parameters of the fully connected layer according to the predicted classification results corresponding to the input second sample electrophoresis images and the corresponding overall positive and negative labels, and uses the multiple second sample electrophoresis images included in the next batch as the multiple second sample electrophoresis images included in the new current batch.
[0247] In some embodiments, the computer device can calculate the loss value of the fully connected layer according to the predicted classification results corresponding to the second sample electrophoresis images included in the current batch and the corresponding overall positive and negative labels, and the output value of the second classification model to be trained, adjust the parameters of the fully connected layer according to the loss value of the fully connected layer, and use the multiple second sample electrophoresis images included in the next batch as the multiple second sample electrophoresis images included in the new current batch until the fully connected layer to be trained meets the convergence condition.
[0248] Since the calculation of the loss value based on the predicted classification result corresponding to the input second sample electrophoresis image and the corresponding overall positive / negative label, as well as the output value of the second classification model to be trained, is similar to the prior art, it will not be elaborated here.
[0249] Step 1026, the computer device obtains the trained second classification model.
[0250] By training the feature extractor and the fully connected layer in stages, the feature extractor can focus more on learning image features to accurately extract the features of the second sample electrophoresis image. The fully connected layer starts training after the feature extractor is trained. Since the feature extractor can already accurately extract the feature vectors of the second sample electrophoresis image, the fully connected layer can focus more on learning the mapping relationship to determine the predicted classification result.
[0251] In the embodiment of the present application, the computer device completes the training of the second classification model by first training the feature extractor of the second classification model and then training the fully connected layer of the second classification model. This can not only reduce the training complexity of the second classification model through staged training, avoid difficulties in parameter optimization, improve the training efficiency, enable the feature extractor to more accurately and stably extract the features of the electrophoresis image, and enable the fully connected layer to more accurately classify the electrophoresis image, but also effectively avoid overfitting, enhance the generalization ability of the second classification model, and ensure the accurate classification of the suspected positive target electrophoresis image by the second classification model.
[0252] Based on the electrophoresis image recognition method provided in the above embodiment, Figure 11 This is a structural block diagram of an electrophoresis image recognition device provided in an embodiment of the present application. As Figure 11 shown, in one embodiment, an electrophoresis image recognition device 1100 is provided. The electrophoresis image recognition device 1100 can be applied to an electronic device. The electrophoresis image recognition device 1100 includes a first classification module 1101, a second classification module 1102, a third classification module 1103, and an image recognition module 1104.
[0253] The first classification module 1101 is configured to recognize a target electrophoresis image through a first classification model, obtain a positive / negative predicted value corresponding to the target electrophoresis image, and determine a first classification result corresponding to the target electrophoresis image according to the positive / negative predicted value; the first classification model is trained according to a first data set, and the first data set includes multiple first sample electrophoresis images and the corresponding overall positive / negative labels for each first sample electrophoresis image.
[0254] The second classification module 1102 is configured to, if the first classification result is suspected positive, analyze the target electrophoresis image through a second classification model to obtain a second classification result corresponding to the target electrophoresis image; the second classification model is trained according to a suspected data set, and the suspected data set includes multiple second sample electrophoresis images labeled as suspected data, and the overall positive and negative labels corresponding to each second sample electrophoresis image.
[0255] The third classification module 1103 is configured to, if the first classification result or the second classification result is positive, identify each lane included in the target electrophoresis image through a third classification model to obtain a third classification result corresponding to each lane; the third classification model is trained according to a second data set, and the second data set includes multiple third sample electrophoresis images, and the positive and negative labels corresponding to each lane in each third sample electrophoresis image.
[0256] The image recognition module 1104 is configured to determine a target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result or the third classification result.
[0257] In some embodiments, the electrophoresis image recognition device 1100 includes a threshold detection module.
[0258] The threshold detection module is configured to, if the positive and negative prediction value is greater than a first prediction threshold, determine that the first classification result corresponding to the target electrophoresis image is positive.
[0259] The threshold detection module is further configured to, if the positive and negative prediction value is less than a second prediction threshold, determine that the first classification result corresponding to the target electrophoresis image is negative.
[0260] The threshold detection module is further configured to, if the positive and negative prediction value is not greater than the first prediction threshold and not less than the second prediction threshold, determine that the first classification result corresponding to the target electrophoresis image is suspected positive.
[0261] In some embodiments, the electrophoresis image recognition device 1100 includes an image acquisition module, an image enhancement module, a feature acquisition module, a sample pair acquisition module, a calculation module, and a model training module.
[0262] The image acquisition module is configured to acquire multiple second sample electrophoresis images included in the current batch.
[0263] The image enhancement module is configured to perform image enhancement processing on each second sample electrophoresis image in the current batch through two different image enhancement algorithms to obtain two enhanced training images corresponding to each second sample electrophoresis image.
[0264] The feature acquisition module is configured to generate embedding vectors corresponding to the two enhanced training images respectively through a second classification model to be trained according to the two enhanced training images corresponding to each second sample electrophoresis image.
[0265] A sample pair acquisition module, configured to combine multiple enhanced training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs.
[0266] A calculation module, configured to calculate the similarity of each group of sample pairs according to the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs.
[0267] The calculation module is further configured to calculate a loss value according to the similarities of multiple groups of sample pairs.
[0268] A model training module, configured to adjust the parameters of the second classification model to be trained according to the loss value until the second classification model to be trained meets the convergence condition.
[0269] In some embodiments, the sample pair acquisition module includes a tensor conversion sub-module and a tensor splicing sub-module.
[0270] The tensor conversion sub-module is configured to convert the embedding vectors corresponding to multiple enhanced training images corresponding to the current batch into a first tensor.
[0271] The tensor splicing sub-module is configured to combine and splice the first tensors corresponding to multiple enhanced training images in pairs to obtain second tensors corresponding to multiple groups of sample pairs.
[0272] In some embodiments, the calculation module is further configured to perform a dot product calculation according to the embedding vectors corresponding to the two enhanced training images included in the target sample pair to obtain the initial similarity corresponding to the target sample pair, where the target sample pair is any sample pair.
[0273] The calculation module is further configured to adjust the initial similarity according to the temperature coefficient to obtain the similarity corresponding to the target sample pair.
[0274] In some embodiments, the electrophoresis image recognition device 1100 includes a label acquisition module and a mask generation module.
[0275] The label acquisition module is configured to acquire the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs.
[0276] The mask generation module is configured to generate a mask according to the overall positive and negative labels corresponding to the two enhanced training images included in each group of sample pairs.
[0277] The calculation module is further configured to determine the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs included in the current batch according to the mask and the similarity of each group of sample pairs; a positive sample pair is a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included are the same, and a negative sample pair is a sample pair in which the overall positive and negative labels corresponding to the two enhanced training images included are different.
[0278] The calculation module is further configured to calculate the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs, where the first positive sample pair is any positive sample pair.
[0279] The calculation module is further configured to obtain an average value according to the ratio corresponding to each group of positive sample pairs, and calculate a loss value according to the average value.
[0280] In some embodiments, the feature acquisition module is configured to extract the features of the two enhanced training images corresponding to each second sample electrophoresis image through a feature extractor to be trained, and obtain the embedding vectors corresponding to the two enhanced training images respectively.
[0281] The model training module is further configured to adjust the parameters of the feature extractor to be trained according to the loss value until the feature extractor to be trained meets the convergence condition, and the feature extractor training is completed.
[0282] The model training module is further configured to, when the feature extractor training is completed, generate a feature vector corresponding to the input second sample electrophoresis image through the trained feature extractor, and determine the predicted classification result corresponding to the input second sample electrophoresis image through the fully connected layer to be trained according to the feature vector corresponding to the input second sample electrophoresis image.
[0283] The model training module is further configured to adjust the parameters of the fully connected layer according to the predicted classification result corresponding to the input second sample electrophoresis image and the corresponding overall positive and negative label until the fully connected layer to be trained meets the convergence condition, then the fully connected layer training is completed, and a trained second classification model is obtained.
[0284] Figure 12 It is a structural block diagram of an electronic device provided by an embodiment of the present application. As Figure 12 shown, the electronic device 1200 may include a memory 1202 and a processor 1201. A computer program is stored in the memory 1202. When the computer program is executed by the processor 1201, the electronic device 1200 implements the electrophoresis image recognition method described in the above embodiments.
[0285] The processor 1201 may include one or more processing cores. The processor 1201 connects various parts within the entire electronic device using various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking data stored in the memory, it performs various functions of the electronic device and processes data. Optionally, the processor 1201 may be implemented in at least one hardware form of digital signal processing, field programmable gate array, or programmable logic array. The processor 1201 may integrate a combination of one or several of a central processing unit (CPU for short), a graphics processing unit (GPU for short), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1201 and may be implemented separately through a communication chip.
[0286] The memory 1202 may include random access memory and may also include read-only memory. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created during the use of the electronic device.
[0287] An embodiment of the present application discloses a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, the processor is caused to implement the electrophoresis image recognition method described in the above various embodiments.
[0288] An embodiment of the present application discloses a computer program product that includes a computer program. When the computer program is executable by a processor, the processor is caused to implement the electrophoresis image recognition method described in the above various embodiments.
[0289] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disc, a ROM, etc.
[0290] The above are only specific examples of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An electrophoresis image recognition method, characterized in that, Including: The computer device identifies the target electrophoresis image through the first classification model, obtains the positive and negative prediction value corresponding to the target electrophoresis image, and determines the first classification result corresponding to the target electrophoresis image according to the positive and negative prediction value; The first classification model is trained according to a first data set, and the first data set includes multiple first sample electrophoresis images and the overall positive and negative labels corresponding to each of the first sample electrophoresis images; If the first classification result is suspected positive, the computer device analyzes the target electrophoresis image through the second classification model to obtain the second classification result corresponding to the target electrophoresis image; The second classification model is trained according to a suspected data set, and the suspected data set includes multiple second sample electrophoresis images marked as suspected data and the overall positive and negative labels corresponding to each of the second sample electrophoresis images; If the first classification result or the second classification result is positive, the computer device identifies each lane included in the target electrophoresis image through the third classification model to obtain the third classification result corresponding to each lane; The third classification model is trained according to a second data set, and the second data set includes multiple third sample electrophoresis images and the positive and negative labels corresponding to each lane in each of the third sample electrophoresis images; The computer device determines the target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result or the third classification result.
2. The method according to claim 1, wherein The determining the first classification result corresponding to the target electrophoresis image according to the positive and negative prediction value includes: If the positive and negative prediction value is greater than the first prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is positive; If the positive and negative prediction value is less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is negative; If the positive and negative prediction value is not greater than the first prediction threshold and not less than the second prediction threshold, the computer device determines that the first classification result corresponding to the target electrophoresis image is suspected positive.
3. The method according to claim 1, wherein The method further includes: The computer device obtains multiple second sample electrophoresis images included in the current batch; The computer device performs image enhancement processing on each second sample electrophoresis image in the current batch through two different image enhancement algorithms to obtain two enhanced training images corresponding to each second sample electrophoresis image; The computer device generates embedding vectors corresponding to the two enhanced training images respectively through the second classification model to be trained according to the two enhanced training images corresponding to each second sample electrophoresis image; The computer device combines the multiple enhanced training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs; The computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two enhanced training images included in each group of sample pairs; The computer device calculates the loss value according to the similarities of the multiple groups of sample pairs; The computer device adjusts the parameters of the second classification model to be trained according to the loss value until the second classification model to be trained meets the convergence condition.
4. The method according to claim 3, wherein The computer device combines the multiple augmented training images corresponding to the current batch in pairs to obtain multiple groups of sample pairs, including: The computer device converts the embedding vectors corresponding to the multiple augmented training images corresponding to the current batch into a first tensor; The computer device combines and splices the first tensors corresponding to the multiple augmented training images in pairs to obtain second tensors corresponding to the multiple groups of sample pairs.
5. The method according to claim 3, characterized in that, The computer device calculates the similarity of each group of sample pairs according to the embedding vectors corresponding to the two augmented training images included in each group of sample pairs, including: The computer device performs a dot product calculation on the embedding vectors corresponding to the two augmented training images included in the target sample pair to obtain the initial similarity corresponding to the target sample pair, where the target sample pair is any one of the sample pairs; The computer device adjusts the initial similarity according to the temperature coefficient to obtain the similarity corresponding to the target sample pair.
6. The method according to claim 3, wherein Before the computer device calculates the loss value according to the similarity of the positive sample pairs, the method further includes: The computer device obtains the overall positive and negative labels corresponding to the two augmented training images included in each group of sample pairs; The computer device generates a mask according to the overall positive and negative labels corresponding to the two augmented training images included in each group of sample pairs; The computer device calculates the loss value according to the similarity of the multiple groups of sample pairs, including: The computer device determines the similarity of each group of positive sample pairs and the similarity of each group of negative sample pairs included in the current batch according to the mask and the similarity of each group of sample pairs; the positive sample pair is a sample pair in which the overall positive and negative labels corresponding to the two augmented training images included are the same, and the negative sample pair is a sample pair in which the overall positive and negative labels corresponding to the two augmented training images included are different; The computer device calculates the ratio between the similarity of the first positive sample pair and the similarity of each group of negative sample pairs, where the first positive sample pair is any one of the positive sample pairs; The computer device obtains the average value according to the ratio corresponding to each group of positive sample pairs and calculates the loss value according to the average value.
7. The method according to claim 3, wherein The second classification model to be trained includes a feature extractor and a fully connected layer; the computer device generates the embedding vectors corresponding to the two augmented training images according to the two augmented training images corresponding to each second sample electrophoresis image through the second classification model to be trained, including: The computer device extracts the features of the two augmented training images corresponding to each second sample electrophoresis image through the feature extractor to be trained to obtain the embedding vectors corresponding to the two augmented training images; The computer device adjusts the parameters of the second classification model to be trained according to the loss value until the second classification model to be trained meets the convergence condition, including: The computer device adjusts the parameters of the feature extractor to be trained according to the loss value until the feature extractor to be trained meets the convergence condition, and the training of the feature extractor is completed; The method further includes: When the training of the feature extractor is completed, the computer device generates a feature vector corresponding to the input second sample electrophoresis image through the trained feature extractor, and determines a predicted classification result corresponding to the input second sample electrophoresis image through the fully connected layer to be trained according to the feature vector corresponding to the input second sample electrophoresis image; The computer device adjusts the parameters of the fully connected layer according to the predicted classification result corresponding to the input second sample electrophoresis image and the corresponding overall positive and negative label until the fully connected layer to be trained meets the convergence condition, and then the training of the fully connected layer is completed, and a trained second classification model is obtained.
8. An electrophoresis image recognition device, characterized in that, The device includes: A first classification module, configured to identify a target electrophoresis image through a first classification model, obtain a positive and negative predicted value corresponding to the target electrophoresis image, and determine a first classification result corresponding to the target electrophoresis image according to the positive and negative predicted value; the first classification model is trained according to a first data set, and the first data set includes multiple first sample electrophoresis images and the overall positive and negative label corresponding to each first sample electrophoresis image; A second classification module, configured to, if the first classification result is suspected positive, analyze the target electrophoresis image through a second classification model to obtain a second classification result corresponding to the target electrophoresis image; the second classification model is trained according to a suspected data set, and the suspected data set includes multiple second sample electrophoresis images marked as suspected data and the overall positive and negative label corresponding to each second sample electrophoresis image; A third classification module, configured to, if the first classification result or the second classification result is positive, identify each lane included in the target electrophoresis image through a third classification model to obtain a third classification result corresponding to each lane; the third classification model is trained according to a second data set, and the second data set includes multiple third sample electrophoresis images and the positive and negative label corresponding to each lane in each third sample electrophoresis image; An image recognition module, configured to determine a target classification result corresponding to the target electrophoresis image according to the first classification result, the second classification result or the third classification result.
9. An electronic device, characterized in that, It includes a memory and a processor, and a computer program is stored in the memory. When the computer program is executed by the processor, the processor implements the electrophoresis image recognition method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the processor implements the electrophoresis image recognition method according to any one of claims 1-7.