Reagent identification method and system based on self-supervised learning, and computer program product
By applying a self-supervised learning method in the field of reagent recognition, using VICReg loss function and ResNet50 network, we learn features from label-free data, and improve the robustness of the model through image enhancement technology, solving the problem of inaccurate identification of existing reagent recognition methods in complex environments, achieving efficient and accurate reagent recognition.
Patent Information
- Application Number
- CN202411851847.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
The existing reagent identification methods are inefficient and poorly robust, making it difficult to achieve efficient and accurate reagent identification in complex environments.
Using a self-supervised learning method, VICReg loss function and ResNet50 network are used to learn features from label-free data, and the robustness of the model is improved through image enhancement technology.
It significantly improves the accuracy, robustness and generalization ability of reagent identification, and can maintain efficient recognition performance in a diverse reagent environment.
Smart Images

Figure CN119942173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of laboratory intelligent management, and more specifically, to a reagent identification method, system and computer program product based on self-supervised learning. Background Art
[0002] In biochemical experiments, accurate identification and classification of various reagents are key steps to ensure the smooth progress of the experimental process. However, traditional reagent identification methods mainly rely on manual operation and empirical judgment, which are inefficient and easily interfered by external factors, resulting in inconsistent identification results. With the development of computer vision and artificial intelligence technology, visual recognition technology has gradually been applied to the automatic identification of reagents. However, existing visual recognition methods still have many problems and are difficult to meet the requirements for efficient and accurate identification of biochemical reagents in complex environments.
[0003] Early visual recognition methods mainly rely on traditional image processing techniques based on feature extraction and classifiers, such as edge detection, shape matching, and color histogram analysis. These methods usually require manual design of features and matching and classification through heuristic algorithms. Although these methods have certain effects in simple environments, they have many shortcomings in practical applications. First, the manually designed features are very sensitive to environmental changes, such as lighting conditions, shooting angles, and background clutter, resulting in poor robustness of the algorithm. Second, the feature matching method relies on manually set thresholds and is easily interfered by noise, especially when the appearance of reagent bottles is similar or there are stains and occlusions, the misrecognition rate is high. In addition, these methods are too dependent on the appearance features of the reagents, such as shape, color, and text, and it is difficult to handle reagent bottles with similar appearance but different contents. This limitation makes the reagent recognition method based on traditional image processing perform poorly in complex scenes, especially when there are many types of reagents and the environment is complex and changeable. The performance of traditional methods is far from meeting actual needs. Visual recognition technology based on machine learning has gradually emerged, and reagents are classified by using traditional machine learning algorithms such as support vector machines (SVM) and random forests. Such methods no longer rely on manually designed features, but learn features that can distinguish different reagents from data. However, these methods also have obvious limitations. First, machine learning algorithms usually require a large amount of labeled data for training, and it is often difficult to obtain enough labeled samples for reagent identification tasks in biochemical experiments, especially for some rare or special reagents. Secondly, traditional machine learning methods still show poor robustness when facing complex backgrounds and noise interference. Models usually only perform well in specific environments and are difficult to cope with diverse reagent identification scenarios. Finally, with the rise of deep learning, although deep learning methods based on convolutional neural networks (CNNs) have achieved great success in image classification and target detection tasks, they have extremely high demands on data volume, and the model training process is very dependent on large-scale labeled data sets. In actual reagent identification tasks, obtaining enough and high-quality labeled data is often a huge challenge. Therefore, it is of great practical significance to develop a new method to solve these problems. Summary of the invention
[0004] In view of the problems existing in the prior art, the present invention proposes a reagent identification method, system and computer program product based on self-supervised learning, which adopts a self-supervised learning method based on the VICReg (Variance-Invariance-CovarianceRegularization) loss function and combines the ResNet50 architecture as a feature extraction network to learn features from unlabeled data. At the same time, image enhancement technology is used to improve the robustness of the model, thereby achieving efficient and accurate identification of biochemical reagents.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a reagent identification method based on self-supervised learning, comprising:
[0006] Step S1, image acquisition and preprocessing: capturing a large number of reagent images from the laboratory environment or the reagent storage area, and preprocessing the acquired images to make the size and resolution of all input images consistent to meet the input requirements of the feature extraction network;
[0007] Step S2, image enhancement and contrast learning sample acquisition: perform random transformation on each reagent image to generate multiple image enhancement versions; different image enhancement versions of the same reagent are considered as positive samples, while images of different reagents are considered as negative samples;
[0008] Step S3, construction of a self-supervised learning model: using a ResNet50 network as a feature extraction network, the first few layers are responsible for extracting low-level features of the image, the middle layers are responsible for extracting intermediate features, and the last few layers are responsible for extracting high-level semantic features of the image; the output of the ResNet50 network is passed to a fully connected classifier;
[0009] Step S4, training, supervised fine-tuning and evaluation of the self-supervised learning model: the loss function used in the training of the self-supervised learning model is the VICReg loss function; the overall expression of the VICReg loss function is the weighted sum of the variance constraint, the invariance constraint and the covariance constraint:
[0010]
[0011] in, is the variance constraint; is the invariance constraint; is the covariance constraint; and are the corresponding weight coefficients. By adjusting these coefficients, the influence of the three constraints on model learning can be controlled;
[0012] Variance Constraint To avoid the feature representation space collapsing into a constant output, the formula is:
[0013]
[0014] in, It is The feature representation of a sample is is its variance, is a positive constant threshold, which is used to ensure that the variance will not be lower than a certain level, and N is the number of samples.
[0015] Invariance Constraints The formula for making different enhanced versions of the same reagent image have the same representation in the feature space is:
[0016]
[0017] in, and It is the feature representation of the same reagent image after different enhancements;
[0018] Covariance Constraints It is used to reduce the redundant information between different features so that each feature dimension can express different information independently. The formula is as follows:
[0019]
[0020] in, and are different dimensions of the feature vector, and d is the number of dimensions of the feature; the supervised fine-tuning uses a small amount of labeled data to fine-tune the features extracted by the model, during which the weights of the fully connected classifier are optimized through the cross entropy loss function, and the weights of the last few layers of the ResNet50 network are slightly adjusted;
[0021] Step S5: deploying the trained self-supervised learning model for reagent identification.
[0022] The present invention adopts ResNet50 as the feature extraction network, which is a classic deep residual network with strong feature extraction capabilities, especially in image classification and detection tasks. VICReg helps the model learn effective features that do not depend on labels through three main constraints: variance, invariance, and covariance. Through the combination of these three constraints, the VICReg loss function can effectively regularize the feature representation, so that the model can learn useful representations without relying on label data. The present invention significantly improves the accuracy, robustness and generalization ability of reagent identification by combining ResNet50 with the VICReg architecture.
[0023] In some embodiments of the first aspect, in step S1, different illuminations, viewing angles and backgrounds are involved in the image acquisition process; the image preprocessing process also includes contrast enhancement and filtering denoising. Different illuminations, viewing angles and backgrounds may be involved in the acquisition process, which provides a basis for robust training of the model in a real environment. Standardized adjustment of contrast and brightness helps to improve the visibility of reagent labels and appearance in the image, and further enhances the feature extraction capability. Noise removal reduces interference in the image through filtering and other techniques.
[0024] In some embodiments of the first aspect, in step S2, the random transformation includes but is not limited to rotation, horizontal flipping, random cropping, brightness adjustment, and adding Gaussian noise.
[0025] In some embodiments of the first aspect, in step S3, the ResNet50 network uses transfer learning technology and is initialized using ResNet50 weights pre-trained on the ImageNet dataset. In the case of limited computing resources or a small dataset, this technical solution can greatly improve the convergence speed and final performance of the model.
[0026] In some embodiments of the first aspect, in step S3, the fully connected classifier uses one or more fully connected layers, and uses a Softmax activation function in the final output layer to estimate the probability of the reagent category. Through the Softmax function, the model can generate a probability distribution for each category, and select the category with the largest probability as the identification result of the reagent.
[0027] In some embodiments of the first aspect, in step S4, during the supervised fine-tuning process, a dynamic learning rate strategy is adopted, where a larger learning rate is initially used to accelerate model convergence, and the learning rate is gradually reduced as the training progresses to ensure that the model can fully learn the details in the final stage.
[0028] In some embodiments of the first aspect, in step S4, the model evaluation indicators include but are not limited to accuracy, precision, recall, and F1 score. Through these evaluation indicators, the performance of the supervised fine-tuning model can be verified, and further optimization can be performed, such as adjusting the learning rate, optimizing the data enhancement strategy, etc., to improve the accuracy and robustness of the reagent classification.
[0029] In some embodiments of the first aspect, in step S5, the trained self-supervised learning model is pruned and quantized before deployment. Pruning is to reduce the size and computation of the model by reducing redundant network connections and neurons. Quantization is to accelerate the reasoning process by converting model parameters from floating point numbers to smaller data types. These optimizations can significantly improve the operating efficiency of the model, especially on devices with limited resources.
[0030] In a second aspect, the present invention provides a reagent identification system based on self-supervised learning, comprising:
[0031] The image preprocessing module makes the size and resolution of all acquired reagent images consistent to meet the input requirements of the feature extraction network;
[0032] The image enhancement module performs random transformations on each reagent image to generate multiple enhanced versions of the image; different enhanced versions of the same reagent are considered positive samples, while images of different reagents are considered negative samples;
[0033] The self-supervised learning model uses the ResNet50 network as the feature extraction network. The first few layers are responsible for extracting low-level features of the image, the middle layers are responsible for extracting intermediate features, and the last few layers are responsible for extracting high-level semantic features of the image. The output of the ResNet50 network is passed to a fully connected classifier.
[0034] The model training module adopts the VICReg loss function as the loss function; the expression of the VICReg loss function is the weighted sum of the variance constraint, invariance constraint and covariance constraint as described above; the self-supervised learning model is trained, supervised, fine-tuned and evaluated to meet the reagent recognition performance requirements.
[0035] In a final aspect, the present invention provides a computer program product for reagent identification; when the computer program product is run on a computer or device, the computer or device executes the reagent identification method based on self-supervised learning as described above.
[0036] Compared with the prior art, the present invention has the following technical effects:
[0037] 1. Comparison with traditional convolutional neural network (CNN) methods
[0038] Traditional convolutional neural networks (CNNs) have achieved remarkable success in image classification and recognition tasks, especially when there is sufficient labeled data. CNNs can extract spatial features in images through layer-by-layer convolution and pooling operations. However, the limitations of traditional CNNs are:
[0039] a. Data dependence: Traditional CNN models usually require a large amount of labeled data for supervised learning. For complex tasks, such as the identification of biochemical reagents, it is very difficult and expensive to obtain a large amount of labeled data. Therefore, traditional CNNs perform poorly in the absence of sufficient labeled data and are prone to overfitting.
[0040] b Limitations of feature learning: CNN extracts image features layer by layer. The low-level layers are responsible for extracting low-level features such as edges and colors, while the high-level layers are responsible for abstracting high-level semantic information. However, when faced with complex environments (such as lighting changes, noise interference, etc.), the features extracted by CNN often lack robustness and are sensitive to noise and background interference, resulting in a decrease in recognition accuracy.
[0041] In contrast, the present invention adopts a self-supervised learning method and realizes effective learning of unlabeled data through the VICReg loss function. VICReg uses three major constraints, namely variance, covariance, and invariance, to enable the model to extract more robust and universal features under multi-view data enhancement, thus overcoming the traditional CNN's dependence on large-scale labeled data.
[0042] In addition, in terms of architecture, this technology uses ResNet50 as the feature extraction network. Compared with ordinary CNN, ResNet50 introduces residual connections to solve the gradient vanishing problem in deep networks, allowing the model to maintain the ability to learn deep features while avoiding degradation during model training. This enables ResNet50 to learn deep features more effectively and show better performance in self-supervised learning.
[0043] 2. Comparison with Self-Supervised Learning Methods Using Other Loss Functions
[0044] Self-supervised learning technology has gradually attracted attention in recent years. Commonly used self-supervised loss functions include SimCLR, BYOL, MoCo, etc. These methods usually achieve unsupervised feature learning by maximizing the feature similarity of the same data under different perspectives and minimizing the feature similarity of different data through contrastive learning. Although these methods perform well in self-supervised learning, they have the following problems:
[0045] a. Disadvantages of contrastive learning: Contrastive learning methods such as SimCLR rely on a large number of negative sample pairs in order to distinguish the features of positive and negative samples. However, improper selection of negative samples will cause the model to produce "false negative samples", which will interfere with the learning process of the model and affect the expression effect of the features. BYOL avoids the use of negative samples by comparing different enhanced versions of the same sample, but its optimization process is prone to feature collapse, resulting in a lack of diversity in the features learned by the model.
[0046] b. Limitations of loss function design: SimCLR and other methods only focus on the distance and similarity between feature vectors, while ignoring the statistical properties of the feature vectors themselves, such as variance and covariance. This may lead to the collapse of the feature expression space, that is, the model cannot fully express the diversity between data, affecting the generalization ability of the features.
[0047] In contrast, the VICReg loss function used in the present invention is more comprehensive in design. By introducing three major constraints, VICReg can maintain the diversity and effectiveness of features when processing unlabeled data, avoiding the negative sample selection problem and feature collapse phenomenon in contrastive learning methods such as SimCLR. Therefore, the self-supervised learning method using VICReg shows higher accuracy and robustness in complex environments.
[0048] 3. Superiority of Classification Accuracy
[0049] In the task of reagent identification, the accuracy of traditional convolutional neural networks and self-supervised learning methods is often challenged in complex environments. The present invention significantly improves the accuracy, robustness and generalization ability of reagent identification by combining ResNet50 with the VICReg architecture. In a set of experimental evaluations, the present invention shows better performance than traditional CNN and SimCLR methods:
[0050] aImproved recognition accuracy: Under a variety of different lighting, viewing angles and backgrounds, the self-supervised learning model based on the VICReg loss function showed higher accuracy in the reagent recognition task. Experimental results show that compared with traditional CNN, the recognition accuracy is improved by more than 10%, and compared with self-supervised learning methods such as SimCLR, the recognition accuracy is improved by 5%-7%.
[0051] b Enhanced feature generalization capability: VICReg's multiple constraint mechanism enables the model to better adapt to diverse reagent environments and improves its performance in complex scenarios, especially in the presence of background noise and occlusion, the model can still maintain a high level of accuracy.
[0052] c Improved robustness: Compared with other methods, this technology can better handle image changes caused by lighting changes, angle offsets, and reagent appearance stains, showing stronger robustness and ensuring the practical application effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 The figure is a flow chart of a reagent identification method based on self-supervised learning in one embodiment of the present invention.
[0054] Figure 2 Schematic diagram of a reagent identification system based on self-supervised learning in one embodiment of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0056] In the following detailed description, many specific details are set forth to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that well-known algorithms or models are not shown in detail to avoid obscuring the subject matter of the present invention.
[0057] In addition, the order of execution of actions, steps, etc. in the devices and methods shown in the claims, specifications and drawings can be implemented in any order as long as there is no special explicit limitation on the order and the output of the previous processing is not used in the subsequent processing.
[0058] Example 1
[0059] See also Figure 1 This embodiment provides a reagent identification method based on self-supervised learning, comprising:
[0060] Step S1, image acquisition and preprocessing: Capture a large number of reagent images from the laboratory environment or the reagent storage area, and preprocess the acquired images to make the size and resolution of all input images consistent to meet the input requirements of the feature extraction network.
[0061] As a more detailed and preferred technical solution, in order to ensure the generalization ability of the model, the image acquisition process first needs to obtain high-definition images of biochemical reagents from multiple different environments, including laboratories, reagent storage rooms and other scenes. By using a high-resolution camera to capture images, ensure that the labels, appearance and other features of the reagent bottles are clearly displayed in the image. In order to cope with the changes in lighting and different shooting angles in actual operations, try to cover various lighting conditions, different angles and background interference factors during image acquisition, such as different shelves and bottle placement methods. The purpose of image acquisition is to provide a diverse data source for subsequent model training to enhance the robustness of the model in complex environments.
[0062] The image preprocessing steps include cropping, scaling, denoising, and normalization to ensure that the input image meets the input requirements of the ResNet50 network. Each acquired image is cropped to an area containing only the reagent bottle and label to avoid interference from redundant background information on the model. The image is then scaled to 224×224 pixels, which is the standard input size of the ResNet50 model.
[0063] In addition, the image can be de-noised and contrast enhanced. Noise removal is mainly achieved by applying a Gaussian filter to reduce random interference in the image. Contrast enhancement adjusts the brightness and contrast of the image to make the text and color features on the reagent bottle label clearer, thereby improving the accuracy of feature extraction.
[0064] After preprocessing, the data format, size, and brightness of all images are standardized to ensure that they can be used as a unified input source for the self-supervised learning model training phase.
[0065] Step S2, image enhancement and contrast learning sample acquisition: Perform random transformation on each reagent image to generate multiple image enhancement versions; different image enhancement versions of the same reagent are considered as positive samples, while images of different reagents are considered as negative samples.
[0066] This step plays a key role in this recognition method, especially in the self-supervised learning stage, where positive and negative samples are generated through a variety of enhancement strategies to help the model learn stable and discriminative features.
[0067] a Enhancement strategy:
[0068] The core idea of image enhancement is to perform a series of random transformations on each reagent image to simulate various possible changes of the image in different environments, thereby enhancing the model's adaptability to images from different perspectives. Specific enhancement strategies include:
[0069] •Rotation: Randomly rotate a certain angle to simulate the reagent image at different viewing angles;
[0070] •Horizontal Flip: Randomly flip the image horizontally;
[0071] •Brightness adjustment: randomly change the brightness of the image to simulate image changes under different lighting conditions;
[0072] •Gaussian noise addition: Randomly add Gaussian noise to the image to simulate the noise interference that may occur in reality.
[0073] Through these enhancement strategies, the system can generate multiple versions of each reagent image. These different versions of images are used as positive or negative samples in subsequent self-supervised learning, helping the model to distinguish stable and discriminative features during the learning process.
[0074] b. Construction of positive and negative samples:
[0075] After image enhancement, the system considers different enhanced versions of the same reagent as "positive samples", that is, these images all represent the same reagent bottle, so their features should remain consistent. Enhanced versions of different reagents are considered "negative samples", that is, these images should have significantly different feature representations. By comparing positive and negative samples, the model can learn how to extract robust features from the multi-view representation of the reagent that are not affected by the enhancement transformation.
[0076] Contrastive learning of positive and negative samples is the basis of self-supervised learning, which enables the model to learn effective feature representations without annotations. Through contrastive learning, the model can learn robust features from multi-view representations of reagents that are not affected by augmented transformations.
[0077] Step S3, construction of self-supervised learning model. First, the model uses ResNet50 network as the feature extraction network. The first few layers are responsible for extracting low-level features of the image, such as edges, textures, etc.; the middle layers are responsible for extracting intermediate features, such as shapes and contours; and the latter layers are responsible for extracting high-level semantic features of the image, such as the label information of the reagent bottle, color distribution, etc.
[0078] The ResNet50 network is a classic deep residual network with 50 layers. Its characteristic is that residual connections are introduced in each layer to solve the gradient vanishing problem in deep neural network training. The ResNet50 network has strong feature extraction capabilities, especially in image classification and detection tasks.
[0079] aNetwork structure: ResNet50 solves the gradient vanishing problem in deep neural networks through residual connections, ensuring that the model can effectively learn deep features. In self-supervised learning, the first few layers of ResNet50 are used to extract low-level features of the image (such as edges, textures, etc.), the middle layers are responsible for extracting intermediate features such as shapes and contours; the latter layers are responsible for extracting high-level semantic features (such as the shape and labels of reagent bottles). The basic residual module of ResNet50 consists of two or three convolutional layers. The input of each module is directly jump-connected to the output to form a "residual". This design allows the network to avoid the gradient vanishing phenomenon while maintaining a deeper level. The output formula of the residual module is as follows:
[0080]
[0081] in, represents the feature map after convolution, batch normalization and ReLU activation function, is the input of the module, is the weight of the convolution kernel.
[0082] b Transfer learning: The weights of ResNet50 can be initialized through transfer learning, and the pre-trained weights can be used to accelerate the training process, especially when computing resources are limited or the data set is small. This method can greatly improve the convergence speed and final performance of the model. In order to accelerate the training process and improve the feature extraction ability of the model, this embodiment adopts the technology of transfer learning, and uses the weights of ResNet50 pre-trained on the ImageNet data set for initialization. Since the ImageNet data set contains a large number of natural image categories, the pre-trained ResNet50 has learned a wealth of low-level and mid-level image features, which can be effectively transferred to the reagent identification task, thereby shortening the model convergence time and improving the recognition accuracy.
[0083] In the self-supervised learning stage, the model learns rich image features, but has not yet performed classification for specific tasks. Therefore, these features are then sent to the fully connected layer for classification tasks.
[0084] Classifier design: In the self-supervised learning stage, the last layer of features of ResNet50 is passed to a fully connected classifier. The task of this classifier is to map the extracted features to specific reagent categories. In this embodiment, the classifier design uses one or more fully connected layers, and uses the Softmax activation function in the final output layer to estimate the probability of the reagent category.
[0085] The formula of the Softmax function is as follows:
[0086]
[0087] in, is the probability that sample x is classified as class i, is the score of the i-th category output by the classifier, and C is the total number of categories.
[0088] Through the Softmax function, the model can generate a probability distribution for each category and select the category with the highest probability as the reagent identification result. Through supervised learning, the classifier can map the extracted features to specific reagent categories.
[0089] Step S4, training, supervised fine-tuning and evaluation of the self-supervised learning model. The enhanced version of the reagent image is used as the training set and validation set of the self-supervised learning model; the VICReg loss function trained by the self-supervised learning model is the core technology of self-supervised learning in the present invention. VICReg helps the model learn effective features that do not depend on labels through three main constraints: variance, invariance and covariance. Through the combination of these three constraints, the VICReg loss function can effectively regularize the feature representation, so that the model can learn discriminative and robust features without relying on labeled data.
[0090] Variance constraint: The purpose of variance constraint is to prevent the feature representation space from collapsing into a constant output. Specifically, it requires that the feature vectors in the same batch maintain sufficient variance to ensure the diversity of features. The formula for variance constraint is as follows:
[0091]
[0092] in, It is The feature representation of a sample is is its variance, is a positive constant threshold, which is used to ensure that the variance will not be lower than a certain level, and N is the number of samples.
[0093] Invariance constraint: The invariance constraint ensures that different enhanced versions of the same reagent image have the same representation in the feature space, thereby achieving the robustness of the model to image transformation. Its formula is:
[0094]
[0095] in, and It is the feature representation of the same reagent image after different enhancement transformations, and the constraint model maintains consistent representation of the same reagent from different perspectives.
[0096] Covariance constraint: Covariance constraint is used to reduce redundant information between different features and ensure that each feature dimension expresses different information independently. The formula is as follows:
[0097]
[0098] in, and are the different dimensions of the feature vector, and d is the number of feature dimensions. Through this constraint, the model can remove redundancy between features and improve the expressiveness of features.
[0099] The overall expression of the VICReg loss function is:
[0100] The overall expression of the VICReg loss function is the weighted sum of three constraints:
[0101]
[0102] in, and are the corresponding weight coefficients. By adjusting these coefficients, the influence of the three constraints on model learning can be controlled.
[0103] In order to enable the model to have the ability to classify reagents and further improve the recognition accuracy, it is necessary to further supervise fine-tune these self-supervised learned features. The supervised fine-tuning uses a small amount of labeled data to fine-tune the features extracted by the model. During this process, the weights of the fully connected classifier are optimized through the cross entropy loss function, and the weights of the last few layers of the ResNet50 network are slightly adjusted to improve the accuracy and precision of reagent recognition.
[0104] More specifically, after designing the classifier, this embodiment uses a small amount of labeled data to fine-tune the features extracted by self-supervised learning. Supervised fine-tuning optimizes the classifier through a cross entropy loss function so that it can better classify on the labeled data.
[0105] The expression of the cross entropy loss function is:
[0106]
[0107] Where N is the total number of samples, C is the number of categories, is the label of whether the i-th sample belongs to the c-th class (1 means it belongs, 0 means it does not belong), is the probability that the model predicts that the sample belongs to the cth class.
[0108] During the supervised fine-tuning process, the cross-entropy loss function was used to optimize the weights of the fully connected layers, and the weights of the last few layers of the ResNet50 network were slightly adjusted to ensure that the model can perform better on the reagent classification task.
[0109] After supervised fine-tuning, the performance of the model will be quantitatively analyzed through a series of evaluation indicators, including accuracy, precision, recall, and F1 score, etc. These evaluation indicators can reflect the performance of the model in the reagent classification task and help further optimization.
[0110] The formulas for these indicators are as follows:
[0111] Accuracy:
[0112] Indicates the proportion of samples correctly classified by the model to the total samples:
[0113]
[0114] Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative.
[0115] Precision:
[0116] Indicates the proportion of samples predicted by the model to be positive that are actually positive:
[0117]
[0118] Recall:
[0119] Indicates the proportion of samples that are actually positive and are correctly identified by the model:
[0120]
[0121] F1 score:
[0122] The harmonic mean of precision and recall is used to measure the overall performance of the model in the classification task:
[0123]
[0124] Through these evaluation indicators, the performance of the supervised fine-tuning model can be verified and further optimized, such as adjusting the learning rate, optimizing the data augmentation strategy, etc., to improve the accuracy and robustness of reagent classification.
[0125] Once the model training and evaluation are completed, the reagent identification system of this embodiment can be deployed in a laboratory management system or a reagent management system for automatic identification and classification of reagents.
[0126] aDeployment environment:
[0127] The model can be deployed in the **Laboratory Information Management System (LIMS)**. The system acquires reagent images in real time by connecting to a camera device and inputs the images into the trained model for recognition and classification. The recognized reagent information can be automatically recorded in the experimental management database to facilitate subsequent experimental operations and reagent tracking.
[0128] When deployed, the model can run on a server, edge computing device, or local computer, depending on the equipment and network conditions in the lab.
[0129] b Actual application scenarios:
[0130] In actual operation, the experimenter only needs to place the reagent bottle under the camera, and the system will automatically identify the type of reagent and output relevant information, such as the reagent name, usage precautions, storage requirements, etc. This process does not require manual input, reduces human operation errors, and improves experimental efficiency and the accuracy of reagent management.
[0131] During model deployment and application, the model may need to be further optimized according to specific application scenarios. This embodiment also provides the following commonly used optimization strategies:
[0132] a. Optimization of data enhancement:
[0133] Although the image enhancement strategy has been used in the training phase, in actual applications, it can be further adjusted according to the specific environment of the laboratory. For example, if some environmental lighting conditions are special, the brightness adjustment in image enhancement can be optimized to make the model more robust to lighting changes.
[0134] b. Adjustment of learning rate:
[0135] In the process of fine-tuning, the choice of learning rate is crucial. Usually, a dynamic learning rate strategy is adopted, that is, the learning rate is adjusted according to the change of loss during training. At the beginning, a larger learning rate is used to accelerate the convergence of the model, and the learning rate is gradually reduced as the training progresses to ensure that the model can fully learn the details in the final stage.
[0136] c Model pruning and quantization:
[0137] To make the model more efficient in practical applications, the trained model can be pruned and quantized. Pruning reduces the size and computation of the model by reducing redundant network connections and neurons. Quantization speeds up the inference process by converting model parameters from floating point numbers to smaller data types (such as INT8). These optimizations can significantly improve the operating efficiency of the model, especially on devices with limited resources.
[0138] The present invention proposes an efficient, robust and highly generalized biochemical reagent identification method by combining VICReg self-supervised learning with the ResNet50 network architecture. Compared with traditional convolutional neural networks and other self-supervised learning methods, the present invention can automatically learn the effective features of reagents without the need for a large amount of labeled data, and improve classification accuracy through supervised fine-tuning. In practical applications, the system can automatically identify and classify reagents, reduce human errors, improve experimental efficiency, and provide efficient technical support for biochemical experiments.
[0139] Example 2
[0140] See also Figure 2 This embodiment provides a reagent identification system based on self-supervised learning, which is used to implement the reagent identification method as described in Example 1, including:
[0141] The image preprocessing module 100 makes the size and resolution of all acquired reagent images consistent to meet the input requirements of the feature extraction network;
[0142] The image enhancement module 200 performs random transformation on each reagent image to generate multiple image enhancement versions; different image enhancement versions of the same reagent are considered as positive samples, while images of different reagents are considered as negative samples;
[0143] The self-supervised learning model 300 uses a ResNet50 network as a feature extraction network, wherein the first several layers are responsible for extracting low-level features of the image, the middle layers are responsible for extracting intermediate features, and the last several layers are responsible for extracting high-level semantic features of the image; the output of the ResNet50 network is passed to a fully connected classifier;
[0144] Model training module 400, the image enhancement version obtained by image enhancement module 200 is used as the training set and verification set of self-supervised learning model 300; VICReg loss function is used as the loss function for model training; the expression of the VICReg loss function is the weighted sum of the variance constraint, invariance constraint and covariance constraint as described in Example 1; the self-supervised learning model is trained, supervised fine-tuned and evaluated to meet the reagent recognition performance requirements.
[0145] The above-mentioned reagent identification method can be embodied in the form of a computer program product or a software functional unit. If the above-mentioned reagent identification method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Therefore, the technical solution is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable an electronic system (which can be a personal computer, a server, or a network system, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk.
[0146] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0147] Those skilled in the art should understand that those skilled in the art can implement variations by combining the prior art and the above embodiments, which will not be described in detail here. Such variations do not affect the essential content of the present invention, and will not be described in detail here.
[0148] The above describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the systems and structures that are not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention are still within the scope of protection of the technical solutions of the present invention.
Claims
1. A reagent identification method based on self-supervised learning, characterized in that: include: Step S1, image acquisition and preprocessing: capturing a large number of reagent images from the laboratory environment or the reagent storage area, and preprocessing the acquired images to make the size and resolution of all input images consistent to meet the input requirements of the feature extraction network; Step S2, image enhancement and contrast learning sample acquisition: perform random transformation on each reagent image to generate multiple image enhancement versions; different image enhancement versions of the same reagent are considered as positive samples, while images of different reagents are considered as negative samples; Step S3, construction of self-supervised learning model: ResNet50 network is used as feature extraction network, the first few layers are responsible for extracting low-level features of the image, the middle layers are responsible for extracting intermediate features, and the last few layers are responsible for extracting high-level semantic features of the image; The output of the ResNet50 network is passed to a fully connected classifier; Step S4, training, supervised fine-tuning and evaluation of the self-supervised learning model: the loss function used in the training of the self-supervised learning model is the VICReg loss function; the overall expression of the VICReg loss function is the weighted sum of the variance constraint, the invariance constraint and the covariance constraint: in, is the variance constraint; is the invariance constraint; is the covariance constraint; and are the corresponding weight coefficients. By adjusting these coefficients, the influence of the three constraints on model learning can be controlled; Variance Constraint To avoid the feature representation space collapsing into a constant output, the formula is: in, It is The feature representation of a sample is is its variance, is a positive constant threshold, which is used to ensure that the variance will not be lower than a certain level, and N is the number of samples. Invariance Constraints The formula for making different enhanced versions of the same reagent image have the same representation in the feature space is: in, and It is the feature representation of the same reagent image after different enhancements; Covariance Constraints It is used to reduce the redundant information between different features so that each feature dimension can express different information independently. The formula is as follows: in, and are the different dimensions of the feature vector, and d is the number of dimensions of the feature; The supervised fine-tuning uses a small amount of labeled data to fine-tune the features extracted by the model. During this process, the weights of the fully connected classifier are optimized through the cross entropy loss function, and the weights of the last few layers of the ResNet50 network are slightly adjusted; Step S5: deploying the trained self-supervised learning model for reagent identification.
2. The reagent identification method based on self-supervised learning according to claim 1, characterized in that: In step S1, the image acquisition process involves different illuminations, viewing angles and backgrounds; the image preprocessing process also includes contrast enhancement and filtering denoising.
3. The reagent identification method based on self-supervised learning according to claim 1, characterized in that: In step S2, the random transformation includes but is not limited to rotation, horizontal flipping, random cropping, brightness adjustment, and adding Gaussian noise.
4. The self-supervised learning driven reagent identification method according to claim 1, wherein the core feature is During the implementation of step S3, a transfer learning strategy is adopted. Specifically, this method uses the weights of the ResNet50 network pre-trained on the ImageNet Large Scale Visual Recognition Challenge dataset as initialization parameters to improve the network's performance in the reagent identification task.
5. The reagent identification method based on self-supervised learning according to claim 1 or 4, characterized in that: In step S3, the fully connected classifier uses one or more fully connected layers, and uses a Softmax activation function in the final output layer to estimate the probability of the reagent category.
6. The reagent identification method based on self-supervised learning according to claim 1, characterized in that: In step S4, during the supervised fine-tuning process, a dynamic learning rate strategy is adopted, where a larger learning rate is used at the beginning to accelerate the convergence of the model, and the learning rate is gradually reduced as the training proceeds.
7. The reagent identification method based on self-supervised learning according to claim 1, characterized in that: In step S4, the model evaluation indicators include but are not limited to accuracy, precision, recall, and F1 score.
8. The reagent identification method based on self-supervised learning according to claim 1, characterized in that: In step S5, the trained self-supervised learning model is pruned and quantized before deployment.
9. A reagent identification system based on self-supervised learning, characterized in that: include: The image preprocessing module makes the size and resolution of all acquired reagent images consistent to meet the input requirements of the feature extraction network; The image enhancement module performs random transformations on each reagent image to generate multiple enhanced versions of the image; different enhanced versions of the same reagent are considered positive samples, while images of different reagents are considered negative samples; The self-supervised learning model uses the ResNet50 network as the feature extraction network. The first few layers are responsible for extracting low-level features of the image, the middle layers are responsible for extracting intermediate features, and the last few layers are responsible for extracting high-level semantic features of the image. The output of the ResNet50 network is passed to a fully connected classifier. A model training module, wherein the training adopts a VICReg loss function as a loss function; the expression of the VICReg loss function is a weighted sum of the variance constraint, the invariance constraint and the covariance constraint as described in claim 1; and the self-supervised learning model is trained, supervised, fine-tuned and evaluated to meet the reagent recognition performance requirements.
10. A computer program product, characterized in that Used for reagent identification; when the computer program product runs on a computer or device, the computer or device executes the reagent identification method based on self-supervised learning as described in any one of claims 1 to 8.