Reagent board card detection result identification method and model construction method thereof

By building a multi-channel object detection network model based on YOLOv8, combined with the NMS algorithm and non-local attention mechanism module, the shortcomings of manual comparison in the identification of reagent board detection results and the accuracy and efficiency of existing image recognition models in multi-point recognition are solved, and the automatic recognition effect with high accuracy and high efficiency is achieved.

CN120014240APending Publication Date: 2025-05-16QINGDAO HIGHTOP BIOTECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510134403.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing reagent board detection results recognition methods mainly rely on manual comparison, which is large in workload and error-prone. The existing image recognition models have problems of poor recognition accuracy and low efficiency when identifying multi-point positions.

Method used

A reagent board detection result recognition model is constructed, and a multi-channel object detection network model based on YOLOv8 is adopted. In combination with the NMS algorithm module, the non-local attention mechanism module NonLocalBlockND is added to capture the relationship between all positions in the input feature map and improve the accuracy and robustness of the detection.

Benefits of technology

Automatic recognition of 24-point reagent board detection results is realized, which improves recognition efficiency and saves human resources. Through targeted training sets, preprocessing and post-processing methods, the anti-interference ability is enhanced and the accuracy of the recognition results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014240A_ABST
    Figure CN120014240A_ABST
Patent Text Reader

Abstract

The invention provides a reagent board card detection result identification method and a model construction method thereof, and belongs to the technical field of detection based on computer vision. The method comprises the following steps: firstly, acquiring original images and cut images of reagent board cards at N point locations, preprocessing the original images and cut images, constructing a data set, and then designing and constructing a reagent board card detection result identification model which comprises a multi-channel target detection network model based on YOLOv8 improved design and an NMS algorithm module connected in series with the multi-channel target detection network model, and training the model through the data set to obtain an optimal model, and carrying out real-time reagent board card detection result identification through deployment of the optimal model. According to the invention, improvement and design are carried out based on a YOLOv8 detection algorithm, a brand new reagent board card detection result identification model is constructed in combination with an NMS algorithm, automatic identification of 24-site reagent board card detection results is realized, and the results are directly output. Compared with a manual identification method, the identification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision-based detection technology, and in particular relates to a reagent board detection result recognition method and a model construction method thereof. Background Art

[0002] Food intolerance is a complex allergic disease. The human immune system regards one or more foods that enter the human body as harmful substances, thus producing an excessive protective immune response against these substances and producing food-specific IgG antibodies. IgG antibodies form immune complexes with food particles (type III allergy), which can cause inflammatory reactions in all tissues (including blood vessels) and manifest as symptoms and diseases in all systems of the body.

[0003] Food Specific IgG Antibody Detection Kit (Western Blot Method) (abbreviated as: Reagent Board) is used for qualitative detection of food specific IgG antibodies in human serum, plasma and whole blood.

[0004] Through food intolerance testing, we can find out the possible causes of the disease and avoid letting inappropriate food continue to damage the body. By formulating a reasonable diet management plan and adopting the method of fasting or eating less intolerant foods, we can control the source of the disease, curb the continuous development of the disease, and ultimately significantly improve the quality of life.

[0005] The identification of reagent plate results currently mainly uses the manual comparison method to compare the color of the plate after detection with the standard colorimetric card and read the result. The manual recognition method is not only labor-intensive, but also prone to errors. The image recognition method is fast and easy to use, but most of the image recognition models on the market are general models, such as face recognition, text recognition, license plate recognition, etc., or they are used to identify a single point. The reagent plate has as many as 24 recognition points. If the existing image recognition model is used to identify multiple points on the reagent plate, problems such as poor recognition accuracy and low recognition efficiency are likely to occur. Summary of the invention

[0006] In view of the above problems, the first aspect of the present invention provides a method for constructing a reagent card test result recognition model, comprising the following process: Step 1, obtaining the original image of the reagent card at N points and its cropped image; the cropped image automatically crops the target range to be identified according to prior information; Step 2, preprocessing the image obtained in step 1, wherein the preprocessing includes grayscale image processing, and constructing a training set and a test set together with the cropped color RGB image; Step 3, constructing a reagent board test result recognition model, the model includes a multi-channel target detection network model improved based on YOLOv8 and an NMS algorithm module connected in series with it; The multi-channel target detection network model includes a backbone network, a Neck network, and a Head network; wherein a non-local attention mechanism module NonLocalBlockND is added before the spatial pyramid pooling module SPPF in the backbone network to capture the relationship between all positions in the input feature map. This module better captures the color features of each small reaction area on the reagent board by obtaining the interaction between each position and all other positions in the feature map, and extracts the global context information of multiple small reaction areas. This global context information includes the overall layout of the reagent board, the background color and texture, the relationship between the quality control position and the reaction position, and the influence of external environmental factors. By integrating this information, the model can more accurately identify and classify the state of each reaction position, thereby improving the accuracy and robustness of detection; The NMS algorithm module selects N non-overlapping bounding boxes based on the position, range, recognition results of a single point and confidence of the M bounding boxes output by the multi-channel target detection network model. Through the NMS algorithm, redundant and overlapping bounding boxes can be removed, and the most representative N bounding boxes can be retained. Then, according to the maximum range of these bounding boxes, the image to be identified is evenly divided, and the coordinate range of the N points in the image to be identified is calculated, so as to obtain the color depth recognition results of each detection position; Step 4: Train the constructed reagent board test result recognition model and select the model with the best performance as the final model.

[0007] Preferably, the multi-channel target detection network model processes the input image by branching according to the channel; the backbone network part is composed of two branches: the first branch is the RGB channel, and the second branch is the grayscale channel; the information of different channels is taken out through the SilenceChannel module (used to separate and select different channel information of the input image, which greatly facilitates the calling and other work of the dual backbone) and put into two channels respectively, and extracts different features of the picture through a series of convolution operations, and then the two features are fused through the Concat module, so as to effectively separate the foreground and background of the task and better capture the target; Each branch consists of 5 convolutional modules and 4 C2f (i.e. Convolutional-2-Feature modules, which consist of two convolutional layers, multiple bottleneck layers, and residual structures) modules; the convolutional module is used to extract basic features, including edges and textures. Each module contains one or more convolutional layers, as well as activation functions and normalization layers, gradually increasing the abstract level of features and providing rich and diverse feature representations for subsequent modules; the C2f module reduces redundant parameters through optimized design, and contains multiple convolutional layers, residual connections, and bottleneck structures, which helps to transmit more information in the network; At the same time, the backbone network adopts the CSPDarkNet (Cross Stage Partial Darknet) structure, uses a series of convolution and deconvolution layers, and introduces residual connections and bottleneck structures; in the feature fusion part, the features of the two branches are spliced ​​together through a Concat module, and the spatial pyramid pooling SPPF module and the non-local attention mechanism NonLocalBlockND module are applied to the spliced ​​feature map.

[0008] Preferably, the input size of the non-local attention mechanism module NonLocalBlockND is T×H×W×1024, where T represents the number of frames or the length of the time series, H×W is the spatial dimension of the feature map, and 1024 is the number of channels; the module generates feature representations respectively through 1x1 convolution operations, which can be understood as query vectors, generated key vectors, and value vectors from left to right. These three branches are all reduced to 512 channels through 1×1 convolutions, and the output is T×H×W×512; then the two left branches are matrix multiplied to calculate the similarity between all positions relationship; then it is normalized through softmax to generate an attention weight matrix, and the output is THW×THW, which represents the similarity between each position in the input feature map; the output value is matrix multiplied with the third branch to obtain a weighted global feature representation with a size of T×H×W×512; the weighted features are restored to the original number of channels 1024 through 1x1 convolution, and a residual structure is performed with the original input to output a feature map with a size of T×H×W×1024; the residual structure can retain the original feature information and superimpose the globally weighted features.

[0009] Preferably, the Neck network is responsible for multi-scale feature fusion, and adopts a path aggregation network-feature pyramid network PAN-FPN to fuse the feature maps obtained by the backbone network at different stages. PAN-FPN includes two PAN modules for path aggregation of features at different levels, and enhances the representation ability of feature maps through bottom-up and top-down paths; the outputs P3, P4, and P5 of the backbone network are input into the PAN-FPN network structure to realize the fusion of multi-scale feature maps; specifically as follows: First, the P5 feature map is upsampled and fused with the P4 feature map to obtain F1; Next, F1 passes through a C2f layer and is upsampled again and fused with the P3 feature map to obtain T1; Then, T1 is passed through a convolutional layer and fused with the output of F1 to obtain F2; Next, F2 passes through the C2f layer to obtain T2; T2 goes through another convolutional layer and is fused with the P5 feature map to obtain F3; Finally, F3 passes through the C2f layer to obtain T3.

[0010] Preferably, the Head network includes a detection head and a classification head; the detection head contains a series of convolutional layers and deconvolutional layers for generating detection results; these layers are responsible for predicting the bounding box regression value of each anchor box and the confidence of the target. Through decoupling, the detection head is more focused on the regression task and improves the accuracy of bounding box prediction; the classification head classifies each feature map through a global pooling layer, and outputs the probability distribution of each category by reducing the dimension of the feature map; The output of each detection head is a feature map whose size is a certain ratio of the input image size; each position on each feature map predicts the bounding box, category, and confidence, and finally merges them together to form the final result, outputting a four-dimensional vector representing the upper left corner coordinates x, y and lower right corner coordinates x, y of the target box.

[0011] Preferably, the specific processing process of the NMS algorithm module is: Use the NMS algorithm to select 24 bounding boxes from the input 840 bounding boxes; According to the maximum range of the bounding box, the image to be identified is evenly divided, and the coordinate range of 24 points of the image to be identified, that is, 6 rows and 4 columns, is calculated; At this time, the points on the image obtained and the 24 bounding boxes obtained are processed in three cases: if there are multiple bounding boxes at a certain point, the bounding box with the highest confidence is selected as the detection result of the current point; if there is no bounding box at a certain point, its result is recognized as 0 negative; if the number of bounding boxes obtained is significantly different from the actual number of 24 points, the recognition fails; The results at the 24 points are sorted according to the order on the test kit, which is the identified test results of the reagent board.

[0012] The second aspect of the present invention provides a method for identifying the test results of a reagent board card, which converts the test result identification model of the reagent board card constructed by the construction method described in the first aspect into ONNX format and deploys it to a backend server, and includes the following process: Step 1, capture and obtain the reagent board image and cropping range parameters; Step 2, calculate the grayscale channel of the image, adjust the size of the grayscale image, normalize the grayscale image, expand the dimension, and add channel processing to form 4-channel picture information to obtain a preprocessed image; Step 3, inputting the preprocessed image into the reagent board test result recognition model; Step 4: post-process the board recognition results, obtain the actual detection results and output them.

[0013] The third aspect of the present invention provides a reagent board test result identification device, the device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of the reagent board test result identification model constructed by the construction method described in the first aspect; when the processor executes the computer execution program stored in the memory, the processor executes a reagent board test result identification method.

[0014] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer execution program of the reagent board card test result recognition model constructed by the construction method described in the first aspect. When the computer execution program is executed by a processor, the processor executes a reagent board card test result recognition method.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention is improved and designed based on the YOLOv8 detection algorithm, and a new reagent board detection result recognition model is constructed in combination with the NMS algorithm, which realizes the automatic recognition of the 24-point reagent board detection results and directly outputs the results. Compared with the manual recognition method, the recognition efficiency is improved and human resources are saved; compared with other image recognition methods, the use of targeted training sets, pre-processing and post-processing methods enhances the anti-interference ability and improves the accuracy of the recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a picture of the original board captured by the WeChat applet in the embodiment of the present invention.

[0017] Figure 2 This is a cropped picture in the embodiment.

[0018] Figure 3 It is a color gradient diagram in the embodiment.

[0019] Figure 4 This is the internal structure diagram of the NonLocalBlockND module added in the present invention.

[0020] Figure 5 This is a structural diagram of the improved YOLOv8 model used in the present invention.

[0021] Figure 6 It is the data flow direction of the Neck part in the network structure of the present invention.

[0022] Figure 7 It is the data flow direction of the Head part in the network structure of the present invention.

[0023] Figure 8The parameters of the PyTorch model trained by the present invention after conversion to the ONNX model.

[0024] Fig. 9 An example of a bounding box identified for an image of a reagent cartridge.

[0025] Fig.10 This is a flow chart for identifying the test results of the reagent board.

[0026] Fig.11 This is a simplified diagram of the equipment for identifying test results on the reagent board. DETAILED DESCRIPTION

[0027] The invention will be further described below in conjunction with specific embodiments.

[0028] This example further illustrates the method proposed in the present invention through a specific experimental process.

[0029] 1. Get the raw data Use the camera function of the WeChat applet to take photos of various scenes in different experimental environments (sufficient light, insufficient light, flash on, etc.) and different angles of the same scene to obtain the image data of the test kit and the image cropping parameters, such as Figure 1 Then, the image data and the image cropping parameters are uploaded to the backend server through the http protocol.

[0030] 2. Data Preprocessing (1) Image Cropping Since the detection position accounts for a small proportion of the entire image, in order to obtain better detection results, the image is cropped and enlarged according to prior information, so as to focus on and extract the target area to be identified. Figure 2 shown.

[0031] (2) Image Annotation The processed image is labeled according to the actual color gradient category. Figure 3 shown.

[0032] 3. Model building Construct a reagent board test result recognition model, which includes a multi-channel target detection network model improved based on YOLOv8 and a NMS algorithm module in series; The multi-channel target detection network model includes a backbone network, a neck network, and a head network. A non-local attention mechanism module NonLocalBlockND is added before the spatial pyramid pooling module SPPF in the backbone network to capture the relationship between all positions in the input feature map. It calculates the interaction between each position in the feature map and all other positions to better capture the color features of each small reaction area on the reagent board and capture the global context information of multiple small reaction areas. The NMS algorithm module selects N bounding boxes based on the position, range, recognition results and confidence of a single point of the M bounding boxes output by the multi-channel target detection network model, evenly divides the image to be recognized according to the maximum range of the bounding boxes, and calculates the coordinate range of the N points of the image to be recognized, thereby obtaining the recognition result.

[0033] The multi-channel target detection network model consists of Backbone (backbone network), Neck (feature enhancement network), and Head (detection head) to form the target detection network YOLOv8, and the backbone network is improved. A NonLocalBlockND (non-local attention mechanism) module is added before the SPPF (spatial pyramid pooling) module in the backbone network.

[0034] This module can effectively capture the relationship between all positions in the input feature map. It enhances the global perception ability of the model by calculating the interaction between each position in the feature map and all other positions. This is very helpful for the yin and yang detection task of the reagent board. Each small area (reaction position) on the reagent board will have color depth changes. NonLocalBlockND can help the model better capture the color characteristics of these local areas. In addition, there is a certain correlation between multiple areas on the reagent board (for example, the reaction in the quality control area may affect the judgment of other areas). NonLocalBlockND can capture this global contextual information, thereby improving the accuracy of detection. In complex background conditions (such as noise, uneven lighting, etc.), NonLocalBlockND can also help the model better distinguish between the background and target areas and reduce false detections. The internal structure diagram of the NonLocalBlockND module is shown below. Figure 4As shown in Figure 1, the input size of the module is T×H×W×1024, where T represents the number of frames (or time series length, or batch in YOLOv8), H×W is the spatial dimension of the feature map, and 1024 is the number of channels. These modules generate feature representations through 1x1 convolution operations, which can be understood as query vectors, generated key vectors, and value vectors from left to right. All three branches are reduced to 512 channels through 1×1 convolutions, and the output is T×H×W×512. The two branches on the left are then matrix multiplied to calculate the similarity relationship between all positions. They are then normalized through softmax to generate an attention weight matrix. The output is THW×THW, which represents the similarity between each position in the input feature map. This output value is matrix multiplied with the third branch to obtain a weighted global feature representation (size T×H×W×512). The weighted features are restored to the original channel number 1024 through 1x1 convolution, and the residual structure is performed with the original input to output a feature map of size T×H×W×1024. The residual structure can retain the original feature information and superimpose the global weighted features, thereby retaining local details and combining global information.

[0035] In the reagent board detection scenario, since it is a monochrome type of detection, color images cannot extract features of a single color depth very well, and grayscale gradient information will be better for single color level recognition. The pyramid structure originally used for multi-scale detection is used to fuse two feature layers, so that this feature layer takes into account the detection position, and at the same time, it is better for key tasks and distinguishing easily confused color levels.

[0036] In the input stage of the model, the image needs to be preprocessed, including grayscale conversion and adaptive equalization. The color RGB image channel and the grayscale image channel are fused and then input into the model. The structure diagram of the improved multi-channel target detection network model of the present invention is as follows: Figure 5 As shown. The structure mainly consists of the following parts: The neural network in the present invention processes the input image by branch according to the channel. The Backbone part consists of two branches: the first branch is the RGB channel, and the second branch is the grayscale channel. The information of different channels is taken out through the SilenceChannel module (used to separate and select different channel information of the input image, which greatly facilitates the calling and other work of the dual backbone), and is put into two channels respectively. After a series of operations such as convolution, the different features of the image are extracted, and then the two features are fused through the Concat module, which can effectively separate the foreground and background of the task and better capture the target. The two features complement each other and provide more abundant feature information, thereby reducing the possibility of missed detection and false detection.

[0037] Each branch consists of 5 convolutional modules and 4 C2f (i.e., Convolutional-2-Feature modules, which consist of two convolutional layers, multiple bottleneck layers, and residual structures) modules. Convolutional modules are used to extract basic features, such as edges and textures. Each module usually contains one or more convolutional layers, as well as activation functions (such as ReLU) and normalization layers (such as BatchNorm), gradually increasing the abstract level of features and providing rich and diverse feature representations for subsequent modules. C2f modules reduce redundant parameters through optimized design, thereby reducing computational complexity. These modules usually contain multiple convolutional layers, residual connections, and bottleneck structures, which help to transmit more information in the network while reducing the amount of computation.

[0038] In addition, the backbone network adopts the CSPDarkNet (Cross Stage Partial Darknet) structure, uses a series of convolution and deconvolution layers, and introduces residual connections and bottleneck structures to reduce the network size and improve performance. In order to further enhance the feature extraction capability, the backbone network combines technologies such as deep separable convolution and dilated convolution. In the feature fusion part, the features of the two branches are spliced ​​together through a Concat module, and the SPPF (spatial pyramid pooling) module and the NonLocalBlockND (non-local attention mechanism) module are applied to the spliced ​​feature map.

[0039] The NonLocalBlockND module ensures that the model can capture global information while extracting local features, thereby enhancing the understanding of complex scenes. The SPPF module further improves the model's multi-scale feature extraction capabilities through multi-scale pyramid pooling, enabling it to have stronger detection capabilities for targets of different sizes. Through these designs, the entire Backbone part, from preliminary feature extraction (StemLayer) to multi-level feature screening and enhancement (StageLayer), and finally to the SPPF module to enhance the detection capabilities of targets of different scales, forms an efficient and powerful feature extraction framework.

[0040] The Neck part, also known as the feature enhancement network, is responsible for multi-scale feature fusion. It adopts the idea of ​​PAN-FPN (path aggregation network-feature pyramid network) to fuse the feature maps obtained by the backbone network at different stages to enhance the representation ability of the features. The PAA (Progressive Anchor Assignment) module is used to optimize the allocation of anchor boxes and the selection of positive and negative samples, thereby improving the training effect of the model. PAN-FPN contains two PAN modules, which are used for path aggregation of features at different levels, and enhance the representation ability of feature maps through bottom-up and top-down paths.

[0041] In the Neck part of YOLOv8, the data flows as follows Figure 6 As shown. The outputs P3, P4, and P5 of the feature extraction network (Backbone) are input into the PAN-FPN network structure to achieve the fusion of multi-scale feature maps. Through this design, the model can better utilize multi-scale information and improve the detection ability of targets of different sizes. The specific steps are as follows: S1. First, the P5 feature map is upsampled and fused with the P4 feature map to obtain F1.

[0042] S2. Next, F1 passes through a C2f layer and is upsampled again and fused with the P3 feature map to obtain T1.

[0043] S3. Then, T1 is passed through a convolutional layer and fused with the output of F1 to obtain F2.

[0044] S4. Next, F2 passes through the C2f layer to obtain T2.

[0045] S5.T2 goes through another convolution layer and is fused with the P5 feature map to obtain F3.

[0046] S6.Finally, F3 passes through the C2f layer to obtain T3.

[0047] In summary, T1, T2, and T3 are the products of three different processing processes and are important outputs of the entire Neck part. This design allows the network to respond to different feature levels from the bottom up and finally fuse them together to provide a more comprehensive and multi-level feature representation for subsequent detection tasks.

[0048] The head part is responsible for the final detection and classification tasks. It adopts the Decoupled-Head idea to decouple the regression branch and the classification branch, thereby simplifying the model structure. The head includes a detection head and a classification head. The detection head contains a series of convolutional layers and deconvolutional layers to generate detection results. These layers are responsible for predicting the bounding box regression value of each anchor box and the confidence of the target. Through decoupling, the detection head can focus more on the regression task and improve the accuracy of the bounding box prediction. The classification head classifies each feature map through the global pooling layer, and outputs the probability distribution of each category by reducing the dimension of the feature map. This design enables the classification head to focus more on the classification task and improve the accuracy of classification. The head data flow is as follows: Figure 7As shown in the figure, the three outputs of Neck (T1, T2, T3) are feature maps at different levels to capture target information of different scales. The output of each detection head is a feature map, whose size is usually a certain ratio of the input image size. If the input image size is 640x640, then the output feature map sizes of the three detection heads may be 80x80, 40x40, and 20x20 respectively. Each position on each feature map predicts the bounding box, category, and confidence. Finally, they are merged together to form the final result, outputting a four-dimensional vector, which represents the upper left corner coordinates x, y and the lower right corner coordinates x, y of the target box respectively.

[0049] 4. Model Training The image data with completed color gradient annotation is divided into a training set for the training of the target detection model YOLOv8. During the training process, the model parameters are iteratively updated through the back propagation algorithm, including the parameters of the convolution layer, activation layer, and normalization layer. When the network training reaches a certain stage and the accuracy of the verification set is no longer improved, the training process automatically stops and the best performing YOLOv8 target detection model on the verification set is saved.

[0050] 5. Model Deployment Convert the trained model to an ONNX (Open Neural Network Exchange) file format to facilitate the model call. ONNX model related parameters can be viewed using Netron software, such as Figure 8 As shown, the input is a tensor of [1, 4, 640, 640] and the output is a tensor of [1, 12, 8400].

[0051] 6. Post-processing The output of the ONNX model stores the location, range, recognition results and confidence of 840 bounding boxes, which are used as input for the post-processing process. After post-processing, the test results of the reagent board are obtained. The specific steps are: Use the NMS algorithm to select about 24 bounding boxes from the 840 input bounding boxes; According to the maximum range of the bounding box, the image to be identified is evenly divided, and the coordinate range of 24 points (6 rows and 4 columns) of the image to be identified is calculated; At this time, there are generally three situations between the points on the image obtained in step 2 and the 24 or so bounding boxes obtained in step 1, which are processed separately. If there are multiple bounding boxes at a certain point, the bounding box with the highest confidence is selected as the detection result of the current point; if there is no bounding box at a certain point, its result is recognized as 0 (negative); if the number of bounding boxes obtained in step 1 is significantly different from the actual number of 24 points, the recognition fails.

[0052] The results at the 24 points are sorted according to the order on the test kit, which is the identified test results of the reagent board.

[0053] 7. Image recognition process Pass in the image data and crop the image according to the range of the passed in cropping box. Perform grayscale processing on the cropped image and convert it into a grayscale image. Perform adaptive histogram equalization on the grayscale image, then adjust the image size and expand the boundaries so that the image meets the model input requirements. Convert the grayscale image and RGB image data into input tensors (Tensor), and use the ONNX model file to build an inference session (Inference Session). Run the inference session to obtain the coordinates, range, and confidence of the recognition box. Use the NMS (Non-Maximum Suppression) algorithm to post-process and filter the results to obtain the final recognition box and recognition results. Example pictures are as follows Fig. 9 Finally, the corresponding result is generated according to the kit colorimetric card and displayed to the user. The entire recognition process is as follows Fig.10 shown.

[0054] like Fig.11 As shown, the present invention also provides a reagent board test result identification device, the device includes at least one processor and at least one memory, and also includes a communication interface and an internal bus; a computer execution program is stored in the memory; a computer execution program of a reagent board test result identification model constructed by the construction method as described above is stored in the memory; when the processor executes the computer execution program stored in the memory, the processor can be made to execute a reagent board test result identification method. Wherein the internal bus can be an industry standard architecture (IndustryStandard Architecture, ISA) bus, a peripheral component interconnection (Peripheral Component, PCI) bus or an extended industry standard architecture (.XtendedIndustryStandardArchitecture, EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the accompanying drawings of the present application is not limited to only one bus or a type of bus. Wherein the memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and can also be a U disk, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.

[0055] The device may be provided as a terminal, a server or other forms of device. In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0056] The present invention also provides a computer-readable storage medium, in which a computer execution program of the reagent board card detection result recognition model constructed by the construction method as described above is stored. When the computer execution program is executed by a processor, the processor can execute a reagent board card detection result recognition method.

[0057] Specifically, a system, device or equipment equipped with a readable storage medium may be provided, on which a software program code for implementing the functions of any of the above-mentioned embodiments is stored, and a computer or processor of the system, device or equipment reads and executes the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any of the above-mentioned embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0058] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0059] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A method for constructing a reagent card test result recognition model, characterized in that: The process includes: Step 1, obtaining the original image of the reagent card at N points and its cropped image; the cropped image automatically crops the target range to be identified according to prior information; Step 2, preprocessing the image obtained in step 1, wherein the preprocessing includes grayscale image processing, and constructing a training set and a test set together with the cropped color RGB image; Step 3, constructing a reagent board test result recognition model, the model includes a multi-channel target detection network model improved based on YOLOv8 and an NMS algorithm module connected in series therewith; The multi-channel target detection network model includes a backbone network, a Neck network, and a Head network; wherein a non-local attention mechanism module NonLocalBlockND is added before the spatial pyramid pooling module SPPF in the backbone network to capture the relationship between all positions in the input feature map. It better captures the color features of each small reaction area on the reagent board by calculating the interaction between each position and all other positions in the feature map, and captures the global context information of multiple small reaction areas; The NMS algorithm module selects N non-overlapping bounding boxes based on the positions, ranges, recognition results and confidences of the M bounding boxes output by the multi-channel target detection network model, evenly divides the image to be recognized according to the maximum range of the bounding boxes, and calculates the coordinate ranges of the N points of the image to be recognized, thereby obtaining the color depth recognition results of each detection position; Step 4: Train the constructed reagent board test result recognition model and select the model with the best performance as the final model.

2. A method for constructing a reagent card test result recognition model as claimed in claim 1, characterized in that: The multi-channel target detection network model processes the input image by branch according to the channel; the backbone network part is composed of two branches: the first branch is the RGB channel, and the second branch is the grayscale channel; The SilenceChannel module extracts information from different channels and puts them into two channels respectively. After a series of convolution operations, different features of the image are extracted. The two features are then fused through the Concat module to effectively separate the foreground and background of the task and better capture the target. Each branch consists of 5 convolution modules and 4 C2f modules. The convolution module is used to extract basic features, including edges and textures. Each module contains one or more convolution layers, as well as activation functions and normalization layers, gradually increasing the abstract level of features and providing rich and diverse feature representations for subsequent modules. The C2f module reduces redundant parameters through optimized design, and contains multiple convolution layers, residual connections, and bottleneck structures, which helps to transmit more information in the network. At the same time, the backbone network adopts the CSPDarkNet structure, uses a series of convolution and deconvolution layers, and introduces residual connections and bottleneck structures; in the feature fusion part, the features of the two branches are spliced ​​together through a Concat module, and the spatial pyramid pooling SPPF module and the non-local attention mechanism NonLocalBlockND module are applied to the spliced ​​feature map.

3. A method for constructing a reagent card test result recognition model as claimed in claim 1, characterized in that: The input size of the non-local attention mechanism module NonLocalBlockND is T×H×W×1024, where T represents the number of frames or the length of the time series, H×W is the spatial dimension of the feature map, and 1024 is the number of channels; the module generates feature representations through 1x1 convolution operations, which can be understood as query vectors, generated key vectors, and value vectors from left to right. These three branches are all reduced to 512 channels through 1×1 convolutions, and the output is T×H×W×512; then the two branches on the left are matrix multiplied to calculate the similarity relationship between all positions. ; Then it is normalized through softmax to generate an attention weight matrix, and the output is THW×THW, which represents the similarity between each position in the input feature map; the output value is matrix multiplied with the third branch to obtain a weighted global feature representation with a size of T×H×W×512; the weighted features are restored to the original number of channels 1024 through 1x1 convolution, and a residual structure is performed with the original input to output a feature map of size T×H×W×1024; the residual structure can retain the original feature information and superimpose the globally weighted features.

4. A method for constructing a reagent card test result recognition model as claimed in claim 1, characterized in that: The Neck network is responsible for multi-scale feature fusion. It uses the path aggregation network-feature pyramid network PAN-FPN to fuse the feature maps obtained by the backbone network at different stages. PAN-FPN contains two PAN modules for path aggregation of features at different levels, and enhances the representation ability of feature maps through bottom-up and top-down paths. The outputs P3, P4, and P5 of the backbone network are input into the PAN-FPN network structure to realize the fusion of multi-scale feature maps. The details are as follows: First, the P5 feature map is upsampled and fused with the P4 feature map to obtain F1; Next, F1 passes through a C2f layer and is upsampled again and fused with the P3 feature map to obtain T1; Then, T1 is passed through a convolutional layer and fused with the output of F1 to obtain F2; Next, F2 passes through the C2f layer to obtain T2; T2 goes through another convolutional layer and is fused with the P5 feature map to obtain F3; Finally, F3 passes through the C2f layer to obtain T3.

5. A method for constructing a reagent card test result recognition model as claimed in claim 1, characterized in that: The Head network includes a detection head and a classification head; the detection head contains a series of convolutional layers and deconvolutional layers to generate detection results; these layers are responsible for predicting the bounding box regression value of each anchor box and the confidence of the target. Through decoupling, the detection head focuses more on the regression task and improves the accuracy of bounding box prediction; the classification head classifies each feature map through a global pooling layer, and outputs the probability distribution of each category by reducing the dimension of the feature map; The output of each detection head is a feature map whose size is a certain ratio of the input image size; each position on each feature map predicts the bounding box, category, and confidence, and finally merges them together to form the final result, outputting a four-dimensional vector representing the upper left corner coordinates x, y and lower right corner coordinates x, y of the target box.

6. A method for constructing a reagent card test result recognition model as claimed in claim 1, characterized in that: The specific processing process of the NMS algorithm module is as follows: Use the NMS algorithm to select 24 bounding boxes from the input 840 bounding boxes; According to the maximum range of the bounding box, the image to be identified is evenly divided, and the coordinate range of 24 points of the image to be identified, that is, 6 rows and 4 columns, is calculated; At this time, the points on the image obtained and the 24 bounding boxes obtained are processed in three cases: if there are multiple bounding boxes at a certain point, the bounding box with the highest confidence is selected as the detection result of the current point; if there is no bounding box at a certain point, its result is recognized as 0 negative; if the number of bounding boxes obtained is significantly different from the actual number of 24 points, the recognition fails; The results at the 24 points are sorted according to the order on the test kit, which is the identified test results of the reagent board.

7. A method for identifying test results of a reagent board, characterized in that: The reagent board test result recognition model constructed by the construction method according to any one of claims 1 to 6 is converted into ONNX format and deployed to the back-end server, and includes the following process: Step 1, capture and obtain the reagent board image and cropping range parameters; Step 2, calculate the grayscale channel of the image, adjust the size of the grayscale image, normalize the grayscale image, expand the dimension, and add channel processing to form 4-channel picture information to obtain a preprocessed image; Step 3, inputting the preprocessed image into the reagent board test result recognition model; Step 4: post-process the board recognition results, obtain the actual detection results and output them.

8. A reagent card test result identification device, characterized in that: The device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of a reagent board test result recognition model constructed by the construction method according to any one of claims 1 to 6; when the processor executes the computer execution program stored in the memory, the processor executes a reagent board test result recognition method.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer execution program of the reagent card test result recognition model constructed by the construction method according to any one of claims 1 to 6. When the computer execution program is executed by the processor, the processor executes a reagent card test result recognition method.

Citation Information

Cited By

  • Computer board card testing method and system

    CN120743651A