Sample image processing method and device based on multi-channel sample tube recognition model

By performing grid segmentation, adversarial patch removal, and perturbation cleanup on batch sample tube images, combined with a multi-channel recognition model and credibility verification, the problem of low accuracy in batch sample tube recognition was solved, and efficient information storage was achieved.

CN121095104AActive Publication Date: 2025-12-09FUDAN (SHANGHAI) TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511639093.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2025-12-09
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

When acquiring images of batch sample tubes, the uniform acquisition conditions lead to reflections or pixel disturbances in some areas. Differences or occlusions in the style and placement of sample tube labels result in low recognition accuracy and increased time for data entry.

Method used

By segmenting the sample tube images into a grid, a set of single-tube images is generated. Then, adversarial patch removal and adversarial perturbation cleanup are performed. A multi-channel sample tube recognition model is used for recognition, combined with credibility verification, to ensure that the information is accurately stored in the database.

Benefits of technology

It improves the recognition accuracy of batch sample tubes, reduces the time spent on information entry into the database, and effectively reduces background interference and cross-tube occlusion problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095104A_ABST
    Figure CN121095104A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a sample image processing method and device based on a multi-channel sample tube recognition model. A specific embodiment of the method comprises the following steps: performing grid segmentation on a sample tube image to obtain a single tube image set; for each single tube image, the following processing steps are executed: performing adversarial patch removal processing on the single tube image to generate a removed single tube image; performing anti-disturbance cleaning on the removed single tube image to generate a cleaned single tube image; performing sample tube identification on the cleaned single tube image through a multi-channel sample tube identification model to generate sample tube identification information; carrying out credibility verification on the identification information of each sample tube to obtain an information verification result; and in response to determining that the information verification result characterizes that the verification is passed, performing information storage on the identification information of each sample tube and the sample tube image. According to the embodiment, the recognition accuracy of the batch sample tubes can be improved, so that the information storage time consumption of the batch sample tubes is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and particularly to a sample image processing method and device based on a multi-channel sample tube identification model. BACKGROUND

[0002] In laboratory and detection scenarios, batch recognition of whole rows, whole boxes or whole racks of sample tubes is often required to achieve rapid sample warehousing, de-warehousing and inventorying. Currently, when processing batch sample tubes, the commonly used method is to process and identify the collected images containing batch sample tubes through an image recognition model, and to store information.

[0003] However, when the above method is used to process batch sample tubes, the following technical problems often exist: When collecting images of batch sample tubes, it is difficult to simultaneously satisfy the best imaging for each sample tube due to the limitation of uniform collection conditions, so there are some areas of reflected light or pixel disturbance in the collected images of batch sample tubes, and the label style and pasting position of each sample tube are different or blocked, resulting in low recognition accuracy of batch sample tubes, and further increasing the time consumption of information storage. SUMMARY

[0004] The summary part of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments part. The summary part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of the present disclosure propose a sample image processing method, device, electronic equipment and readable medium to solve the technical problems mentioned in the background part.

[0006] In a first aspect, some embodiments of the present disclosure provide a sample image processing method based on a multi-channel sample tube identification model, the method comprising: performing grid segmentation on an obtained sample tube image to obtain a single tube image set, wherein the sample tube image is obtained after image acquisition and transmission of a batch of sample tubes placed on a grid template, and each single tube image in the single tube image set represents a local area of a sample tube and corresponds to grid coordinate information; for each single tube image in the single tube image set, performing the following processing steps: performing an adversarial patch removal process on the single tube image to generate a removed single tube image; performing an adversarial perturbation cleaning process on the removed single tube image to generate a cleaned single tube image; performing sample tube identification on the cleaned single tube image by using a pre-trained multi-channel sample tube identification model to generate sample tube identification information; performing credibility verification on each sample tube identification information to obtain an information verification result; and in response to determining that the information verification result indicates that the verification is passed, performing information storage of each generated sample tube identification information and the sample tube image.

[0007] In a second aspect, some embodiments of the present disclosure provide a sample image processing device based on a multi-channel sample tube identification model, the device comprising: a grid segmentation unit configured to perform grid segmentation on an obtained sample tube image to obtain a single tube image set, wherein the sample tube image is obtained after image acquisition and transmission of a batch of sample tubes placed on a grid template, and each single tube image in the single tube image set represents a local area of a sample tube and corresponds to grid coordinate information; a processing unit configured to, for each single tube image in the single tube image set, perform the following processing steps: perform an adversarial patch removal process on the single tube image to generate a removed single tube image; perform an adversarial perturbation cleaning process on the removed single tube image to generate a cleaned single tube image; perform sample tube identification on the cleaned single tube image by using a pre-trained multi-channel sample tube identification model to generate sample tube identification information; a verification unit configured to perform credibility verification on each sample tube identification information to obtain each information verification result; and a storage unit configured to, in response to determining that the information verification result indicates that the verification is passed, perform information storage of each generated sample tube identification information and the sample tube image.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0010] The above various embodiments of the present disclosure have the following beneficial effects: the sample image processing method based on the multi-channel sample tube recognition model of some embodiments of the present disclosure can improve the recognition accuracy of batch sample tubes, thereby reducing the information storage time of batch sample tubes. Specifically, the reason for the low recognition accuracy of related batch sample tubes and the long information storage time is that, when image acquisition is performed on batch sample tubes, it is difficult to simultaneously satisfy the best imaging for each sample tube due to the limitation of uniform acquisition conditions, so there will be some areas of reflection or pixel disturbance in the acquired batch sample tube images, and the label style and pasting position of each sample tube are different or blocked, thereby resulting in low recognition accuracy of batch sample tubes, and further increasing the information storage time. Based on this, the sample image processing method based on the multi-channel sample tube recognition model of some embodiments of the present disclosure first performs grid segmentation on the obtained sample tube image to obtain a single tube image set. The above sample tube image is obtained by image acquisition and transmission of batch sample tubes placed on a grid template. Each single tube image in the above single tube image set represents a local area of a sample tube and corresponds to grid coordinate information. In this way, the overall sample tube image can be segmented into various local single tube images, so that subsequent processing is performed on single tube images, thereby reducing the problems of background interference and cross-tube blocking. Then, for each single tube image in the above single tube image set, the following processing steps are performed: first, the above single tube image is subjected to an adversarial patch removal process to generate a removed single tube image. In this way, the adversarial patch removal process can reduce the structured blocking of the sample tube barcode area and the text area in the single tube image and the high saturation light. Second, the above removed single tube image is subjected to an adversarial disturbance cleaning process to generate a cleaned single tube image. In this way, the adversarial disturbance cleaning process can suppress pixel interference caused by reflection, pixel-level noise or malicious adversarial disturbance in the single tube image, restore the true texture of the single tube image, and thus enable the subsequent recognition model to still extract stable features even in a high-noise environment. Third, the multi-channel sample tube recognition model pre-trained is used to perform sample tube recognition on the above cleaned single tube image to generate sample tube recognition information. In this way, for different information carriers (such as sample tube barcodes, label texts in labels, and sample tube appearances) in the cleaned single tube image, through the matching input domain and sample tube recognition model channel, the recognition accuracy of different information of the sample tube can be improved. Then, the credibility of each sample tube recognition information is verified to obtain an information verification result. In this way, abnormal sample tube recognition information can be filtered out to avoid the storage of incorrect or low-credibility recognition information. Finally, in response to determining that the above information verification result represents a verification pass, the generated sample tube recognition information and the above sample tube image are stored in the information storage.Therefore, the partial area reflection or pixel disturbance existing in the batch sample tube image can be effectively reduced by using the anti-patch removal and anti-disturbance cleaning, and different information in the single tube image can be identified by the multi-channel sample tube identification model, so as to realize accurate identification of the batch sample tube information, and further reduce the information storage time of the batch sample tube. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0012] Figure 1 is a schematic diagram of one application scenario of a sample image processing method based on a multi-channel sample tube identification model of some embodiments of the present disclosure; Figure 2 is a top view perspective schematic diagram of batch sample tubes placed on a grid template in the sample image processing method based on a multi-channel sample tube identification model according to the present disclosure; Figure 3 is a flowchart of some embodiments of the sample image processing method based on a multi-channel sample tube identification model according to the present disclosure; Figure 4 is a structural schematic diagram of some embodiments of a sample image processing method device based on a multi-channel sample tube identification model according to the present disclosure; Figure 5 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0014] It should also be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0015] It should be noted that the terms "first", "second", and the like in the present disclosure are merely intended to distinguish different devices, modules or units, and do not imply the sequence of the functions performed by these devices, modules or units or the mutual dependency of these devices, modules or units.

[0016] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one" or "multiple" should be understood as "one or more" unless otherwise explicitly indicated in the context.

[0017] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0019] Figure 1 is a schematic diagram of one application scenario of a sample image processing method based on a multi-channel sample tube identification model of some embodiments of the present disclosure.

[0020] In Figure 1In the application scenario, first, the computing device 107 can perform grid segmentation on the obtained sample tube image 106 to obtain a single tube image set. The sample tube image 106 is obtained by image acquisition and transmission of a batch of sample tubes 101 placed on a grid template. Each single tube image in the single tube image set represents a local area of a sample tube (for example, the single tube image 108) and corresponds to grid coordinate information. The batch of sample tubes 101 can be composed of at least one row of sample tubes, and each row of sample tubes is placed on a sample tube rack. Taking a single sample tube 102 as an example, each sample tube includes a sample tube barcode 103 and a sample tube label 104. The image acquisition device 105 can acquire images of the batch of sample tubes 101 at a 45-degree overhead view to obtain the sample tube image 106. For the single tube image in the single tube image set (taking the single tube image 108 obtained after segmentation as an example), the computing device 107 performs the following processing steps: first, the single tube image 108 is subjected to an adversarial patch removal process to generate a removed single tube image. Second, the removed single tube image is subjected to an adversarial perturbation cleaning process to generate a cleaned single tube image. Third, a pre-trained multi-channel sample tube recognition model is used to recognize the cleaned single tube image to generate sample tube recognition information 109 corresponding to the single tube image 108. Then, the computing device 107 can verify the credibility of the generated sample tube recognition information 109 to obtain an information verification result. Finally, the computing device 107 can store the generated sample tube recognition information 109 and the sample tube image 106 in the target database 110 in response to determining that the information verification result indicates that the verification is passed.

[0021] Figure 2 is a schematic diagram of an overhead view of a batch of sample tubes placed on a grid template in a sample image processing method based on a multi-channel sample tube recognition model according to the present disclosure.

[0022] In practice, the batch of sample tubes can be placed on the grid template according to the positions of the individual sample tubes. A schematic diagram of an overhead view of a batch of sample tubes placed on a grid template is shown in Figure 2 Figure 2 which includes sample tubes 102, a test tube rack 201, and black grids 202 in the grid template. One black grid in the grid template can correspond to one sample tube. The distance between each black grid in the grid template can be set according to the distance between each sample tube in the actual sample tube rack.

[0023] ​It should be noted that the computing device 107 described above can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules, such as to provide distributed services, or as a single software or software module. Herein, no specific limitation is made. It should be understood that Figure 1 The number of computing devices in the system can have any number according to the needs of implementation. With reference to Figure 3 , a flow 300 of some embodiments of a sample image processing method based on a multi-channel sample tube recognition model according to the present disclosure is shown. The sample image processing method based on the multi-channel sample tube recognition model includes the following steps: Step 301, performing grid segmentation on the obtained sample tube image to obtain a single tube image set.

[0024] In some embodiments, the execution subject (for example, the computing device 107 shown in Figure 1 ) of the sample image processing method based on the multi-channel sample tube recognition model can obtain the sample tube image by wired or wireless connection and perform grid segmentation on the obtained sample tube image to obtain a single tube image set. The sample tube image can be obtained by image acquisition equipment performing image acquisition and transmission on a batch of sample tubes placed on a grid template. Each single tube image in the single tube image set represents a local area of a sample tube and corresponds to grid coordinate information. The grid template can be a layout template composed of multiple rows of black grids. The number of rows of black grids in the grid template can be adjusted according to the number of sample tubes in the batch of sample tubes. For example, the grid template includes 6 black grids in each row, and the batch of sample tubes includes 22 sample tubes, so the grid template can be adjusted to 4 rows (i.e., the ceiling of the ratio of the number of sample tubes to the number of grids in each row). The batch of sample tubes can be composed of at least one row of sample tubes, and each row of sample tubes is placed on a sample tube rack. The image acquisition equipment can be a camera. The grid coordinates corresponding to each single tube image can represent the position of the corresponding sample tube in the grid template. For example, the grid coordinates corresponding to a single tube image are (2, 6), which can represent the corresponding sample tube in the second row and the sixth column of the grid template. Each sample tube in the batch of sample tubes has a sample tube label and a sample tube barcode attached thereto. The sample tube image can be an RGB image.

[0025] It should be noted that one black grid in the grid template can correspond to the placement of one sample tube. The distance between each black grid in each row of the grid template can be set according to the distance between each sample tube in the actual sample tube rack, which is not limited herein. In actual business scenarios, the distance between adjacent sample tubes in the sample tube rack is fixed and the same, so the distance between adjacent black grids in each row of the grid template is also the same. The distance between each black grid in each row of the grid template can be flexibly adjusted according to a preset layout condition. The preset layout condition can be that, under a preset image acquisition angle, the sample tube barcodes and sample tube labels of the sample tubes in the front row (i.e., the side close to the image acquisition device) do not block the sample tubes in the rear row (i.e., the side away from the image acquisition device). During image acquisition, the batch of sample tubes placed on the grid template are located at the center of the picture, and the image acquisition device can acquire images at the preset image acquisition angle, obtaining sample tube images. As an example, the preset image acquisition angle can be 45 degrees.

[0026] In practice, the execution subject can use a segmentation template pre-generated according to the grid template and the preset image acquisition angle to perform grid segmentation on the obtained sample tube images, obtaining each single tube image as a single tube image set. The segmentation template can be pixel-level cropping information formed by mapping the grid template to the actual pixel plane according to the preset image acquisition angle, including the pixel center, pixel width and height, and cropping frame coordinates (the cropping frame size can be set according to the image size of a single sample tube mapped to the actual pixel plane according to the preset image acquisition angle) of each black grid in the image, which can be directly used for image segmentation.

[0027] It should be noted that the wireless connection mode can include, but is not limited to, 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other now known or future developed wireless connection modes.

[0028] Step 302, for each single tube image in the single tube image set, the following processing steps are performed: Step 3021, performing adversarial patch removal processing on the single tube image to generate a removed single tube image.

[0029] In some optional implementations of some embodiments, the execution subject can perform adversarial patch removal processing on the single tube image to generate a removed single tube image by the following steps: First, normalizing the single tube image to generate a single tube image tensor. In practice, the execution subject can normalize the pixel values in the single tube image one by one to generate a single tube image tensor.

[0030] In the second step, the single tube image tensor is input into a lightweight image recognizer to generate removal image recognition information. The removal image recognition information represents the reliability of the image recognition result. The lightweight image recognizer can be a model that takes the single tube image tensor as input and outputs the removal image recognition information, and is used to identify information in the single tube image. The lightweight image recognizer can include a barcode decoding tool and a text recognition network. As an example, the barcode decoding tool can include various decoding tools (e.g., pyzbar decoding tool, zxing-cpp decoding tool) for identifying barcode information in the image. The text recognition network can include a DBNet (Differentiable Binarization Network) network and a CRNN (Convolutional Recurrent Neural Network) network, which are respectively used to detect and identify text information in the predicted single tube image. The removal image recognition information includes decoding confidence and text confidence. The decoding confidence can be a numerical value determined by the barcode decoding tool, representing the consistency of the stripe decoding. The text confidence can be the average of the maximum character probability of each time of the identified text sequence output by the text recognition network.

[0031] It should be noted that the decoding confidence is the ratio of the number of consistent decodings to the total number of decodings. In practice, the execution subject can rotate the single tube image tensor by a preset rotation angle for a preset number of times, and perform stripe decoding on the single tube image tensor obtained after each rotation in parallel through each decoder included in the barcode decoder, and verify whether the decoding information output by each decoder is consistent, thereby determining the number of consistent decodings (i.e., the number of times when the decoding information is all the same). When decoding in parallel, if one decoder fails to decode, it is determined that this decoding is inconsistent. As an example, the preset rotation angle is 15 degrees, and the preset number of times can be 10.

[0032] In the third step, the single tube image is subjected to patch removal processing according to the generated removal image recognition information to generate a single tube image after removal.

[0033] In some optional implementations of some embodiments, the execution subject can perform patch removal processing on the single tube image according to the generated removal image recognition information to generate a single tube image after removal by the following steps: In the first step, based on the single tube image, the following patch removal steps are performed: The first sub-step is to generate a masked single tube image according to the generated image mask and the single tube image. The generated masked single tube image is a single tube image with a part of the area covered by the mask. In practice, the execution subject can cover the area in the single tube image with the generated image mask to obtain the masked single tube image.

[0034] It should be noted that when performing the above patch removal step for the first time, the execution subject can randomly generate a mask with a random size (for example, 16x16 patch) as an image mask to cover a random area in the single tube image.

[0035] The second sub-step is to perform image reconstruction on the masked single tube image according to a pre-constructed image reconstruction model to obtain a region reconstruction image. The region reconstruction image can be a reconstruction image of the area covered by the mask. The image reconstruction model can be a neural network model that takes the masked single tube image as input and outputs the region reconstruction image, which is used to reconstruct the image area covered by the mask. As an example, the image reconstruction model can be a MAE (Masked Autoencoder) model. For example, the image reconstruction model can be a ViT (Visual Transformer) model or a U-Net model.

[0036] The third sub-step is to perform region removal and replacement on the single tube image by using the region reconstruction image to obtain a replaced single tube image. In practice, the execution subject can replace the area in the single tube image covered by the image mask with the region reconstruction image to obtain the replaced single tube image.

[0037] The fourth sub-step is to generate a pixel residual image corresponding to the replaced single tube image. In practice, the execution subject can perform pixel-by-pixel subtraction between the replaced single tube image and the single tube image to obtain the pixel residual image.

[0038] The fifth sub-step is to generate removal image recognition information according to the replaced single tube image and the lightweight image recognizer. In practice, the execution subject can input the replaced single tube image into the lightweight image recognizer to generate removal image recognition information corresponding to the replaced single tube image.

[0039] In the sixth sub-step, the image mask is generated according to the generated pixel residual image and the preset coverage coefficient sequence, wherein the generated image mask corresponds to the high residual area of the residual image. The preset coverage coefficient in the preset coverage coefficient sequence can be the area proportion of the single tube image covered by the image mask corresponding to the execution times. The preset coverage coefficients in the preset coverage coefficient sequence are in increasing order, indicating that the coverage proportion of the image mask is increased during the execution of the patch removal step. In practice, first, the execution subject can determine each high residual area in the pixel residual image by the residual threshold and generate the area mask corresponding to each high residual area. The high residual area can be a pixel residual value area whose each pixel residual value is higher than the residual threshold and the number of pixel residual values contained is greater than or equal to the area pixel number threshold. Then, the execution subject can select the preset coverage coefficient corresponding to the bit sequence as the target coverage coefficient from the preset coverage coefficient sequence according to the execution times of the patch removal step. After that, the execution subject can generate a random mask in addition to each area mask according to the target coverage coefficient. Finally, the execution subject can determine the generated each area mask and the random mask as the image mask.

[0040] In the second step, in response to determining that the generated each removal image recognition information does not satisfy the effective removal condition, the replaced single tube image is determined as the removed single tube image. The effective removal condition can be that the text confidence and the decoding confidence included in the removal image recognition information generated in this execution of the patch removal step are less than or equal to the text confidence and the decoding confidence included in the removal image recognition information generated in the last execution of the patch removal step.

[0041] Optionally, the patch removal step can further include the following step: in response to determining that the generated each removal image recognition information satisfies the effective removal condition, the replaced single tube image is taken as the single tube image, and the patch removal step is executed again.

[0042] The above-mentioned related content serves as one of the invention points of the present disclosure, and solves the technical problem that in unified image acquisition and image transmission of batch sample tubes, there are often structured patch interference (such as reflection, blur and occlusion) and large imaging noise interference, which reduces the clarity of the sample tube barcode or text in the image, and generates false edges and adversarial textures, thereby causing the image quality to decrease. If the above factors are solved, the effect of improving the image quality can be achieved. In order to achieve this effect, through the prior steps of multiple mask coverings of the single tube image and reconstruction of the covered area, the covering of the originally existing reflection, blur and occlusion in the image on the sample tube image is gradually reduced, the barcode structure of the sample tube barcode and the text coherence in the sample tube label are restored, and the optimal imaging conditions of the corresponding single sample tube are simulated through multiple mask coverings and reconstruction of the covered area, thereby greatly reducing the interference of the stripes or moire in the single tube image, and further improving the image quality of the single tube image.

[0043] Step 3022, performing adversarial disturbance cleaning on the removed single tube image to generate a cleaned single tube image.

[0044] In some optional implementations of some embodiments, the above-mentioned execution subject can perform adversarial disturbance cleaning on the above-mentioned removed single tube image to generate a cleaned single tube image by the following steps: First, generating a noise distribution map and an image multi-dimensional feature matrix according to the above-mentioned removed single tube image.

[0045] Second, based on the image multi-dimensional feature matrix, the following disturbance cleaning steps are performed: First sub-step, generating a denoised single tube image according to the preset texture similarity number, the above-mentioned image multi-dimensional feature matrix, the above-mentioned removed single tube image and the above-mentioned noise distribution map.

[0046] Second sub-step, generating denoised image recognition information according to the denoised single tube image and the above-mentioned lightweight image recognizer. In practice, the above-mentioned execution subject can input the denoised single tube image into the above-mentioned lightweight image recognizer to obtain the denoised image recognition information. The above-mentioned denoised image recognition information also includes decoding confidence and text confidence.

[0047] Third sub-step, determining the image confidence according to the removed image recognition information corresponding to the above-mentioned removed single tube image and the denoised image recognition information. In practice, the above-mentioned execution subject can subtract the decoding confidence and the text confidence included in the denoised image recognition information from the decoding confidence and the text confidence included in the removed image recognition information to obtain the decoding confidence difference and the text confidence difference as the image confidence.

[0048] A fourth sub-step, in response to determining that the image credibility satisfies the iterative optimization condition, updating the preset texture similarity quantity according to the target residual image, and executing the disturbance cleaning step again. Wherein, the generated target residual image is determined by the denoised single tube image and the removed single tube image. The iterative optimization condition can be that the decoding confidence difference and the text confidence difference included in the image credibility are both positive values, indicating that this time of executing the disturbance cleaning step improves the image quality and optimizes the removed single tube image. In practice, first, the execution subject can determine the high residual area quantity in the target residual image and the pixel residual image (i.e. the pixel residual image generated in the last execution of the patch removal step) corresponding to the removed single tube image respectively as the first high residual area quantity and the second high residual area quantity through the residual threshold. Then, the normal state can add the absolute value of the difference between the second high residual area quantity and the first high residual area quantity to the preset texture similarity quantity to update the preset texture similarity quantity.

[0049] A fifth sub-step, in response to determining that the image credibility does not satisfy the iterative optimization condition, determining the denoised single tube image as the cleaned single tube image.

[0050] In some optional implementations of some embodiments, the execution subject can generate a noise distribution map and an image multi-dimensional feature matrix according to the removed single tube image by the following steps: First, normalize the removed single tube image to generate a removed single tube image tensor. In practice, the execution subject can normalize the pixel values in the removed single tube image one by one to generate a removed single tube image tensor.

[0051] Second, determine the noise of the removed single tube image tensor pixel by pixel to generate a noise distribution map, wherein the noise distribution map has the same size as the removed single tube image tensor. In practice, first, the execution subject can first perform inverse gamma transform (for example, the gamma parameter is 2.2) on the removed single tube image tensor channel by channel, map the pixel from sRGB domain to linear domain, and get a linear domain image. Then, the execution subject can determine the mean and variance of the linear domain image using a sliding window to fit the Poisson-Gaussian noise model (for example, σ²(x)=αx+β), and determine the noise standard deviation σ(x, y) by pixel, and normalize it to [0, 1] to get the noise distribution map.

[0052] Thirdly, the removed single tube image and the noise distribution map are synchronously blocked according to a preset block size and a preset step size, to obtain a removed single tube image block set and a noise distribution map block set. In practice, the execution subject can take the block size as a window size, and perform sliding window processing on the removed single tube image and the noise distribution map respectively according to the preset step size, to obtain the removed single tube image block set and the noise distribution map block set.

[0053] Fourthly, each block multi-dimensional feature information is generated according to the removed single tube image block set and the noise distribution map block set. The generated block multi-dimensional feature information corresponds to the removed single tube image block.

[0054] Fifthly, an image multi-dimensional feature matrix is determined according to the generated each block multi-dimensional feature information. In practice, the execution subject can arrange the each block multi-dimensional feature information in sequence according to the position and sequence of each removed single tube image block in the removed single tube image, to obtain the image multi-dimensional feature matrix.

[0055] In some optional implementations of some embodiments, the execution subject can generate each block multi-dimensional feature information according to the removed single tube image block set and the noise distribution map block set by the following steps: Firstly, for each removed single tube image block in the removed single tube image block set, the following steps are performed: Firstly, the removed single tube image block is subjected to frequency domain feature extraction to generate image frequency domain feature information. In practice, first, the execution subject can perform two-dimensional discrete cosine transform on the removed single tube image in the luminance channel in units of 16x16 sub-images, to obtain a coefficient matrix. Then, the execution subject can extract K low-frequency coefficients (for example, between 8 and 16 dimensions) from the top left corner in sequence in the coefficient matrix as the image frequency domain feature information.

[0056] Secondly, the removed single tube image block is subjected to spatial domain feature extraction to generate image spatial domain feature information. In practice, the execution subject can determine the Sobel gradient in the luminance channel of the removed single tube image block to obtain a gradient histogram. Then, the execution subject can determine the luminance local variance and average gradient amplitude of the removed single tube image block from the gradient histogram as the image spatial domain feature information.

[0057] A third sub-step is to perform lightweight coding on the removed single tube image block to generate image coding feature information. In practice, the execution subject can input the removed single tube image block into a convolution network composed of three convolution layers to generate the image coding feature information. Each convolution layer includes a 3x3 convolution kernel, a ReLU activation function, and a BN layer.

[0058] A fourth sub-step is to generate image noise feature information according to the noise distribution block corresponding to the removed single tube image block. In practice, the execution subject can crop the area corresponding to the removed single tube image block from the noise distribution map, and determine the noise mean and noise variance of the cropped area as the image noise feature information.

[0059] A fifth sub-step is to sequentially splice the generated image frequency domain feature information, image spatial domain feature information, image coding feature information, and image noise feature information to obtain patch multi-dimensional feature information.

[0060] In some optional implementations of some embodiments, the execution subject can generate a denoised single tube image according to the preset texture similarity number, the image multi-dimensional feature matrix, the removed single tube image, and the noise distribution map by the following steps: For each two patch multi-dimensional feature information in the image multi-dimensional feature matrix, the execution subject can determine the cosine similarity between the two patch multi-dimensional feature information as the feature similarity.

[0061] For each patch multi-dimensional feature information in the image multi-dimensional feature matrix, the execution subject can perform the following steps: A first sub-step is to sort each feature similarity corresponding to the patch multi-dimensional feature information to obtain a feature similarity sequence. In practice, the execution subject can sort each feature similarity corresponding to the patch multi-dimensional feature information (i.e., related) from large to small to obtain a feature similarity sequence.

[0062] A second sub-step is to select, according to the feature similarity sequence, patch multi-dimensional feature information satisfying the preset texture similarity number in the image multi-dimensional feature matrix as each connected patch feature information. In practice, first, the execution subject can select the first preset texture similarity number of each feature similarity from the feature similarity sequence as each target feature similarity. Then, the execution subject can determine the patch multi-dimensional feature information corresponding to each target feature similarity as each connected patch feature information. As an example, the preset texture similarity number can be 8.

[0063] The third sub-step is to determine the connection weight between each connected patch feature information. In practice, for each connected patch feature information, first, the above execution subject can determine the noise distribution mean of the connected patch feature information and the patch multi-dimensional feature information in the corresponding area in the noise distribution map. Then, the above execution subject can determine the exponential function value of the difference between the two noise distribution means as the noise similarity weight through the exponential function (for example, the exp() function in Python). Then, the above execution subject can determine the exponential function value corresponding to the Euclidean distance between the connected patch feature information and the patch multi-dimensional feature information as the feature weight through the exponential function (for example, the exp() function in Python). Finally, the above execution subject can determine the product of the feature weight and the noise similarity weight as the connection weight between the patch multi-dimensional feature information and the connected patch feature information. In this way, through the mapping of the exponential function, the sub-blocks with smaller Euclidean distance between them can get a weight value close to 1, while the sub-blocks with larger Euclidean distance between them tend to have a weight close to 0, thereby realizing the effect of "the more similar the texture, the greater the weight; the greater the texture difference, the smaller the weight".

[0064] The third step is to construct a feature adjacency graph with the patch multi-dimensional feature information in the above image multi-dimensional feature matrix as the node and the determined connection weight as the edge length. When there is a connection weight between any two patch multi-dimensional feature information in the above image multi-dimensional feature matrix, there is a connection relationship between the two nodes, and if there is no connection weight, there is no connection relationship between the two nodes.

[0065] The fourth step is to perform channel splicing on the above removed single-tube image and the above noise distribution map to obtain a spliced single-tube image. In practice, the above execution subject can splice the above noise distribution map as a channel to the above removed single-tube image to obtain a spliced single-tube image through a related library function (for example, the concat() function in Python).

[0066] The fifth step is to perform local feature extraction on the above spliced single-tube image to generate image local feature information. In practice, the above execution subject can perform local feature extraction on the above spliced single-tube image through a convolutional network composed of four convolutional layers to generate image local feature information. Each convolutional layer in the convolutional network includes a 3x3 convolutional kernel and a ReLU activation function, with a convolution step of 1 and a padding of 1.

[0067] In the sixth step, the non-local feature information of the image is generated by performing non-local feature extraction on the feature adjacency graph based on the image multi-dimensional feature matrix. The dimension of the non-local feature information of the image is aligned with the dimension of the local feature information of the image. In practice, the execution subject can input the image multi-dimensional feature matrix as the graph node feature into the graph neural network, perform information propagation on the graph feature adjacency graph, and obtain the node embedding (i.e., the non-local feature information of the image).

[0068] Specifically, the execution subject can use two layers of graph convolutional layers as the graph neural network, and perform weighted aggregation using the pre-calculated connection weight (i.e., edge length). Then, the execution subject can perform overlap up-sampling on each node embedding according to the corresponding position of the removed single-pipe image block and backfill it to the image coordinate domain to generate a non-local feature map consistent in size and alignment with the non-local feature information of the image as the non-local feature information of the image. As an example, the two layers of graph convolutional layers can be composed of two GCNConv (Graph Convolutional Network) layers or GATConv (Graph Attention Network) layers. In addition, when backfilling, weighted averaging or Hanning window normalization can be performed on the overlapping area to eliminate the block boundary effect.

[0069] In the seventh step, the denoised single-pipe image is obtained by fusing and decoding the non-local feature information of the image and the local feature information of the image through the pre-constructed decoder. In practice, the execution subject can concatenate the non-local feature information of the image and the local feature information of the image in the channel dimension to obtain the fused feature information. Then, the execution subject can input the fused feature information into the decoding network (i.e., the decoder) to generate a residual map, and perform residual learning reconstruction (i.e., subtract the removed single-pipe image from the residual map) on the residual map to generate the denoised single-pipe image. The decoding network (i.e., the decoder) can be composed of one layer of convolutional layer containing 3x3 convolution and ReLU activation function, each residual block (each residual block is composed of two layers of 3x3 convolution and identity skip connection), and one layer of convolutional layer containing 3x3 convolution kernel. As an example, the number of residual blocks can be 4.

[0070] The above first step to the seventh step is related to the content of the present disclosure as one of the application points, which solves the technical problem that "in the uniform image acquisition condition, the single tube image often contains fine-grained disturbance and real noise (such as stripes or moire, photosensitive noise, etc.), and these disturbances are repeated across positions and overlap with the fine line band of the sample tube label text and barcode in the image, resulting in reduced image quality and reduced recognition accuracy". If the above factors are solved, the effect of improving the image quality and information recognition accuracy can be achieved. In order to achieve this effect, first, the noise level map of the cleaned single tube image is estimated, the image blocks with similar textures are connected through the near neighbor graph, and the information is propagated on the graph to suppress the fine-grained disturbance such as stripes / moire repeated across positions. Then, local convolution is used to suppress random noise points and preserve the pixel details of the text structure and barcode edge in the image. Finally, the noise components that should be removed from the sample tube image are removed by residual decoding, and the effective details are preserved. Thus, the adversarial disturbance and real noise in the image can be significantly weakened, the sample tube text strokes and barcode module boundaries are clearer, the decoding success rate of the barcode and the text recognition rate are improved, and the image quality and image information recognition accuracy are improved.

[0071] In step 3023, the pre-trained multi-channel sample tube recognition model is used to recognize the sample tube of the cleaned single tube image to generate sample tube recognition information.

[0072] In some embodiments, the above execution subject can use the pre-trained multi-channel sample tube recognition model to recognize the sample tube of the cleaned single tube image to generate sample tube recognition information. The sample tube recognition information includes stripe code decoding information, sample tube appearance information, and sample tube text information. The stripe code decoding information, sample tube appearance information, and sample tube text information correspond to recognition confidence, which represents the accuracy of the recognized content. The stripe code decoding information can be the information obtained by decoding the stripe code on the sample tube. The sample tube appearance information can be the appearance attribute (such as color, cap type, tube type) of the sample tube. The sample tube text information can be the text content on the sample tube label.

[0073] It should be noted that the determination method of the recognition confidence corresponding to the sample tube barcode recognition content can refer to the determination method of the decoding confidence, which will not be repeated here.

[0074] In some optional implementations of some embodiments, the above execution subject can use the pre-trained multi-channel sample tube recognition model to recognize the sample tube of the cleaned single tube image to generate sample tube recognition information by the following steps: First, the single tube image after cleaning is pre-processed to generate a text channel image tensor, an appearance channel image tensor, and a barcode channel image tensor. The multi-channel sample tube identification model includes a text detection and recognition model, an appearance attribute detection model, and a stripe decoding model. In practice, first, the execution subject can perform pixel-by-pixel normalization on the single tube image after cleaning to obtain a single tube image tensor after cleaning. Then, the execution subject can take the image tensor after the single tube image tensor after cleaning is grayed (for example, 0.299R+0.587G+0.114B) as the text channel image tensor. After that, the execution subject can take the image tensor after the single tube image tensor after cleaning is processed by contrast enhancement (for example, adaptive histogram equalization) as the barcode channel image tensor. Finally, the execution subject can take the image tensor after the single tube image tensor after cleaning is processed by slight gamma correction (for example, gamma parameter is 0.9 to 1.1) as the barcode appearance channel image tensor.

[0075] Second, according to the text channel image tensor and the text detection and recognition model, the sample tube text information is generated. The text detection and recognition model can be a neural network model for detecting and recognizing text content existing in the image. The text detection and recognition model includes a text detection model and a text recognition model. As an example, the text detection model can be but not limited to a DBNet (Differentiable Binarization Network) model or a CTPN (Connectionist Text Proposal Network) model. The text recognition model can be a CRNN (Convolutional Recurrent Neural Network) model or a ViT-Seq2Seq (Vision Transformer based Sequence-to-Sequence Model) model. In practice, the execution subject can input the text channel image tensor into the text detection and recognition model to generate the sample tube text information.

[0076] Thirdly, generating sample tube appearance information according to the appearance channel image tensor and the appearance attribute detection model. The appearance attribute detection model can be a multi-classification neural network model for detecting sample tube appearance attributes (such as color, cap type, and tube type). As an example, the appearance attribute detection model can be, but is not limited to, a YOLO series model, a ResNet model, or a Transformer model. In practice, the execution subject can input the appearance channel image tensor into the appearance attribute detection model to generate sample tube text information.

[0077] Fourthly, generating sample tube stripe code decoding information according to the barcode channel image tensor and the stripe decoding model. The stripe decoding model can be a model for decoding stripe codes in images. The stripe decoding model can include various decoding tools, and the barcode channel image tensor is rotated by a preset rotation angle for a preset number of times, each rotated image tensor is decoded in parallel, and the number of consistent decoding is determined. The stripe decoding model can include, but is not limited to, zxing-cpp decoding tool, zbar decoding tool, and pyzbar decoding tool. In practice, the execution subject can input the barcode channel image tensor into the stripe decoding model to obtain sample tube stripe code decoding information output by each decoding tool included in the stripe decoding model (with the corresponding sample tube stripe code decoding information with the highest consistency rate as the output) and the corresponding recognition confidence (i.e. the ratio of the number of consistent decoding to the total number of decoding).

[0078] Fifthly, determining the sample tube text information, the sample tube appearance information, and the sample tube stripe code decoding information as sample tube recognition information.

[0079] Step 303, verifying the credibility of each sample tube recognition information to obtain an information verification result.

[0080] In some embodiments, the execution subject can perform credibility verification on each sample tube identification information to obtain an information verification result. The information verification result can be a Boolean variable. For example, when the information verification result is TRUE, it means that the verification is passed; when the information verification result is FALSE, it means that the verification is failed. In practice, first, for each sample tube identification information, the execution subject can determine whether the recognition confidence of the stripe code decoding information, the sample tube appearance information and the sample tube text information included in each sample tube identification information is greater than or equal to a preset confidence threshold to obtain a single information verification result. The single information verification result can be a Boolean variable. For example, when the single information verification result is TRUE, it means that the recognition confidence of the sample tube barcode recognition content, the sample tube appearance attribute and the sample tube label content included in the corresponding sample tube identification information are all greater than or equal to the preset confidence threshold; when the single information verification result is FALSE, it means that at least one of the recognition confidence of the sample tube barcode recognition content, the sample tube appearance attribute and the sample tube label content included in the corresponding sample tube identification information is less than the preset confidence threshold. As an example, the preset confidence can be 0.989. Finally, the execution subject can perform a Boolean AND operation on each single information verification result to obtain the information verification result. Thus, when there is a sample tube identification information included in the recognition confidence less than the preset confidence threshold, each sample tube identification information of the batch of sample tubes will be determined as failed verification, and the sample tube image corresponding to the batch of sample tubes can be processed again or verified by a human.

[0081] At step 304, in response to determining that the information verification result indicates that the verification is passed, the generated each sample tube identification information and the sample tube image are stored in a database.

[0082] In some embodiments, the execution subject can store the generated each sample tube identification information and the sample tube image in a database in response to determining that the information verification result indicates that the verification is passed. In practice, after the information verification is passed, the execution subject can store each sample tube identification information and the sample tube image in a target database. The target database can be a database for storing sample tube identification information. Further referring to Figure 4 , as an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a sample image processing device based on a multi-channel sample tube identification model. The device embodiments correspond to the method embodiments shown in Figure 3 , and the sample image processing device based on the multi-channel sample tube identification model can be applied in various electronic devices.

[0083] AsFigure 4 As shown, a sample image processing device 400 based on a multi-channel sample tube recognition model in some embodiments includes: a grid segmentation unit 401, a processing unit 402, a verification unit 403, and a storage unit 404. The grid segmentation unit 401 is configured to perform grid segmentation on the acquired sample tube images to obtain a set of single-tube images. The sample tube images are obtained by acquiring and transmitting images of a batch of sample tubes placed on a grid template. Each single-tube image in the set represents a local region of a sample tube and corresponds to grid coordinate information. The processing unit 402 is configured to perform the following processing steps on each single-tube image in the set: perform adversarial patch removal processing on the single-tube image to generate a removed single-tube image; perform adversarial perturbation cleaning on the removed single-tube image to generate a cleaned single-tube image; perform sample tube recognition on the cleaned single-tube image using a pre-trained multi-channel sample tube recognition model to generate sample tube recognition information; the verification unit 403 is configured to perform credibility verification on the recognition information of each sample tube to obtain the verification results of each piece of information; and the storage unit 404 is configured to store the generated sample tube recognition information and the sample tube images in a database in response to determining that the verification results indicate that the verification has passed.

[0084] It is understandable that the units described in the sample image processing apparatus 400 based on the multi-channel sample tube recognition model are similar to the reference units. Figure 3 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method are also applicable to the sample image processing device 400 based on the multi-channel sample tube recognition model and the units contained therein, and will not be repeated here. The following is for reference. Figure 5 It illustrates electronic devices suitable for implementing some embodiments of the present disclosure (such as...). Figure 1 A schematic diagram of the structure of the computing device 107 shown. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0085] like Figure 5 As shown, the electronic device 500 may include a processing unit 501 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 502 or a program loaded from the storage device 508 into the random access memory 503. The random access memory 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the read-only memory 502, and the random access memory 503 are interconnected via a bus 504. An input / output interface 505 is also connected to the bus 504.

[0086] Generally, the following devices can be connected to the input / output interface 505: input devices 506, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 507, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 508, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 5 Each block shown in the flowcharts can represent a device or multiple devices as needed.

[0087] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the read-only memory 502. When the computer program is executed by the processing devices 501, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.

[0088] In one embodiment, the above processing device is configured to run a computer program stored in the memory to perform the following steps: performing grid segmentation on an acquired sample tube image to obtain a single tube image set, wherein the sample tube image is obtained by image acquisition and transmission of a batch of sample tubes placed on a grid template, each single tube image in the single tube image set represents a local area of a sample tube and corresponds to grid coordinate information; for each single tube image in the single tube image set, performing the following processing steps: performing an adversarial patch removal process on the single tube image to generate a removed single tube image; performing an adversarial perturbation cleaning process on the removed single tube image to generate a cleaned single tube image; performing sample tube recognition on the cleaned single tube image through a pre-trained multi-channel sample tube recognition model to generate sample tube recognition information; performing credibility verification on each sample tube recognition information to obtain an information verification result; in response to determining that the information verification result indicates that the verification is passed, storing the generated each sample tube recognition information and the sample tube image in a database.

[0089] The embodiment of the present disclosure further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, the computer program comprises program instructions, and the method realized by executing the program instructions can refer to each embodiment of the method of the present disclosure.

[0090] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like.

[0091] It should be noted that, in the present document, the term “comprising” or “including” or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without more limitations, the element defined by the statement “comprising a” does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0092] The above description is merely some preferred embodiments of the present disclosure and a description of the principles of the applied technology. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features and the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) are replaced with each other to form a technical solution.

Claims

1. A sample image processing method based on a multi-channel sample tube recognition model, characterized in that, include: The acquired sample tube images are segmented into a grid to obtain a set of single tube images. The sample tube images are obtained by acquiring and transmitting images of a batch of sample tubes placed on a grid template. Each single tube image in the set of single tube images represents a local area of ​​a sample tube and corresponds to grid coordinate information. For each single tube image in the single tube image set, perform the following processing steps: The single-tube image is subjected to adversarial patch removal processing to generate a single-tube image after removal; The removed single-tube image is cleaned against disturbances to generate a cleaned single-tube image; The sample tube identification is performed on the cleaned single tube image using a pre-trained multi-channel sample tube recognition model to generate sample tube identification information. The credibility of the identification information for each sample tube is verified to obtain the information verification results. In response to the determination that the information verification result indicates that the verification is successful, the generated identification information of each sample tube and the sample tube image are stored in the database.

2. The method according to claim 1, characterized in that, The step of performing adversarial patch removal processing on the single-tube image to generate a single-tube image after removal includes: The single-tube image is normalized to generate a single-tube image tensor; The single-tube image tensor is fed into a lightweight image recognizer to generate removed image recognition information, wherein the removed image recognition information characterizes the credibility of the image recognition result; Based on the generated removal image recognition information, the single-tube image is patched and removed to generate a single-tube image after removal.

3. The method according to claim 2, characterized in that, The step of performing anti-perturbation cleaning on the removed single-tube image to generate a cleaned single-tube image includes: Based on the image of the removed single tube, a noise distribution map and a multi-dimensional feature matrix of the image are generated. Based on the image's multidimensional feature matrix, the following perturbation cleanup steps are performed: A denoised single-tube image is generated based on a preset number of texture similarities, the multidimensional feature matrix of the image, the single-tube image after removal, and the noise distribution map. Based on the denoised single-tube image and the lightweight image recognizer, denoised image recognition information is generated; The image credibility is determined based on the image recognition information after removal and the image recognition information after denoising corresponding to the single tube image after removal. In response to determining that the image confidence satisfies the iterative optimization condition, the preset texture similarity number is updated according to the target residual map, and the perturbation cleaning step is performed again, wherein the generated target residual map is determined by the denoised single tube image and the removed single tube image; In response to the determination that the image credibility does not meet the iterative optimization condition, the denoised single-tube image is determined as the cleaned single-tube image.

4. The method according to claim 3, characterized in that, The step of generating a noise distribution map and a multi-dimensional feature matrix of the image based on the removed single tube image includes: The removed single-tube image is normalized to generate a removed single-tube image tensor; The noise of the removed single-tube image tensor is determined pixel by pixel to generate a noise distribution map, wherein the noise distribution map has the same size as the removed single-tube image tensor. Based on the preset block size and preset step size, the removed single tube image and the noise distribution map are synchronously segmented into blocks to obtain a set of removed single tube image blocks and a set of noise distribution map blocks. Based on the set of single-tube image blocks after removal and the set of noise distribution blocks, multi-dimensional feature information of each block is generated, wherein the generated multi-dimensional feature information of the blocks corresponds to the single-tube image blocks after removal. Based on the multidimensional feature information of each generated patch, the multidimensional feature matrix of the image is determined.

5. The method according to claim 4, characterized in that, The step of generating multi-dimensional feature information for each image patch based on the removed single-tube image patch set and the noise distribution patch set includes: For each removed single-tube image block in the set of removed single-tube image blocks, perform the following steps: Frequency domain features are extracted from the removed single-tube image block to generate image frequency domain feature information; Spatial domain features are extracted from the removed single-tube image block to generate image spatial domain feature information; The removed single-tube image block is lightweight encoded to generate image encoding feature information; Based on the noise distribution map block corresponding to the removed single tube image block, image noise feature information is generated; The generated image frequency domain feature information, image spatial domain feature information, image coding feature information, and image noise feature information are spliced ​​together to obtain multidimensional feature information of the image patch.

6. The method according to claim 1, characterized in that, The step of using a pre-trained multi-channel sample tube recognition model to perform sample tube recognition on the cleaned single-tube image to generate sample tube recognition information includes: The cleaned single-tube image is preprocessed to generate a text channel image tensor, an appearance channel image tensor, and a barcode channel image tensor. The multi-channel sample tube recognition model includes a text detection and recognition model, an appearance attribute detection model, and a stripe decoding model. Based on the text channel image tensor and the text detection and recognition model, generate sample tube text information; Based on the appearance channel image tensor and the appearance attribute detection model, generate sample tube appearance information; Based on the barcode channel image tensor and the stripe decoding model, generate sample tube barcode decoding information; The sample tube text information, the sample tube appearance information, and the sample tube barcode decoding information are determined as the sample tube identification information.

7. A sample image processing device, characterized in that, include: The grid segmentation unit is configured to perform grid segmentation on the acquired sample tube image to obtain a set of single tube images. The sample tube image is obtained by acquiring and transmitting images of a batch of sample tubes placed on a grid template. Each single tube image in the set of single tube images represents a local area of ​​a sample tube and corresponds to grid coordinate information. The processing unit is configured to perform the following processing steps for each single tube image in the single tube image set: performing adversarial patch removal processing on the single tube image to generate a removed single tube image; performing adversarial perturbation cleaning on the removed single tube image to generate a cleaned single tube image; and performing sample tube recognition on the cleaned single tube image using a pre-trained multi-channel sample tube recognition model to generate sample tube recognition information. The verification unit is configured to verify the credibility of the identification information of each sample tube and obtain the verification results of each piece of information. The storage unit is configured to store the generated sample tube identification information and the sample tube image in response to determining that the information verification result indicates that the verification has passed.

8. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Test tube and test tube stand identification system

    CN108745444A

  • Body fluid specimen quality detection method and device, transportation device, equipment and medium

    CN112037202A

  • Image classification method and device for multi-scale adversarial patches, storage medium and equipment

    CN117523302A

  • Sample tube batch code reading method, code reading device and storage library

    CN118246463A

  • Adversarial patch removing method based on adversarial patch positioning model

    CN118506074A