A method and system for identifying a reagent bottle
By enhancing and tilt-correcting the reagent bottle surface image, and combining the mapping matrix and text recognition model, the recognition error problem caused by the distortion of the circular reagent bottle label is solved, and high-precision recognition of the reagent bottle is achieved.
Patent Information
- Application Number
- CN202510316464.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-03-18
AI Technical Summary
After attaching a label to a round reagent bottle, the arc extension of the label causes image distortion, resulting in deviation in label information acquisition, which in turn affects the recognition accuracy of the reagent bottle.
By enhancing the surface image of the reagent bottle, adaptively obtaining the segmentation threshold for binary segmentation, performing edge detection to determine the label area, performing tilt correction and constructing a mapping matrix, mapping the label area to the standard area, and using the text recognition model to obtain the label information.
It effectively avoids information errors caused by label distortion, improves the precision and accuracy of reagent bottle identification, and improves the efficiency of text information extraction.
Smart Images

Figure CN119850931B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a reagent bottle recognition method and system. BACKGROUND
[0002] A reagent bottle is a glass or plastic bottle used to store reagents. It is mainly used to preserve reagents to ensure their stability and purity. The design of the reagent bottle usually takes into account the chemical inertness and sealing of the material to prevent the evaporation of reagents, moisture absorption, or reaction with components in the air.
[0003] After the reagent bottle is loaded with reagents, a label corresponding to the reagent bottle is usually attached to the reagent bottle. Then, by reading the information on the label of the reagent bottle, the recognition of the reagent bottle can be completed. The traditional recognition method usually identifies the information on the label manually, which is low in efficiency.
[0004] To improve the recognition efficiency of the reagent bottle, image recognition technology is usually used to directly identify the information on the label. However, the shape of the reagent bottle is usually circular, and after the label is attached, the label is extended in an arc surface, which will cause a certain degree of image distortion in the recognized label image, resulting in deviation in the acquisition of the information on the label, and further causing errors in the recognition of the reagent bottle. SUMMARY
[0005] To overcome the shortcomings of the prior art, the present application provides a reagent bottle recognition method and system, which aims to solve the technical problem that in the prior art, after a label is attached to a circular reagent bottle, the arc surface extension of the label causes a certain degree of distortion in the acquired label image, resulting in deviation in the acquisition of the information on the label, and further causing errors in the recognition of the reagent bottle.
[0006] To achieve the above-mentioned purpose, in a first aspect, the present application provides a reagent bottle recognition method, comprising the following steps:
[0007] When the label face of the reagent bottle is directly opposite the CCD camera, the surface image of the reagent bottle is acquired, and the surface image is subjected to sharpening processing to obtain an enhanced image;
[0008] The segmentation threshold of the enhanced image is adaptively acquired to perform binary segmentation on the enhanced image and acquire a gray-scale image. Edge detection is performed in the gray-scale image to determine the label region in the gray-scale image. Whether the label is completely photographed is judged based on the label region;
[0009] If the label is completely photographed, the top-left corner point coordinates and the bottom-left corner point coordinates of the label region are acquired, and the label region is subjected to tilt correction based on the top-left corner point coordinates and the bottom-left corner point coordinates;
[0010] obtaining a center point between a left top point and a right top point of the label region, obtaining a standard center point between a standard left top point and a standard right top point of a standard region, mapping the center point to the standard center point to obtain a first mapping matrix, mapping the left top point to the standard left top point to obtain a second mapping matrix, and mapping the right top point to the standard right top point to obtain a third mapping matrix;
[0011] mapping the label region to the standard region based on the first mapping matrix, the second mapping matrix and the third mapping matrix to obtain a standard label image, performing text detection in the standard label image to obtain label text information, and completing identification of a reagent bottle based on the label text information.
[0012] Further, the step of performing sharpness enhancement on the surface image to obtain an enhanced image comprises:
[0013] converting the surface image into a total amount of incident light and a total amount of reflected light, performing logarithmic processing and Fourier transform on the total amount of incident light and the total amount of reflected light to convert the total amount of incident light and the total amount of reflected light into a frequency domain respectively;
[0014] constructing a high-pass filter, and performing filtering processing on the total amount of incident light and the total amount of reflected light in the frequency domain based on the high-pass filter to obtain a frequency domain image;
[0015] performing exponential processing and inverse Fourier transform on the frequency domain image to obtain an enhanced image.
[0016] Further, the step of adaptively obtaining a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image and obtaining a gray scale image comprises:
[0017] obtaining a gray scale range of the enhanced image, the gray scale range comprising a plurality of gray scales, and obtaining a first probability of occurrence of a pixel corresponding to each gray scale, and separating the gray scale range into a first gray scale range and a second gray scale range based on an initial threshold value;
[0018] respectively obtaining a second probability and a third probability of occurrence of a pixel in the first gray scale range and the second gray scale range, obtaining a first gray scale average value of the first gray scale range based on the first probability and the second probability, and obtaining a second gray scale average value of the second gray scale range based on the first probability and the third probability;
[0019] obtaining a segmentation threshold based on the second probability, the first gray scale average value, the third probability and the second gray scale average value.
[0020] Further, the calculation formula of the second probability is:
[0021] ,
[0022] wherein, denotes the second probability, denotes the probability of a pixel with gray level i appearing in the first gray level range, denotes the initial threshold value, denotes the minimum gray level in the gray level range;
[0023] The calculation formula of the third probability is:
[0024] ,
[0025] wherein, denotes the third probability, denotes the probability of a pixel with gray level j appearing in the second gray level range, denotes the maximum gray level in the gray level range;
[0026] The calculation formula of the first gray level average value is:
[0027] ,
[0028] wherein, denotes the first gray level average value, denotes the first probability of a pixel with gray level i appearing in the gray level range;
[0029] The acquisition formula of the segmentation threshold value is:
[0030] ,
[0031] wherein, denotes the segmentation threshold value, denotes the second gray level average value.
[0032] Further, the step of determining whether the label is completely photographed based on the label region comprises:
[0033] acquiring the total number of pixels in the label region, and determining whether the total number of pixels is within a pixel number threshold range;
[0034] If the total number of pixels is within the pixel number threshold range, it is determined that the label is completely photographed.
[0035] Further, the top-left corner point coordinate comprises a top-left horizontal coordinate and a top-left vertical coordinate, the bottom-left corner point coordinate comprises a bottom-left horizontal coordinate and a bottom-left vertical coordinate, and the step of performing tilt correction on the label region based on the top-left corner point coordinate and the bottom-left corner point coordinate comprises:
[0036] obtaining a horizontal interval distance through the left upper horizontal coordinate and the left lower horizontal coordinate, and obtaining a vertical interval distance through the left upper vertical coordinate and the left lower vertical coordinate;
[0037] obtaining an offset angle based on the horizontal interval distance and the vertical interval distance;
[0038] comparing the left upper horizontal coordinate with the left lower horizontal coordinate to complete tilt correction through the offset angle.
[0039] Further, the step of mapping the label region to the standard region based on the first mapping matrix, the second mapping matrix and the third mapping matrix comprises:
[0040] obtaining a plurality of fourth mapping matrices from the pixel point between the left upper corner point and the center point to the standard left upper corner point and the standard center point based on the first mapping matrix and the second mapping matrix;
[0041] obtaining a plurality of fifth mapping matrices from the pixel point between the center point and the right upper corner point to the standard center point and the standard right upper corner point based on the first mapping matrix and the third mapping matrix;
[0042] mapping the label region to the standard region through the first mapping matrix, the second mapping matrix, the third mapping matrix, the plurality of fourth mapping matrices and the plurality of fifth mapping matrices.
[0043] Further, the step of performing character detection in the standard label image comprises:
[0044] setting a training image, the training image comprising training characters, obtaining target character features corresponding to the training characters, constructing an initial character recognition model, taking the training image as an input value of the initial character recognition model to obtain training character features in the initial character recognition model, comparing the training character features with a preset word library to select a plurality of reference character clusters from the preset word library, and combining the plurality of reference character clusters as an original sample cluster;
[0045] obtaining distances between the training character features and the plurality of reference character clusters to select a plurality of comparison clusters from the plurality of reference character clusters, and combining the plurality of comparison clusters as a negative sample cluster;
[0046] A first loss function is constructed based on the training text features and the reference text clusters, a second loss function is constructed based on the training text features, the target text features and the negative sample clusters, and a third loss function is constructed based on the training text features, the target text features and the original sample clusters;
[0047] A total loss function is constructed through the first loss function, the second loss function and the third loss function, and the initial text recognition model is trained based on the total loss function to obtain a final text recognition model.
[0048] Further, the expression of the first loss function is:
[0049] ,
[0050] wherein, represents the first loss function, represents the number of reference text clusters, represents the training text features, represents the center vector of the reference text cluster closest to the training text features, represents the center vector of the yth reference text cluster, represents the set of center vectors of the reference text clusters, represents a hyperparameter, represents a dot product operation, represents an exponential function, represents a logarithmic function;
[0051] The expression of the second loss function is:
[0052] ,
[0053] wherein, represents the second loss function, represents the number of negative samples in the negative sample cluster, represents the target text features, represents the pth negative sample in the negative sample cluster, represents the negative sample cluster;
[0054] The expression of the third loss function is:
[0055] ,
[0056] wherein, represents the third loss function, represents the number of samples in the original sample cluster, represents the qth sample in the original sample cluster, represents the original sample cluster.
[0057] In a second aspect, the embodiments of the present application provide a reagent bottle identification system, applied to the reagent bottle identification method in the first aspect, and the system comprises:
[0058] A first processing module is configured to acquire a surface image of a reagent bottle when a label surface of the reagent bottle faces a CCD camera, and perform sharpness enhancement processing on the surface image to acquire an enhanced image.
[0059] A second processing module is configured to adaptively acquire a segmentation threshold of the enhanced image, perform binaryzation segmentation on the enhanced image to acquire a gray-scale image, perform edge detection in the gray-scale image, determine a label region in the gray-scale image, and judge whether the label is completely photographed based on the label region.
[0060] An adjusting module is configured to, if the label is completely photographed, acquire a top-left corner point coordinate and a bottom-left corner point coordinate of the label region, and perform tilt correction on the label region based on the top-left corner point coordinate and the bottom-left corner point coordinate.
[0061] A conversion module is configured to acquire a center point between a top-left corner point and a top-right corner point of the label region, acquire a standard center point between a standard top-left corner point and a standard top-right corner point of a standard region, map the center point to the standard center point to acquire a first mapping matrix, map the top-left corner point to the standard top-left corner point to acquire a second mapping matrix, and map the top-right corner point to the standard top-right corner point to acquire a third mapping matrix.
[0062] An execution module is configured to map the label region to the standard region based on the first mapping matrix, the second mapping matrix, and the third mapping matrix to acquire a standard label image, perform character detection in the standard label image to acquire label character information, and complete reagent bottle identification based on the label character information.
[0063] In a third aspect, the embodiments of the present application provide a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the reagent bottle identification method in the first aspect when executing the computer program.
[0064] In a fourth aspect, the embodiments of the present application provide a storage medium having a computer program stored thereon, and the computer program is executable on a processor to implement the reagent bottle identification method in the first aspect.
[0065] Compared with the prior art, the beneficial effects of the present application are that: by performing the brightening processing on the surface pattern, the noise caused by uneven illumination can be avoided, and the accuracy of information recognition of the label in the later stage is improved; when the label is completely photographed, by performing the tilt correction on the label area, the label area is parallel to the subsequent standard area in the horizontal direction, and the efficiency of subsequent text extraction is improved; further, on the premise that the label area is parallel to the standard area, the first mapping matrix between the center point of the label area and the standard center point, the second mapping matrix between the top left corner point and the standard top left corner point, and the third mapping matrix between the top right corner point and the standard top right corner point are obtained, the mapping relationship of all pixel points of the standard area to the standard area can be derived through the three mapping relationships, and then the label which is in the form of an arc surface is flattened, the information acquisition error caused by label distortion is avoided, and the accuracy of reagent bottle recognition is improved; by constructing the final text recognition model, the distance between the input value and the true value is narrowed by the first loss function, and the difference between the input value and the similar value is improved by the second loss function, and then the text recognition error caused by similar characters is avoided, the accuracy of text information extraction is improved, and the accuracy of reagent bottle recognition is further improved. BRIEF DESCRIPTION OF DRAWINGS
[0066] Figure 1 Flow chart of the reagent bottle recognition method in the first embodiment of the present application;
[0067] Figure 2 Structure block diagram of the reagent bottle recognition system in the second embodiment of the present application;
[0068] The following specific embodiments will further illustrate the present application in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0069] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings. The drawings show several embodiments of the present application. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0070] It should be noted that when an element is referred to as being "fixedly attached" to another element, it can be directly on the other element or there can be an intervening element. When an element is referred to as being "connected" to another element, it can be directly connected to the other element or intervening elements can be present. The terms "vertical", "horizontal", "left", "right", and the like as used herein are for purposes of illustration and description only.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. The use herein of the terms "and / or" includes a set of one or more associated listed items.
[0072] Referring to Figure 1 The reagent bottle identification method provided by the first embodiment of the application comprises the following steps:
[0073] S10: When the label surface of the reagent bottle faces the CCD camera, the surface image of the reagent bottle is acquired, and the surface image is subjected to a sharpening process to obtain an enhanced image;
[0074] The step S10 comprises:
[0075] S110: The surface image is converted into the total amount of incident light and the total amount of reflected light, and the total amount of incident light and the total amount of reflected light are subjected to logarithmic processing and Fourier transform to convert the total amount of incident light and the total amount of reflected light into the frequency domain, respectively;
[0076] The relationship between the surface image and the total amount of incident light and the total amount of reflected light is:
[0077] ,
[0078] wherein, represents the surface image, represents the total amount of incident light, represents the total amount of reflected light.
[0079] After the logarithmic processing is performed, the product relationship between the total amount of incident light and the total amount of reflected light is changed into an additive relationship, the separation between the total amount of incident light and the total amount of reflected light is realized, and then the subsequent filtering processing on the total amount of incident light and the total amount of reflected light in the frequency domain is facilitated.
[0080] S120: A high-pass filter is constructed, and the total amount of incident light and the total amount of reflected light in the frequency domain are subjected to filtering processing based on the high-pass filter to obtain a frequency domain image;
[0081] The expression of the high-pass filter is:
[0082] ,
[0083] wherein, represents the transfer function of the high-pass filter, represents the cutoff frequency, represents the frequency point a distance to a frequency plane origin, denotes an order of a high-pass filter.
[0084] S130: performing exponential processing and inverse Fourier transform on the frequency domain image to obtain an enhanced image;
[0085] It can be understood that after the processing is completed, the image is converted from the frequency domain back to the enhanced image.
[0086] S20: adaptively obtaining a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image and obtain a gray scale image, performing edge detection in the gray scale image to determine a label region in the gray scale image, and judging whether the label is completely photographed based on the label region;
[0087] The step S20 comprises:
[0088] S210: obtaining a gray scale range of the enhanced image, the gray scale range comprising a plurality of gray scales, and obtaining a first probability of occurrence of pixels corresponding to each gray scale, and separating the gray scale range into a first gray scale range and a second gray scale range based on an initial threshold value;
[0089] By obtaining the number of pixels corresponding to each gray scale, the total number of pixels corresponding to the gray scale range can be obtained, and by dividing the number of pixels corresponding to each gray scale by the total number of pixels corresponding to the gray scale range, the first probability of occurrence of pixels corresponding to each gray scale can be obtained. It can be understood that the first gray scale range and the second gray scale range are mixed to form the gray scale range.
[0090] S220: obtaining a second probability and a third probability of occurrence of pixels in the first gray scale range and the second gray scale range, respectively, obtaining a first gray scale average value of the first gray scale range based on the first probability and the second probability, and obtaining a second gray scale average value of the second gray scale range based on the first probability and the third probability;
[0091] The calculation formula of the second probability is:
[0092] ,
[0093] wherein, denotes a second probability, denotes a probability of occurrence of pixels with a gray scale of i in the first gray scale range, denotes an initial threshold value, denotes a minimum gray scale in the gray scale range;
[0094] The calculation formula of the third probability is:
[0095] ,
[0096] wherein, denotes the third probability, denotes the probability of the pixel with the gray level j appearing in the second gray scale range, denotes the maximum gray level in the gray level range;
[0097] The calculation formula of the first gray average value is:
[0098] ,
[0099] wherein, denotes the first gray average value, denotes the first probability of the pixel with the gray level i appearing in the gray level range. The second gray average value is acquired in the same way as the first gray average value, which will not be described here.
[0100] S230: acquiring a segmentation threshold based on the second probability, the first gray average value, the third probability and the second gray average value;
[0101] The acquisition formula of the segmentation threshold is:
[0102] ,
[0103] wherein, denotes the segmentation threshold, denotes the second gray average value.
[0104] By adaptively acquiring the segmentation threshold, different surface images can be adapted to complete accurate binarization processing of the image. The edge detection algorithm is now widely used, which will not be described here. After completing the edge detection, the contour of the label can be acquired, and then the label region can be acquired.
[0105] S240: acquiring the total number of pixels in the label region, and determining whether the total number of pixels is located in the pixel number threshold range;
[0106] S250: if the total number of pixels is located in the pixel number threshold range, it is determined that the label is completely photographed;
[0107] It should be noted that if the label is not completely photographed, the placement position of the reagent bottle is adjusted again until the label is completely photographed.
[0108] S30: if the label is completely photographed, the top-left corner coordinate and the bottom-left corner coordinate of the label region are acquired, and the label region is corrected based on the top-left corner coordinate and the bottom-left corner coordinate.
[0109] The step S30 comprises:
[0110] S310: obtaining a horizontal interval distance by the left upper horizontal coordinate and the left lower horizontal coordinate, and obtaining a vertical interval distance by the left upper vertical coordinate and the left lower vertical coordinate;
[0111] S320: obtaining an offset angle based on the horizontal interval distance and the vertical interval distance;
[0112] S330: comparing the left upper horizontal coordinate with the left lower horizontal coordinate to complete the tilt correction by the offset angle;
[0113] When the offset angle is obtained, if the left upper horizontal coordinate is greater than the left lower horizontal coordinate, taking the left upper horizontal coordinate as a reference point, the left lower horizontal coordinate is rotated to the direction of the left upper horizontal coordinate by the offset angle to complete the tilt correction; if the left upper horizontal coordinate is less than the left lower horizontal coordinate, taking the left lower horizontal coordinate as a reference point, the left upper horizontal coordinate is rotated to the direction of the left lower horizontal coordinate by the offset angle to complete the tilt correction.
[0114] S40: obtaining a center point between a left upper corner point and a right upper corner point of the label region, and obtaining a standard center point between a standard left upper corner point and a standard right upper corner point of a standard region, mapping the center point to the standard center point to obtain a first mapping matrix, mapping the left upper corner point to the standard left upper corner point to obtain a second mapping matrix, and mapping the right upper corner point to the standard right upper corner point to obtain a third mapping matrix;
[0115] The relationship of the center point, the standard center point and the first mapping matrix is:
[0116] ,
[0117] wherein, represents the coordinate of the standard center point, wherein, represents an auxiliary coordinate for normalization processing, represents the coordinate of the center point, represents the first mapping matrix, wherein, , , , respectively represent linear transformation parameters, , all represent translation parameters, , all represent perspective transformation parameters, in the embodiment, , The values of the second mapping matrix and the third mapping matrix are 0. The acquisition manners of the second mapping matrix and the third mapping matrix are consistent with the acquisition manner of the first mapping matrix, which will not be described here.
[0118] If the tilt and the arc surface distortion are directly solved by the mapping processing, a plurality of geometric deformations such as rotation, perspective, and arc surface flattening need to be compensated in a single change, which will significantly increase the parameter optimization difficulty of the mapping matrix. Especially when the tilt angle is large, the interpolation error may be accumulated. By performing the tilt correction, the problem of non-horizontal alignment of the label region caused by the deviation of the shooting angle or the manual pasting deviation is eliminated, and then the arc surface distortion problem is separately processed when the label region is in a horizontal state, which significantly reduces the complexity of the algorithm and improves the mapping accuracy after the tilt correction. If the tilt correction is not performed and the tilt and the arc surface distortion are directly processed by the mapping, the point used for mapping will be offset from the ideal position, which will cause a certain error in the calculation of the mapping matrix. The rotation parameter and the arc surface flattening parameter may also interfere with each other, which will cause the optimization process to fall into a local optimum, and multiple iterations are required to optimize and find the optimal parameter. Compared with the step-by-step processing mode in the present application, the number of key points obtained by direct mapping is also much larger than the number of key points required by the step-by-step processing. If a local optimum problem occurs in the optimization process, the standard label image obtained at the end may also have distortion. For a 500*200 pixel label region, if the direct mapping is used to obtain the standard label image without considering the distortion, four key points are assumed to be set, and the iteration optimization is performed, the time required to obtain the standard label image is about 20 ms. However, by using the initial mode of the present application, the tilt correction is performed first, and then the mapping is performed, and the standard label image can be obtained in only 10 ms, which significantly improves the efficiency of obtaining the standard label image.
[0119] S50: mapping the label region to the standard region based on the first mapping matrix, the second mapping matrix, and the third mapping matrix to obtain a standard label image, performing text detection in the standard label image to obtain label text information, and completing identification of a reagent bottle based on the label text information;
[0120] The step S50 includes:
[0121] S510: obtaining a plurality of fourth mapping matrices from the pixel points between the top-left corner point and the center point to the standard top-left corner point and the standard center point based on the first mapping matrix and the second mapping matrix;
[0122] S520: Obtain a plurality of fifth mapping matrices between the pixel points between the center point and the top-right corner point and the standard center point and the standard top-right corner point based on the first mapping matrix and the third mapping matrix;
[0123] S530: Map the label region to the standard region through the first mapping matrix, the second mapping matrix, the third mapping matrix, the plurality of fourth mapping matrices, and the plurality of fifth mapping matrices;
[0124] It can be understood that, since the label region has been corrected, the pixel points on the same vertical line with the top-left corner point, the center point, and the top-right corner point can be mapped to the standard region by adjusting the parameters of the first mapping matrix, the second mapping matrix, and the third mapping matrix. After obtaining the first mapping matrix and the second mapping matrix, the fourth mapping matrix for each pixel point between the top-left corner point and the center point can be obtained by combining the first mapping matrix and the second mapping matrix. Then, all the pixel points on the same vertical line can be mapped based on the distance between the vertical lines by adjusting the fourth mapping matrix. The fifth mapping matrix is the same, and will not be described here. The fourth mapping matrix and the fifth mapping matrix are obtained through the first mapping matrix, the second mapping matrix, and the third mapping matrix. The center point is used as a dividing point to separate the label region into two parts to obtain more detailed mapping relationship to avoid the label region being rotated too much to the left or too much to the right, resulting in different lengths of the left and right parts, and further causing mapping deviation in the mapping process.
[0125] S540: Set a training image, the training image includes training text, obtain target text features corresponding to the training text, construct an initial text recognition model, take the training image as an input value of the initial text recognition model to obtain training text features in the initial text recognition model, compare the training text features with a preset word library to select a plurality of reference text clusters from the preset word library, and combine the plurality of reference text clusters into an original sample cluster;
[0126] The benchmark text cluster includes a plurality of benchmark text features with high feature similarity. It should be noted that in the process of obtaining the original sample cluster, the center vector of the benchmark text cluster is obtained, the distance between the training text feature and the center vector of the benchmark text cluster is calculated to obtain the distance between the training text feature and the benchmark text cluster, and the benchmark text cluster closest to the training text feature is removed. The remaining benchmark text clusters are aggregated into the original sample cluster, which includes a plurality of original text features, i.e. samples.
[0127] S550: Obtain the distance between the training text feature and a plurality of benchmark text clusters, and select a plurality of comparison clusters from the plurality of benchmark text clusters; and
[0128] Specifically, after the distance is obtained, a plurality of distances are compared with a distance threshold, respectively, the benchmark text cluster corresponding to the distance less than the distance threshold is selected as a standby text cluster, and the standby text cluster closest to the training text feature is removed from the plurality of standby text clusters to obtain a plurality of comparison clusters, which is essentially the same as the method of obtaining the original sample cluster. Because the target text feature exists in the standby text cluster closest to the training text feature, by removing it, the true value can be prevented from being mixed in the negative sample cluster, which can cause errors in the training of the model and affect the accuracy of the label text information acquisition. Understandably, the comparison cluster and the negative sample cluster each include a plurality of benchmark text features, which are negative samples.
[0129] S560: Construct a first loss function based on the training text feature and the benchmark text cluster, a second loss function based on the training text feature, the target text feature and the negative sample cluster, and a third loss function based on the training text feature, the target text feature and the original sample cluster.
[0130] The expression of the first loss function is:
[0131]
[0132] wherein, represents the first loss function, represents the number of benchmark text clusters, represents the training text feature, represents the center vector of the benchmark text cluster closest to the training text feature, represents the center vector of the yth benchmark text cluster, a set of center vectors representing the reference word clusters, representing a hyperparameter, representing a dot product operation, representing an exponential function, representing a logarithmic function; for example, assuming that there are 2 reference word clusters, the cluster containing the word "is" is , and the cluster containing the word "has" is By setting the first loss function, the training word features are distinguished from the approximate clusters according to the similarity of the training word features and all the reference word clusters, and the problem of fuzzy semantic boundaries of different categories is solved.
[0133] The expression of the second loss function is:
[0134] ,
[0135] wherein, representing a second loss function, representing the number of negative samples in a negative sample cluster, representing a target word feature, representing the pth negative sample in a negative sample cluster, representing a negative sample cluster. It should be noted that the target word feature is a collection of variants of the plurality of training word features in the representation space. For example, there may be variants such as rotation, inversion, etc. By setting the second loss function, the purpose is to narrow the distance between the training word features and the target word features, and to solve the problem of fine-grained distinction of different variants.
[0136] The expression of the third loss function is:
[0137] ,
[0138] wherein, representing a third loss function, representing the number of samples in an original sample cluster, representing the qth sample in an original sample cluster, representing an original sample cluster. By setting the third loss function, the purpose is to globally distinguish the training word features from all other characters.
[0139] S570: constructing a total loss function by the first loss function, the second loss function and the third loss function, and training the initial word recognition model based on the total loss function to obtain a final word recognition model;
[0140] The expression of the total loss function is:
[0141] ,
[0142] wherein, represents a total loss function, , , all represent contribution coefficients. By setting the first loss function, the second loss function and the third loss function to obtain the final loss function, the accuracy of character recognition can be effectively improved, thereby ensuring the accuracy of recognition.
[0143] S580: taking the standard label image as an input value of the final character recognition model to complete character detection;
[0144] Through the final character recognition model, real character features can be obtained, and the real character features are converted into the label character information.
[0145] By performing the sharpening processing on the surface graph, noise caused by uneven illumination can be avoided, and the accuracy of information recognition of the label in the later stage is improved. When the label is completely photographed, by performing the tilt correction on the label region, the label region and the subsequent standard region are parallel in the horizontal direction, and the efficiency of subsequent character extraction is improved. Further, on the premise that the label region and the standard region are parallel, a first mapping matrix between the center point of the label region and the standard center point, a second mapping matrix between the top-left corner point and the standard top-left corner point, and a third mapping matrix between the top-right corner point and the standard top-right corner point are obtained. The mapping relationship of all pixel points of the standard region to the standard region can be derived through the three mapping relationships, and then the label which is extended in an arc surface is flattened, the information acquisition error caused by label distortion is avoided, and the accuracy of reagent bottle recognition is improved. By constructing the final character recognition model, the distance between the input value and the real value is narrowed by the first loss function, and the difference between the input value and the similar value is improved by the second loss function, thereby avoiding the occurrence of character recognition errors caused by similar characters, improving the accuracy of character information extraction, and further improving the accuracy of reagent bottle recognition.
[0146] Referring to Figure 2 , the second embodiment of the present application provides a reagent bottle recognition system, which is applied to the reagent bottle recognition method in the above-mentioned embodiments, and details have been described above. As used below, the terms "module", "unit", "sub-unit", and the like can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware, or a combination of software and hardware can also be implemented and conceived.
[0147] The system comprises:
[0148] The first processing module 10 is used to obtain a surface image of the reagent bottle when the label surface of the reagent bottle faces the CCD camera, and perform a clarity enhancement process on the surface image to obtain an enhanced image;
[0149] The first processing module 10 includes:
[0150] a first unit configured to convert the surface image into a total amount of incident light and a total amount of reflected light, and perform logarithmic processing and Fourier transform on the total amount of incident light and the total amount of reflected light, so as to convert the total amount of incident light and the total amount of reflected light into a frequency domain respectively;
[0151] The second unit is configured to construct a high-pass filter, and perform filtering processing on the total amount of incident light and the total amount of reflected light in the frequency domain based on the high-pass filter to obtain a frequency domain image;
[0152] A third unit is configured to perform exponential processing and inverse Fourier transform on the frequency domain image to obtain an enhanced image;
[0153] a second processing module 20 for adaptively acquiring a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image, acquiring a grayscale image, performing edge detection within the grayscale image to determine a label region in the grayscale image, and determining whether the label is completely captured based on the label region;
[0154] The second processing module 20 includes:
[0155] a fourth unit configured to obtain a grayscale range of the enhanced image, the grayscale range including a plurality of grayscale levels, obtain a first probability of occurrence of a pixel corresponding to each grayscale level, and separate the grayscale range into a first grayscale range and a second grayscale range based on an initial threshold;
[0156] a fifth unit, configured to obtain a second probability and a third probability of occurrence of pixels in the first grayscale range and the second grayscale range, respectively, obtain a first grayscale average value of the first grayscale range based on the first probability and the second probability, and obtain a second grayscale average value of the second grayscale range based on the first probability and the third probability;
[0157] A sixth unit is configured to obtain a segmentation threshold based on the second probability, the first grayscale average value, the third probability, and the second grayscale average value;
[0158] The adjusting module 30 is configured to, if the label is completely captured, acquire a top-left corner point coordinate and a left-bottom corner point coordinate of the label region, and perform tilt correction on the label region based on the top-left corner point coordinate and the left-bottom corner point coordinate.
[0159] The adjusting module 30 comprises:
[0160] A seventh unit is configured to acquire a horizontal interval distance through the top-left horizontal coordinate and the left-bottom horizontal coordinate, and acquire a vertical interval distance through the top-left vertical coordinate and the left-bottom vertical coordinate.
[0161] An eighth unit is configured to acquire an offset angle based on the horizontal interval distance and the vertical interval distance.
[0162] A ninth unit is configured to compare the top-left horizontal coordinate with the left-bottom horizontal coordinate to complete tilt correction through the offset angle.
[0163] The conversion module 40 is configured to acquire a center point between a top-left corner point and a top-right corner point of the label region, acquire a standard center point between a standard top-left corner point and a standard top-right corner point of a standard region, map the center point to the standard center point to acquire a first mapping matrix, map the top-left corner point to the standard top-left corner point to acquire a second mapping matrix, and map the top-right corner point to the standard top-right corner point to acquire a third mapping matrix.
[0164] The execution module 50 is configured to map the label region to the standard region based on the first mapping matrix, the second mapping matrix and the third mapping matrix to acquire a standard label image, perform text detection in the standard label image to acquire label text information, and complete identification of a reagent bottle based on the label text information.
[0165] The execution module 50 comprises:
[0166] A tenth unit is configured to acquire a plurality of fourth mapping matrices between pixel points between the top-left corner point and the center point and between the standard top-left corner point and the standard center point based on the first mapping matrix and the second mapping matrix.
[0167] An eleventh unit is configured to acquire a plurality of fifth mapping matrices between pixel points between the center point and the top-right corner point and between the standard center point and the standard top-right corner point based on the first mapping matrix and the third mapping matrix.
[0168] A twelfth unit is configured to map the label region to the standard region through the first mapping matrix, the second mapping matrix, the third mapping matrix, the plurality of fourth mapping matrices and the plurality of fifth mapping matrices.
[0169] The thirteenth unit is configured to set a training image, obtain a target character feature corresponding to a training character included in the training image, construct an initial character recognition model, take the training image as an input value of the initial character recognition model, obtain a training character feature in the initial character recognition model, compare the training character feature with a preset word library, select a plurality of reference character clusters from the preset word library, and combine the plurality of reference character clusters into an original sample cluster;
[0170] The fourteenth unit is configured to obtain a distance between the training character feature and the plurality of reference character clusters, select a plurality of comparison clusters from the plurality of reference character clusters, and combine the plurality of comparison clusters into a negative sample cluster.
[0171] The fifteenth unit is configured to construct a first loss function based on the training character feature and the reference character cluster, construct a second loss function based on the training character feature, the target character feature, and the negative sample cluster, and construct a third loss function based on the training character feature, the target character feature, and the original sample cluster.
[0172] The sixteenth unit is configured to construct a total loss function based on the first loss function, the second loss function, and the third loss function, train the initial character recognition model based on the total loss function, and obtain a final character recognition model.
[0173] The seventeenth unit is configured to take the standard label image as an input value of the final character recognition model to complete character detection.
[0174] The present application also provides a computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the reagent bottle recognition method described in the above technical solution.
[0175] The present application also provides a storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the reagent bottle recognition method described in the above technical solution.
[0176] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0177] The above embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation on the patent scope of the present application. It should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for identifying a reagent bottle, characterized in that: The following steps are involved: When the label surface of the reagent bottle faces the CCD camera, a surface image of the reagent bottle is acquired, and the surface image is enhanced to acquire an enhanced image; Adaptively obtaining a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image, obtaining a grayscale image, performing edge detection within the grayscale image to determine a label area in the grayscale image, and determining whether the label is completely captured based on the label area; If the label is completely photographed, obtaining the coordinates of the upper left corner point and the lower left corner point of the label area, and performing tilt correction on the label area based on the coordinates of the upper left corner point and the coordinates of the lower left corner point; Obtaining a center point between the upper left corner point and the upper right corner point of the label area, and obtaining a standard center point between the standard upper left corner point and the standard upper right corner point of the standard area, mapping the center point to the standard center point to obtain a first mapping matrix, mapping the upper left corner point to the standard upper left corner point to obtain a second mapping matrix, and mapping the upper right corner point to the standard upper right corner point to obtain a third mapping matrix; Mapping the label area to the standard area based on the first mapping matrix, the second mapping matrix, and the third mapping matrix to obtain a standard label image, performing text detection in the standard label image to obtain label text information, and completing the identification of the reagent bottle based on the label text information; The step of performing text detection in the standard label image includes: Setting a training image, wherein the training image includes training characters, obtaining target character features corresponding to the training characters, constructing an initial character recognition model, using the training image as an input value of the initial character recognition model to obtain training character features in the initial character recognition model, comparing the training character features with a preset vocabulary, selecting a plurality of benchmark character clusters from the preset vocabulary, and combining the plurality of benchmark character clusters into an original sample cluster; Obtaining the distance between the training text feature and the plurality of reference text clusters, selecting a plurality of comparison clusters from the plurality of reference text clusters, and combining the plurality of comparison clusters into a negative sample cluster; Constructing a first loss function based on the training text features and the benchmark text clusters, constructing a second loss function based on the training text features, the target text features, and the negative sample clusters, and constructing a third loss function based on the training text features, the target text features, and the original sample clusters; constructing a total loss function by using the first loss function, the second loss function, and the third loss function, and training the initial text recognition model based on the total loss function to obtain a final text recognition model; The standard label image is used as the input value of the final text recognition model to complete text detection.
2. The method for identifying a reagent bottle according to claim 1, wherein: The step of performing a definition enhancement process on the surface image to obtain an enhanced image comprises: converting the surface image into a total amount of incident light and a total amount of reflected light, performing logarithmic processing and Fourier transform on the total amount of incident light and the total amount of reflected light to convert the total amount of incident light and the total amount of reflected light into a frequency domain respectively; constructing a high-pass filter, and performing filtering processing on the total amount of incident light and the total amount of reflected light in the frequency domain based on the high-pass filter to obtain a frequency domain image; The frequency domain image is subjected to exponential processing and inverse Fourier transformation to obtain an enhanced image.
3. The method for identifying a reagent bottle according to claim 1, wherein: The step of adaptively obtaining a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image and obtain a grayscale image includes: Obtaining a grayscale range of the enhanced image, the grayscale range including a plurality of grayscale levels, obtaining a first probability of occurrence of a pixel corresponding to each grayscale level, and dividing the grayscale range into a first grayscale range and a second grayscale range based on an initial threshold; Obtaining a second probability and a third probability of occurrence of pixels within the first grayscale range and the second grayscale range, respectively, obtaining a first grayscale average value of the first grayscale range based on the first probability and the second probability, and obtaining a second grayscale average value of the second grayscale range based on the first probability and the third probability; A segmentation threshold is obtained based on the second probability, the first grayscale average, the third probability, and the second grayscale average.
4. The method for identifying a reagent bottle according to claim 3, wherein: The calculation formula of the second probability is: , in, represents the second probability, represents the probability that a pixel with gray level i appears in the first gray range, represents the initial threshold, Indicates the minimum gray level in the gray level range; The calculation formula of the third probability is: , in, represents the third probability, represents the probability that a pixel with gray level j appears in the second gray range, Indicates the maximum gray level within the gray level range; The calculation formula of the first grayscale average value is: , in, represents the first grayscale average value, represents the first probability of the occurrence of a pixel with gray level i within the gray level range; The formula for obtaining the segmentation threshold is: , in, represents the segmentation threshold, Represents the second grayscale average value.
5. The method for identifying a reagent bottle according to claim 1, wherein: The step of determining whether the tag is completely photographed based on the tag area includes: Obtaining the total number of pixels in the label area, and determining whether the total number of pixels is within a pixel number threshold range; If the total number of pixels is within the pixel number threshold range, it is determined that the tag is completely photographed.
6. The method for identifying a reagent bottle according to claim 1, wherein: The coordinates of the upper left corner point include an upper left horizontal coordinate and an upper left vertical coordinate, and the coordinates of the lower left corner point include a lower left horizontal coordinate and a lower left vertical coordinate. The step of performing tilt correction on the label area based on the coordinates of the upper left corner point and the coordinates of the lower left corner point includes: Obtain a horizontal spacing distance through the upper left horizontal coordinate and the lower left horizontal coordinate, and obtain a vertical spacing distance through the upper left vertical coordinate and the lower left vertical coordinate; Obtaining an offset angle based on the lateral spacing distance and the longitudinal spacing distance; The upper left horizontal coordinate is compared with the lower left horizontal coordinate to complete the tilt correction through the offset angle.
7. The method for identifying a reagent bottle according to claim 1, wherein: The step of mapping the label area to the standard area based on the first mapping matrix, the second mapping matrix, and the third mapping matrix includes: Acquire, based on the first mapping matrix and the second mapping matrix, a plurality of fourth mapping matrices from pixel points between the upper left corner point and the center point to pixel points between the standard upper left corner point and the standard center point; Acquire a plurality of fifth mapping matrices from pixel points between the center point and the upper right corner point to pixel points between the standard center point and the standard upper right corner point based on the first mapping matrix and the third mapping matrix; The label area is mapped to the standard area through the first mapping matrix, the second mapping matrix, the third mapping matrix, a plurality of the fourth mapping matrices, and a plurality of the fifth mapping matrices.
8. The method for identifying a reagent bottle according to claim 1, wherein: The expression of the first loss function is: , in, represents the first loss function, represents the number of benchmark text clusters, represents the training text features, Represents the center vector of the benchmark text cluster closest to the training text feature, represents the center vector of the y-th benchmark text cluster, represents the set of center vectors of the benchmark text clusters, represents the hyperparameter, represents the dot product operation, represents the exponential function, represents the logarithmic function; The expression of the second loss function is: , in, represents the second loss function, represents the number of negative samples in the negative sample cluster, Indicates the target text features, represents the pth negative sample in the negative sample cluster, Represents negative sample clustering; The expression of the third loss function is: , in, represents the third loss function, represents the number of samples in the original sample cluster, represents the qth sample in the original sample cluster, Represents the original sample clustering.
9. A reagent bottle identification system, applied to the reagent bottle identification method according to any one of claims 1 to 8, characterized in that: The system comprises: A first processing module is configured to obtain a surface image of the reagent bottle when the label surface of the reagent bottle faces the CCD camera, and perform a clarity enhancement process on the surface image to obtain an enhanced image; a second processing module, configured to adaptively obtain a segmentation threshold of the enhanced image to perform binary segmentation on the enhanced image, obtain a grayscale image, perform edge detection within the grayscale image to determine a label area in the grayscale image, and determine whether the label is completely captured based on the label area; an adjustment module, configured to obtain the coordinates of the upper left corner and the lower left corner of the label area if the label is completely photographed, and perform tilt correction on the label area based on the coordinates of the upper left corner and the lower left corner; a conversion module, configured to obtain a center point between the upper left corner point and the upper right corner point of the label area, and obtain a standard center point between the standard upper left corner point and the standard upper right corner point of the standard area, map the center point to the standard center point to obtain a first mapping matrix, map the upper left corner point to the standard upper left corner point to obtain a second mapping matrix, and map the upper right corner point to the standard upper right corner point to obtain a third mapping matrix; an execution module, configured to map the label area to the standard area based on the first mapping matrix, the second mapping matrix, and the third mapping matrix to obtain a standard label image, perform text detection in the standard label image to obtain label text information, and complete the identification of the reagent bottle based on the label text information; The execution module includes: A thirteenth unit is configured to set a training image, wherein the training image includes training characters, obtain target character features corresponding to the training characters, construct an initial character recognition model, use the training image as an input value of the initial character recognition model to obtain training character features in the initial character recognition model, compare the training character features with a preset vocabulary, select a plurality of benchmark character clusters from the preset vocabulary, and combine the plurality of benchmark character clusters into an original sample cluster; A fourteenth unit is configured to obtain a distance between the training text feature and a plurality of the reference text clusters, select a plurality of comparison clusters from the plurality of the reference text clusters, and combine the plurality of comparison clusters into a negative sample cluster; A fifteenth unit is configured to construct a first loss function based on the training text features and the benchmark text clusters, construct a second loss function based on the training text features, the target text features, and the negative sample clusters, and construct a third loss function based on the training text features, the target text features, and the original sample clusters; A sixteenth unit is configured to construct a total loss function using the first loss function, the second loss function, and the third loss function, and train the initial text recognition model based on the total loss function to obtain a final text recognition model; The seventeenth unit is configured to use the standard label image as an input value of the final text recognition model to complete text detection.
Citation Information
Patent Citations
Label identification method and device
CN105844277A