Multi-target UDI consumable identification method and system, medium and product
By generating a non-reflective fused image using dual-polarization imaging technology and combining it with a cascaded recognition model, the problems of recognition stability and robustness in the identification of thin-film high-reflectivity boxed consumables were solved, and high-precision automatic identification of UDI consumables was achieved.
Patent Information
- Application Number
- CN202511756631.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-27
AI Technical Summary
In the identification of highly reflective boxed consumables in thin films, existing technologies cannot effectively suppress strong reflections with single visible light imaging, resulting in overexposure of the barcode area or loss of texture, insufficient identification stability, and poor robustness, especially in complex environments.
The system employs dual-polarization imaging technology to acquire orthogonal polarization images and generate a reflective fused image. By combining a cascaded packaging box and barcode recognition model with coordinate transformation and preset constraint rules for intelligent matching, the system improves recognition accuracy and reliability.
It achieves end-to-end fully automated identification of multi-target, highly reflective boxed UDI consumables, significantly improving the success rate, accuracy and robustness of identification, and meeting the traceability data requirements of medical device regulation.
Smart Images

Figure CN121581083A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method, system, medium, and product for identifying multi-target UDI consumables. Background Technology
[0002] In the field of automated supervision and quality control of medical devices, consumable identification technology based on Unique Device Identification (UDI) is becoming increasingly important. However, thin-film highly reflective boxed consumables (such as laminated paper boxes, thin-film sealed boxes, etc.) are prone to strong specular and Fresnel reflections under traditional imaging conditions due to their high-gloss film coating. This severely obscures the underlying barcode texture, leading to a significant increase in barcode recognition failure rate and seriously affecting the robustness and practicality of the identification system.
[0003] To mitigate reflective interference, existing technologies often employ recognition schemes based on visible light cameras and single-stage target detection networks. This approach uses a single RGB industrial camera to acquire images, utilizes a model to simultaneously detect the positions of the packaging box and the barcode, and then performs simple matching based on spatial overlap. This achieves a certain degree of automated recognition and boasts high processing speed, making it suitable for recognition tasks under general lighting conditions.
[0004] However, the single visible light imaging of this method cannot fundamentally suppress the strong reflection on the film surface, resulting in overexposure or loss of texture in the barcode area, and insufficient recognition stability. The single detection network also has difficulty in taking into account the feature learning of large-scale packaging boxes and small-scale barcodes. Especially in complex geometric structures such as box edges and corners, the problems of barcode missed detection and positioning deviation are prominent, which reduces the robustness of UDI consumable recognition in environments with high reflectivity, multiple targets, and complex spatial layout. Summary of the Invention
[0005] This application provides a multi-target UDI consumable identification method, system, medium, and product to address the technical problem of how to improve the robustness of UDI consumable identification in complex environments.
[0006] In a first aspect, embodiments of this application provide a multi-target UDI consumable identification method, including: Acquire a first polarization image of the target scene at a first polarization angle and a second polarization image at a second polarization angle, wherein the target scene includes multiple boxed consumables, and the first polarization angle and the second polarization angle are orthogonal; Based on the first polarization image and the second polarization image, a reflection-free fused image of the target scene is generated; Based on a preset packaging box recognition model, the target region of interest (ROI) corresponding to the target boxed consumable in the non-reflective fused image and the first position coordinates corresponding to the target ROI are identified. The target boxed consumable is any boxed consumable in the target scene. Based on a preset barcode recognition model, the barcode set in the target ROI is identified, and the second position coordinates of the target barcode in the target ROI are located, wherein the target barcode is any barcode in the barcode set; Based on the first position coordinates and the second position coordinates, the second position coordinates are transformed to obtain the third position coordinates of the target barcode in the non-reflective fused image; Based on the first location coordinates and the third location coordinates, the UDI consumable identification result of each barcode corresponding to each boxed consumable is determined by a preset constraint rule.
[0007] Optionally, generating a reflective fused image of the target scene based on the first polarization image and the second polarization image includes: obtaining a first pixel intensity set of the first polarization image and a second pixel intensity set of the second polarization image, wherein the first pixel intensity set includes multiple first pixel intensities corresponding to the first polarization image, and the second pixel intensity set includes multiple second pixel intensities corresponding to the second polarization image; for each pixel, subtracting the second pixel intensity of the target pixel from the first pixel intensity of the target pixel to obtain a target pixel difference, wherein the target pixel is any of the aforementioned pixels; calculating the sum of the first pixel intensity and the second pixel intensity of the target pixel to obtain a target pixel sum value; calculating the ratio of the target pixel difference to the target pixel sum value to obtain the linear polarization degree of the target pixel; and fusing the first polarization image and the second polarization image based on each linear polarization degree and a preset weighted fusion function to obtain the reflective fused image.
[0008] Optionally, the step of fusing the first polarization image and the second polarization image based on each of the linear polarization degrees and a preset weighted fusion function to obtain the anti-reflective fused image includes: using the difference between 1 and the linear polarization degree of the target pixel as the polarization compensation coefficient of the target pixel; calculating the product of a first preset illumination compensation coefficient, the polarization compensation coefficient, and the first pixel intensity of the target pixel to obtain the first compensated pixel intensity of the compensated target pixel; calculating the product of a second preset illumination compensation coefficient and the second pixel intensity of the target pixel to obtain the second compensated pixel intensity of the target pixel; obtaining the fused target pixel based on the first compensated pixel intensity and the second compensated pixel intensity; and obtaining the anti-reflective fused image based on each of the fused target pixels.
[0009] Optionally, the step of performing coordinate transformation on the second position coordinates based on the first position coordinates and the second position coordinates to obtain the third position coordinates of the target barcode in the anti-reflective fused image includes: calculating the sum of the minimum abscissa of the first position coordinates and the abscissa of the center point of the second position coordinates to obtain the abscissa of the target barcode in the anti-reflective fused image; calculating the sum of the minimum ordinate of the first position coordinates and the ordinate of the center point of the second position coordinates to obtain the ordinate of the target barcode in the anti-reflective fused image; and using the abscissa and ordinate of the barcode as the third position coordinates of the target barcode.
[0010] Optionally, the preset constraint rules include: a first constraint rule: the horizontal coordinate of the barcode is greater than or equal to the minimum horizontal coordinate of the first position coordinate and less than or equal to the maximum horizontal coordinate of the first position coordinate; a second constraint rule: the vertical coordinate of the barcode is greater than or equal to the minimum vertical coordinate of the first position coordinate and less than or equal to the maximum vertical coordinate of the first position coordinate; a third constraint rule: the area intersection-union ratio (IU) of the target barcode bounding box and the packaging box bounding box of the ROI is greater than or equal to a first preset IU threshold and less than or equal to a second preset IU threshold; the preset constraint rules are based on the first position coordinate and the third position coordinate, and the preset constraint rules are used to determine the intersection-union ratio. The UDI consumable identification result for each barcode corresponding to each boxed consumable is determined by setting constraint rules, including: based on the preset constraint rules, taking the target barcodes that conform to the first constraint rule, the second constraint rule, and the third constraint rule as target code segments, wherein the target boxed consumables corresponding to the target barcodes include at least one target code segment; for each target boxed consumable, aggregating all the target code segments of the target boxed consumables to generate a set of UDI code segments corresponding to the target boxed consumables; assembling the correspondence between each set of UDI code segments and each target boxed consumable to obtain the UDI consumable identification result.
[0011] Optionally, both the preset packaging box recognition model and the preset barcode recognition model are sub-models of the preset recognition model. The method further includes: acquiring a target training image and a baseline recognition result of the target training image; recognizing the target training image through an initial recognition model to obtain an initial recognition result; calculating a recognition loss value based on the initial recognition result and the bounding box regression loss function CIoU; and training the initial recognition model based on the recognition loss value to obtain the preset recognition model.
[0012] Optionally, the benchmark recognition result includes multiple benchmark bounding boxes. The step of calculating the recognition loss value based on the initial recognition result and the bounding box regression loss function CIoU includes: obtaining the benchmark coordinates of the target benchmark bounding box and the initial coordinates of the target initial bounding box corresponding to the target benchmark bounding box in the initial recognition result, wherein the target benchmark bounding box is any of the benchmark bounding boxes; calculating the training area intersection-union ratio and Euclidean distance between the target benchmark bounding box and the target initial bounding box based on the benchmark coordinates and the initial coordinates, and obtaining the diagonal distance of the minimum closure region containing the target benchmark bounding box and the target initial bounding box; and determining the recognition loss value based on preset weight parameters and preset size measurement parameters in the following manner: the recognition loss value = 1 - the training area intersection-union ratio + (the Euclidean distance / the diagonal distance)² + the preset weight parameter × the preset size measurement parameter.
[0013] In a second aspect, embodiments of this application provide a multi-target UDI consumable identification system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the multi-target UDI consumable identification system to perform the method described in the first aspect and any possible implementation thereof.
[0014] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a multi-target UDI consumable identification system, cause the multi-target UDI consumable identification system to perform the method described in the first aspect and any possible implementation thereof.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a multi-target UDI consumable identification system, cause the multi-target UDI consumable identification system to perform the method described in the first aspect and any possible implementation thereof.
[0016] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By employing dual-polarization imaging technology to acquire orthogonal polarized images and fusing them to generate anti-reflective images, combined with cascaded packaging box and barcode recognition models, and designing an intelligent matching process based on coordinate transformation and preset constraint rules, end-to-end fully automated identification of multi-target, highly reflective boxed UDI consumables was achieved. This eliminated thin-film reflection interference from an optical perspective, improved the detection accuracy of targets at different scales through a divide-and-conquer strategy, and ensured extremely high reliability of the association and matching between barcodes and consumables through rigorous spatial mapping and multiple verifications. This significantly improved the recognition success rate, accuracy, and robustness of the UDI traceability system in complex industrial scenarios.
[0017] By calculating the linear polarization degree of each pixel based on the Stokes vector principle and constructing a weighted fusion function with polarization degree as an adaptive measure to generate the final image, the specular reflection and diffuse reflection components can be accurately quantified and separated at the physical level. This achieves pixel-level precise suppression of highly reflective areas on the thin film surface, while preserving rich texture details in non-reflective areas. This fundamentally improves the image quality of the barcode area and provides a reliable data foundation for subsequent high-precision detection.
[0018] By accurately mapping the barcode coordinates within the ROI back to the global coordinate system and employing multiple constraint rules combining center point inclusion judgment and intersection-union verification for association determination, this method can effectively handle complex situations such as barcodes located on box edges, corners, and densely stacked consumables. It avoids the risk of erroneous association caused by simple coordinate overlap matching, ensuring that multiple UDI codes can be correctly aggregated and assigned to a unique corresponding consumable. This greatly improves the accuracy and completeness of data association, meeting the stringent requirements of medical device regulation for error-free traceability data.
[0019] By using CIoU as the bounding box regression loss function to train the cascaded detection model, the model not only considers the overlap area between the predicted box and the ground truth box during optimization, but also optimizes the consistency of their center point distance and aspect ratio. This training method significantly improves the model, especially the localization accuracy of the detection network for small target barcodes, and effectively reduces the localization drift caused by perspective distortion or film refraction, thus laying a solid foundation for subsequent accurate coordinate mapping and reliable matching association. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a multi-target UDI consumable identification method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the process for generating a reflective fused image according to an embodiment of this application; Figure 3 This is a schematic diagram of the process for training the preset recognition model provided in the embodiments of this application; Figure 4This is a schematic diagram of the structure of a multi-target UDI consumable identification system provided in an embodiment of this application.
[0021] Explanation of reference numerals in the attached figures: 601, Central Processing Unit; 602, Read-Only Memory; 603, Random Access Memory; 604, Bus; 605, Input / Output Interface; 606, Input Section; 607, Output Section; 608, Storage Section; 609, Communication Section; 610, Driver; 611, Removable Media. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] In the description of the embodiments of this application, words such as "illustrative," "for example," or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "illustrative," "for example," or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "illustrative," "for example," or "for example" is intended to present the relevant concepts in a specific manner.
[0024] In the description of the embodiments of this application, the terms "first, second, third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0025] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0026] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0027] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0028] In related technologies, images are acquired using a single RGB industrial camera, and the positions of the packaging box and barcode are simultaneously detected using a model. Simple matching is then performed based on spatial overlap. However, this approach cannot fundamentally suppress the strong reflection on the film surface, leading to overexposure or texture loss in the barcode area and insufficient recognition stability. A single detection network also struggles to simultaneously learn the features of large-scale packaging boxes and small-scale barcodes. In particular, barcode miss detection and positioning deviations are prominent at complex geometric structures such as box edges and corners, reducing the robustness of UDI consumable identification in highly reflective, multi-target, and complex spatial layout environments. To address these issues, this application provides a multi-target UDI consumable identification method, system, medium, and product, which improves the robustness of multi-target UDI consumable identification in complex industrial scenarios.
[0029] Figure 1 This is a flowchart illustrating a multi-target UDI consumable identification method provided in an embodiment of this application.
[0030] This invention discloses a multi-target UDI consumable identification method, such as... Figure 1 As shown, the steps include the following.
[0031] S101. Obtain a first polarization image of the target scene at a first polarization angle and a second polarization image at a second polarization angle. The target scene includes multiple boxed consumables, and the first polarization angle and the second polarization angle are orthogonal.
[0032] Specifically, the system is configured with two rotatable polarizers with orthogonal polarization directions, i.e., linearly polarized light sources (e.g., 0° and 90° or 45° and 135°, etc.), and a color industrial camera with a linear polarizing mirror mounted at the front end. During acquisition, the light source with the first polarization angle is first turned on, and the polarizing mirror in front of the camera is adjusted to the same polarization angle as the light source to acquire the first polarized image. Then, the light source with the second polarization angle is turned on, and the camera's polarizing mirror is synchronously adjusted to the same polarization angle as the light source to acquire the second polarized image. These two images are captured continuously within a very short time interval of the same target scene containing multiple boxed consumables, thus ensuring the consistency of the scene content. Finally, two scene images with orthogonal polarization states are obtained.
[0033] UDI (Unique Device Identification) refers to a standardized coding system for uniquely identifying medical devices globally. It comprises two parts: the Device Identifier (DI) and the Manufacturing Identifier (PI). The Device Identifier represents static information about the product, uniquely identifying the manufacturer, model, and specifications of the medical device, similar to a product's ID number. The Manufacturing Identifier represents dynamic information about the product, identifying the production batch, serial number, expiration date, and other production process-related information, similar to a product's travel code. The target scene refers to the entire physical area covered by the field of view of the industrial image acquisition system, containing multiple objects to be identified and processed, such as in a pharmaceutical production line or intelligent drug management scenarios. The first polarization angle refers to the first specific azimuth angle of the electric field vector vibration direction of linearly polarized light, used to represent the illumination and acquisition conditions for the first polarization state. The first polarization image refers to the two-dimensional digital image acquired when both the illumination source and the camera polarizer are set to the first polarization angle; this image records the light intensity distribution of the scene under that specific polarization state. The second polarization angle refers to a specific azimuth angle perpendicular to the first polarization angle in direction, used to represent polarized illumination and acquisition conditions orthogonal to the first state. The second polarized image refers to a two-dimensional digital image acquired when both the illumination source and the camera polarizer are set to the second polarization angle. Boxed consumables refer to medical device consumables packaged in rigid or semi-rigid materials such as cardboard or plastic boxes, whose packaging surface is typically covered with a high-gloss film layer.
[0034] Through the above embodiments, orthogonal polarization imaging technology is used to specifically collect information about the scene under different polarization states from the source of optical imaging. Since diffuse reflection light from the object surface will depolarize, while specular reflection and Fresnel reflection generated by the thin film will retain polarization characteristics, the orthogonal polarization image responds very differently to these two types of light. This provides an indispensable and complementary original data foundation for subsequent calculation and fusion to completely suppress thin film reflection interference, which is the primary key technology link to achieve highly robust recognition.
[0035] S102. Based on the first polarization image and the second polarization image, generate a reflection-free fused image of the target scene.
[0036] Specifically, by reading the acquired first and second polarization images, the intensity or brightness values of these two input images can be analyzed and compared point-by-point (i.e., pixel-by-pixel). At each identical pixel coordinate position, the pixel value of the first polarization image and the pixel value of the second polarization image are extracted respectively. Then, these two values are compared, and the one with the smaller intensity or brightness value is selected. The logic behind this selection is that reflected light is usually superimposed on the light of the object itself as additional light, manifesting as excessive brightness or whitening in local areas. At the same position in both images, the point with the lower pixel value is more likely to represent the object itself after the influence of reflected light has been suppressed. The true visual information is used to assign the pixel with the lower value that wins in the comparison to the pixel at the same coordinate position in the final anti-reflective fused image. This operation traverses every pixel in the two input images. By performing the "compare and take the minimum value" operation on all pixels, a completely new image is finally constructed. Each pixel of the final anti-reflective fused image comes from the version with the least reflection effect in the two original polarized images. Thus, while fully preserving the original details and colors of the scene, it effectively eliminates or greatly reduces distracting glare and reflections, making the scene content that was originally obscured by reflections clearly visible.
[0037] Among them, a non-reflective fused image refers to a two-dimensional digital image synthesized through specific image processing techniques. Its visual characteristics are that it basically eliminates specular highlights and Fresnel reflections caused by smooth surfaces (such as thin films), and is used to represent an enhanced image that can more clearly present the underlying texture of objects (such as barcodes) that are covered by reflected light.
[0038] Figure 2 This is a schematic diagram of the process for generating a reflection-free fused image provided in an embodiment of this application.
[0039] Based on the above embodiments, as an optional embodiment, see [link to embodiment]. Figure 2 ,against Figure 1 The illustrated step S102 can be achieved through... Figure 2 The steps S201-S205 are implemented, and will be explained in detail below.
[0040] S201. Obtain the first pixel intensity set of the first polarization image and the second pixel intensity set of the second polarization image. The first pixel intensity set includes multiple first pixel intensities corresponding to the first polarization image, and the second pixel intensity set includes multiple second pixel intensities corresponding to the second polarization image.
[0041] Specifically, the system loads two previously acquired, registered orthogonal polarization image data from the storage medium or memory buffer. It then parses the format of these two digital images and extracts the intensity value of each pixel. For the first polarization image, the system iterates through all its pixels, reading and recording the intensity value of each pixel in a specific color channel (e.g., to retain maximum information, it may convert the brightness value to grayscale, or process the R, G, and B channels separately; no specific restrictions are placed here). All these values are organized into a first pixel intensity set. Completely synchronously, the system performs the same operation on the second polarization image, iterating through all its pixels and reading and recording the intensity values to form a second pixel intensity set. These two sets are one-to-one corresponding and strictly aligned in pixel coordinates, ensuring accurate pixel-level calculations can be performed subsequently.
[0042] The first pixel intensity set is a data structure used to store and represent the light intensity information of all pixels in the first polarization image. The number of elements in this set is equal to the total number of pixels in the image. The first pixel intensity specifically refers to the intensity value of a pixel at a specified coordinate position in the first polarization image. The second pixel intensity set is a data structure with the same structure as the first pixel intensity set, but whose data source is the second polarization image. It is used to store the intensity information of all pixels in the second polarization image. The second pixel intensity specifically refers to the intensity value of a pixel at the same coordinate position in the second polarization image as in the first polarization image.
[0043] S202. For each pixel, subtract the second pixel intensity of the target pixel from the first pixel intensity to obtain the target pixel difference value. The target pixel is any pixel.
[0044] Specifically, after obtaining the first and second pixel intensity sets with strict registration, a traversal loop is started to process each pixel position in the image sequentially. For the target pixel currently being processed, the system retrieves the first pixel intensity (Ii) of that coordinate point from its corresponding first pixel intensity set. ∥ At the same time, the second pixel intensity (I) at the same coordinate point is retrieved from the second pixel intensity set. ⊥Subsequently, the system performs a scalar subtraction operation, subtracting the intensity of the second pixel from the intensity of the first pixel. The result of this subtraction operation is defined as the target pixel difference (I0). ∥ -I ⊥ This process is repeated for each pixel in the image, eventually generating a target pixel difference map of the same size as the original image. Each pixel value in this map represents the quantized result of the light intensity difference at the corresponding location under two orthogonal polarization states.
[0045] Here, the target pixel refers to a single pixel selected and processed within the current computation cycle; it is an arbitrary element in the set of all pixels in the image. The target pixel difference is the direct result of the subtraction operation described above; it is a scalar value used to quantify the difference in the intensity of reflected light from the target pixel under two orthogonal polarization states.
[0046] For example, suppose the system is processing a target pixel with coordinates (150, 300). The I value is retrieved from the first pixel intensity set. ∥ (150, 300) = 220, find I from the second pixel intensity set. ⊥ (150, 300) = 70. Perform a subtraction operation: 220 - 70 = 150. The resulting value 150 is the target pixel difference of the target pixel (150, 300). Store this result 150 at the (150, 300) position in a data structure dedicated to storing difference results. Then, automatically move to the next pixel, for example, (150, 301), and continue to perform the same search and subtraction operation until all pixels of the entire image have been processed, thus obtaining a complete difference map.
[0047] S203. Calculate the sum of the first pixel intensity and the second pixel intensity of the target pixel to obtain the target pixel sum value.
[0048] Specifically, an addition operation will be performed simultaneously, adding the retrieved first pixel intensity (I) to the total intensity. ∥ The value of ) and the intensity of the second pixel (I) ⊥ The values of the target pixels are added together, and the result is defined as the sum of the target pixel values (I). ∥ +I ⊥ This calculation process also iterates through every pixel in the image, eventually generating a target pixel and value map of the same size as the original image. Each pixel value in this map represents an approximate measure of the total light intensity at the corresponding location under two orthogonal polarization states.
[0049] The target pixel sum is the direct result of the above addition operation. It is a scalar value used to represent the sum of the reflected light intensity of the target pixel under two orthogonal polarization states.
[0050] For example, continuing with the target pixel (150, 300), given I... ∥ (150, 300) = 220, I ⊥ (150, 300) = 70, perform addition: 220 + 70 = 290. The resulting value 290 is the target pixel sum of the target pixel (150, 300). This result 290 is stored in the corresponding position (150, 300) of the data structure dedicated to storing sum results. This process traverses all pixels and finally forms a complete sum map.
[0051] S204. Calculate the ratio of the target pixel difference to the target pixel sum to obtain the linear polarization degree of the target pixel.
[0052] Specifically, the difference between the target pixels is used as the numerator, and the sum of the target pixel values is used as the denominator. The two are divided to obtain a ratio, which is defined as the linear polarization degree of the target pixel. This calculation process is applied to every pixel in the image, ultimately generating a linear polarization degree map with the exact same size as the original image. Each pixel value in this map is a dimensionless value between 0 and 1, quantitatively describing the degree of polarization of light at the corresponding location.
[0053] The linear polarization degree (DoLP) is a dimensionless physical quantity used to represent the proportion of fully polarized light in partially polarized light. Its value range is [0,1], where 0 represents completely unpolarized light (natural light) and 1 represents completely linearly polarized light. In this application scenario, a high linear polarization degree value usually indicates a strong specular reflection area, while a low value area corresponds to an object surface dominated by diffuse reflection.
[0054] S205. Based on each linear polarization degree and a preset weight fusion function, fuse the first polarization image and the second polarization image to obtain a reflection-free fused image.
[0055] Specifically, the calculated full-image linear polarization degree map, the first polarization image, and the second polarization image are used as core inputs. When the linear polarization degree value of a pixel is high, the preset weight fusion function will assign a very low weight to the pixel in the image containing stronger reflection (usually the brighter one), while assigning a very high weight to the pixel in the other image. Conversely, if the linear polarization degree value of a pixel is low (indicating almost no reflection), the preset weight fusion function may assign similar weights to the pixels in the two images (e.g., 0.5 each), essentially averaging them. This is achieved through a linear polarization fusion function... The weights, dynamically determined by the degree of polarization, are used to perform a pixel-by-pixel weighted average fusion of the first and second polarized images. The value of each new pixel is determined by the formula: fused pixel value = first weight × first polarized image pixel value + second weight × second polarized image pixel value. Since the weights change smoothly according to the intensity of reflection, this method can suppress reflection very naturally and avoid harsh edges or loss of detail that may occur at the boundary between reflective and non-reflective areas. The resulting reflection-free fused image not only eliminates reflection but also performs better in terms of detail preservation and smoothness of transition, and is more in line with physical reality.
[0056] The preset weighted fusion function is a mathematical formula that is determined in advance through theoretical derivation or experimental optimization. Its input is the linear polarization degree, and its output is the weight coefficient used to fuse the intensity values of the first polarization image and the second polarization image, thus establishing a mapping relationship from polarization characteristics to image processing strategies.
[0057] Through the above embodiments, linear polarization is used as a precise guide, and adaptive, pixel-level image fusion is achieved through a preset weighted fusion function. This is not a simple averaging, but rather intelligently suppresses the noisy specular reflection component in reflective areas and effectively preserves and enhances useful texture information in non-reflective areas. It effectively removes film reflections at the algorithm level, generating high-quality, high-definition enhanced images. This provides highly reliable input for subsequent target detection steps and directly solves the problem of feature loss caused by reflection. It is one of the decisive links in achieving high robustness recognition in the entire technical solution.
[0058] Based on the above embodiments, as an optional embodiment, see [link to embodiment]. Figure 2 ,against Figure 1 The step S205 shown can be achieved through... Figure 2 The steps S2051-S2055 are implemented, and will be explained in detail below.
[0059] S2051. The difference between 1 and the linear polarization degree of the target pixel is used as the polarization compensation coefficient of the target pixel.
[0060] Specifically, when the linear polarization degree value of the target pixel has been obtained and it is ready to be substituted into the preset weighted fusion function for image fusion, this compensation coefficient needs to be calculated first. A subtraction operation is performed on the linear polarization degree value of the target pixel, that is, the linear polarization degree value is subtracted from the value 1. The result of this subtraction operation is defined as the polarization compensation coefficient (i.e., 1-DoLP) of the target pixel. This operation is performed on each pixel in the final fused image one by one, thereby generating a polarization compensation coefficient map with the same size as the linear polarization degree map. Each value in this map will be used as a key parameter to directly participate in the subsequent pixel intensity weighted fusion calculation.
[0061] The polarization compensation coefficient, which is the direct result of the subtraction operation, is a scalar value ranging from 0 to 1, used to dynamically adjust the first polarization image (I) during image fusion. ∥ The contribution weight of the image is determined by the value of the image itself; the larger the value, the higher the degree of adoption of the image information.
[0062] S2052. Calculate the product of the first preset illumination compensation coefficient, the polarization compensation coefficient, and the first pixel intensity of the target pixel to obtain the first compensated pixel intensity of the target pixel after compensation.
[0063] Specifically, after obtaining the polarization compensation coefficient of the target pixel, a fixed first preset illumination compensation coefficient is read from the system configuration, and the first pixel intensity corresponding to the target pixel in the first polarization image is retrieved from the storage. The system performs a continuous multiplication operation, multiplying the three values of the first preset illumination compensation coefficient, the polarization compensation coefficient, and the first pixel intensity. This product is defined as the first compensated pixel intensity. The product of the first preset illumination compensation coefficient and the polarization compensation coefficient can be determined as the first weight in step S205. This process is performed one by one for each pixel in the fused image. The output result represents the effective light intensity contribution value extracted from the first polarization image after calibration and weighting, taking into account the ambient lighting conditions and local polarization characteristics.
[0064] The first preset illumination compensation coefficient is a pre-defined constant weighting factor used to globally adjust the contribution level of the first polarization image in the entire fusion result and compensate for light intensity changes caused by ambient lighting conditions. It is a constant determined experimentally, for example, 0.5. The first compensated pixel intensity refers to the final result of the above multiplication operation. It is a scalar value representing the effective intensity value of the target pixel in the first polarization image after illumination and polarization compensation and weighting, which is considered usable for the final fusion.
[0065] Through the above embodiments, by combining the static environmental compensation coefficient with the dynamic local polarization compensation coefficient, the original intensity of the first polarization image is finely weighted, ensuring that the contribution of the first polarization image is significantly suppressed in areas with strong reflection (low polarization compensation coefficient), thereby reducing the introduction of specular reflection noise. In areas with weak reflection and useful texture (high polarization compensation coefficient), the information of the first polarization image is moderately preserved to supplement scene details.
[0066] S2053. Calculate the product of the second preset illumination compensation coefficient and the second pixel intensity of the target pixel to obtain the second compensated pixel intensity of the target pixel.
[0067] Specifically, when processing the same target pixel, a fixed second preset illumination compensation coefficient is read from the configuration parameters, and the second pixel intensity corresponding to the target pixel in the second polarization image is retrieved from the storage. The second preset illumination compensation coefficient is multiplied by the second pixel intensity, and this product is defined as the second compensated pixel intensity. The second preset illumination compensation coefficient can be determined as the second weight in step S205. This calculation process is performed in parallel or sequentially for each pixel in the image, and its output represents the light intensity contribution value extracted from the second polarization image as the basis for fusion after considering the global illumination conditions.
[0068] The second preset illumination compensation coefficient is a pre-defined constant weighting factor used to globally adjust the baseline contribution level of the second polarization image in the entire fusion result and compensate for light intensity changes caused by ambient lighting conditions. It is a constant determined experimentally, for example, 1.0. The second compensated pixel intensity refers to the final result of the above multiplication operation. It is a scalar value representing the baseline intensity value of the target pixel in the second polarization image, which is considered usable for the final fusion after illumination compensation weighting.
[0069] The above embodiments establish the fundamental role of the second polarization image in the fusion result. The setting of the second preset illumination compensation coefficient allows the system to flexibly adjust the intensity level of the second polarization image (which mainly carries diffuse reflection texture information unaffected by polarization) in the final output according to the actual application scenario. Using the intensity of the second compensated pixel as the basis for fusion ensures that even in highly reflective areas, the final fused image can clearly retain the inherent surface texture features of the object (such as barcodes). This provides a fundamental guarantee that the entire method can effectively eliminate reflections without losing key identification information. This step is the stable cornerstone for constructing high-quality, reflection-free fused images.
[0070] S2054. Based on the first compensated pixel intensity and the second compensated pixel intensity, the fused target pixel is obtained.
[0071] Specifically, after the calculation of the first and second compensated pixel intensities of the same target pixel is completed, the two intensity values, which are from different polarization images and processed by their respective compensation paths, are summed. The result of this addition operation, namely the sum of the first and second compensated pixel intensities, is directly defined as the final intensity value of the fused target pixel. It can be referred to formula (1). This addition operation traverses every pixel position in the image and combines all the fused target pixels in the original coordinate order, thus finally generating the desired anti-reflective fused image.
[0072] I final (x,y)=α×(1−DoLP(x,y))×I ∥ (x,y)+β×I (x,y)(1) Among them, I final (x,y) represents the target pixel after fusion, which is the final result of the above addition operation. It is a scalar value used to represent the intensity value of the pixel located at the current target pixel coordinate position in the final generated anti-reflection fused image. α is the first preset illumination compensation coefficient, and (1−DoLP(x,y)) is the polarization compensation coefficient. α×(1−DoLP(x,y)) can be used as the first weight in step S205. ∥ (x,y) represents the first pixel intensity of the target pixel, and β is the second preset illumination compensation coefficient, which can be used as the second weight in step S205. (x,y) represents the second pixel intensity of the target pixel.
[0073] Through the above embodiments, optimal information integration is achieved by superimposing the outputs of two independent compensation paths. The second compensation pixel intensity provides a clear and stable diffuse texture base, ensuring the reliable representation of basic surface features (such as barcodes). The first compensation pixel intensity provides calibrated and beneficial detail supplementation in non-reflective areas. This additive fusion method ensures that in highly reflective areas, the first compensation pixel intensity is effectively suppressed, and the final result mainly depends on the pure second compensation pixel intensity, thus eliminating reflection. In weakly reflective areas, both work together to enrich image details. This step is the final output stage of the entire polarization image fusion algorithm. It successfully integrates the results of all previous processing steps, directly producing a high-quality, non-reflective input image that significantly improves the performance of subsequent recognition algorithms.
[0074] S2055. Based on each fused target pixel, a reflection-free fused image is obtained.
[0075] Specifically, after generating the corresponding fused target pixel intensity value for each pixel position in the original image, all the calculated fused target pixel intensity values are systematically and orderly filled into a newly created image data matrix or buffer according to their original two-dimensional spatial coordinate relationship. This new image data matrix is completely consistent with the original first polarization image and second polarization image in terms of dimension, but the data of each pixel has been replaced with the new value after polarization fusion processing. When the last fused target pixel is placed in its correct coordinates, a complete and brand-new digital image is generated. This image is the direct output of this technical solution - the anti-reflection fused image.
[0076] S103. Based on the preset packaging box recognition model, identify the target region of interest (ROI) and the first position coordinates of the target boxed consumable in the non-reflective fusion image. The target boxed consumable is any boxed consumable in the target scene.
[0077] Specifically, the anti-reflective fused image generated in the aforementioned steps is used as input and fed into a pre-trained pre-defined packaging box recognition model (e.g., a deep learning model based on the YOLOv8n architecture and optimized for large targets). This model performs inference analysis on the entire input image, traversing all possible target packaging boxes. For each detected target packaging box, the model outputs its position information in the anti-reflective fused image. This position is usually represented by a rectangular bounding box that tightly surrounds the packaging box. This bounding box is the target region of interest (ROI) corresponding to the target packaging box. The coordinate data describing the specific position and size of this rectangular bounding box in the global coordinate system of the anti-reflective fused image (e.g., the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the width and height of the box; or the minimum x-coordinate of the box) are used. min Minimum ordinate Y min Maximum x-coordinate max Maximum ordinate Y max The central x-coordinate mid and the central ordinate Y mid (etc.), then is defined as the first position coordinate of the target ROI.
[0078] The pre-defined packaging box recognition model refers to a deep learning network pre-trained using a large amount of image data labeled with packaging box bounding boxes. Its function is to automatically locate the packaging box target from the input image and provide its position. The target boxed consumable refers to any individual boxed consumable successfully identified by the model within the current detection period. The target region of interest (ROI) is a sub-image region cropped from the non-reflective fused image, bounded by the aforementioned bounding box. This region contains complete visual information of a specific target boxed consumable. The first position coordinates are a set of data used to accurately locate the spatial position of the target ROI in its source image (i.e., the non-reflective fused image), typically composed of the vertex coordinates or center point coordinates and width and height of the bounding box.
[0079] Through the above embodiments, the system utilizes a deep learning model specifically designed for packaging box detection to efficiently and accurately complete the coarse localization of all boxed consumables in the scene on high-quality input images. By determining the target ROI, the complex global detection problem is decomposed into multiple more easily handled local detection problems, effectively improving the processing efficiency and detection accuracy in multi-target scenes.
[0080] S104. Based on the preset barcode recognition model, identify the set of barcodes in the target ROI and locate the second position coordinate of the target barcode in the target ROI. The target barcode is any barcode in the barcode set.
[0081] Specifically, the target ROI sub-image obtained from the previous step, corresponding to a specific boxed consumable, is used as input data and fed into another specially trained pre-defined barcode recognition model (e.g., a deep learning model based on the YOLOv8n architecture and optimized for small targets). Inference analysis is performed within this cropped sub-image area to search for and identify all barcodes present within the target ROI. All successfully detected barcodes together constitute the barcode set corresponding to the target ROI. For each target barcode in this set, the model outputs its position information in the coordinate system of the target ROI sub-image itself. This position is also represented by a rectangular bounding box tightly surrounding the barcode, describing the specific position and size of this bounding box in the target ROI sub-image coordinate system (e.g., the x-coordinate of the box's center point u). j and the vertical coordinate v j If the coordinates of the top left corner, width, and height of the frame are used, then the second position coordinates of the target barcode are defined.
[0082] The preset barcode recognition model refers to a deep learning network pre-trained using a large amount of image data labeled with barcode bounding boxes. Its function is to automatically locate the barcode target from the input image and provide its position, especially optimized for small targets. The barcode set refers to the collection of all barcode instances successfully recognized by the model within the current target ROI. The target barcode refers to any individual barcode in the barcode set. The second position coordinates are a set of data used to accurately locate the position of the target barcode within the coordinate system of its respective target ROI sub-image.
[0083] Through the above embodiments, the cascaded detection strategy decomposes the complex task of finding small barcodes in the global image into two sub-tasks: first finding large boxes in the global image, and then finding small barcodes in the local images of each box. The barcode recognition model is preset to work on the pure target ROI sub-image, which greatly improves the detection accuracy and recall rate of small target barcodes and effectively avoids missed detections.
[0084] S105. Based on the first position coordinates and the second position coordinates, the second position coordinates are transformed to obtain the third position coordinates of the target barcode in the non-reflective fused image.
[0085] Specifically, after obtaining the first position coordinates of the target ROI in the anti-reflection fused image and the second position coordinates of the target barcode in its respective target ROI sub-image, a translation transformation is used to convert the second position coordinates of the target barcode in the local coordinate system to the global coordinate system of the anti-reflection fused image. The starting offset of the target ROI in the global coordinate system and the position of the target barcode in the local coordinate system can be vector-added. The result of the operation is the absolute position of the target barcode in the unified reference system of the anti-reflection fused image, and this position is defined as the third position coordinate.
[0086] Coordinate transformation refers to the process of changing the coordinate representation of a point in one coordinate system to another coordinate system through specific mathematical operations (such as translation). Third position coordinates refer to the coordinate representation of the target barcode in the global coordinate system of the reflection-free fused image after coordinate transformation, establishing a direct spatial connection between the barcode and the original panoramic image.
[0087] Based on the above embodiments, as an optional embodiment, for Figure 1 The step S105 shown can be implemented through steps S301-S303, which will be explained in detail below.
[0088] S301. Calculate the sum of the minimum abscissa of the first position coordinate and the abscissa of the center point of the second position coordinate to obtain the abscissa of the target barcode in the non-reflective fused image.
[0089] Specifically, the first set of location coordinate data corresponding to the target ROI is read from storage, and the parameter describing the starting position of the ROI bounding box in the horizontal direction—the minimum x-coordinate—is extracted from it. min Simultaneously, from the second position coordinate data set of the target barcode, the parameter describing the horizontal center position of the barcode within its respective ROI sub-map—the x-coordinate u of the center point—is extracted. j The minimum x-coordinate X min The numerical value and the x-coordinate of the center point u j The values are added together to obtain the horizontal coordinates of the target barcode in the global coordinate system of the non-reflective fused image, i.e., the horizontal coordinate x of the barcode. j .
[0090] In this context, the minimum x-coordinate of the first position coordinate refers to the x-axis coordinate of the leftmost edge of the bounding box of the target ROI within the global coordinate system of its corresponding anti-glare blended image, defining the horizontal starting position of the ROI in global space. The x-coordinate of the center point of the second position coordinate refers to the x-axis coordinate of the center point of the bounding box of the target barcode within the local coordinate system of its corresponding target ROI sub-image, representing the horizontal center position of the barcode within the sub-image. The x-coordinate of the target barcode in the anti-glare blended image is a scalar value used to represent the precise horizontal position (x-coordinate) of the target barcode's center point within the global coordinate system of the anti-glare blended image.
[0091] S302. Calculate the sum of the minimum ordinate of the first position coordinate and the ordinate of the center point of the second position coordinate to obtain the ordinate of the target barcode in the non-reflective fused image.
[0092] Specifically, the minimum ordinate Y-coordinate, which describes the starting position of the ROI bounding box in the vertical direction, is extracted from the first location coordinate data set. min Furthermore, from the second position coordinate data set of the target barcode, the parameter describing the vertical center position of the barcode within its respective ROI sub-map, namely the center point ordinate v, is extracted. j The minimum ordinate Y min The numerical value and the ordinate of the center point v j The values are added together to obtain the vertical coordinates of the target barcode in the global coordinate system of the non-reflective fused image, i.e., the barcode's ordinate y. j .
[0093] In this context, the minimum y-coordinate of the first position coordinate refers to the y-axis coordinate of the uppermost edge of the bounding box of the target ROI within the global coordinate system of its corresponding anti-glare fused image, defining the vertical starting position of the ROI in global space. The y-coordinate of the center point of the second position coordinate refers to the y-axis coordinate of the center point of the bounding box of the target barcode within the local coordinate system of its corresponding target ROI sub-image, representing the vertical center position of the barcode within the sub-image. The y-coordinate of the target barcode in the anti-glare fused image is a scalar value used to represent the precise vertical position of the target barcode's center point within the global coordinate system of the anti-glare fused image.
[0094] S303. Use the horizontal and vertical coordinates of the barcode as the third position coordinates of the target barcode.
[0095] Specifically, the horizontal coordinate x of the barcode is... j with barcode y-axis j These two scalar values are logically correlated and combined to create a data structure for representing a point in two-dimensional space, and the horizontal coordinate of the barcode is set to x. j And the barcode's y-axis j The values are assigned to the horizontal and vertical axis fields of the data structure, respectively. This newly constructed two-dimensional point coordinate, containing precise horizontal and vertical coordinate values, is formally defined as the third position coordinate of the target barcode in the global coordinate system of the anti-reflective fused image. This process generates a unique absolute position identifier for each target barcode in a unified spatial reference system.
[0096] The horizontal coordinate of the barcode refers to the x-axis coordinate of the center point of the target barcode in the global coordinate system of the anti-reflective fused image, accurately describing the absolute position of the barcode in the horizontal direction. The vertical coordinate of the barcode refers to the y-axis coordinate of the center point of the target barcode in the global coordinate system of the anti-reflective fused image, accurately describing the absolute position of the barcode in the vertical direction.
[0097] S106. Based on the first position coordinate and the third position coordinate, and through preset constraint rules, determine the UDI consumable identification result of each barcode corresponding to each boxed consumable.
[0098] Specifically, the algorithm iterates through each box of consumables and each barcode combination. For each pair, it activates and strictly enforces a series of preset constraint rules. These rules are based on logical judgments of physical world layout and industry packaging practices. Their complexity and content can be very rich. For example, they may include: First, spatial proximity rules, which determine that a barcode is most likely to belong to the consumable box that is physically closest to it. The algorithm calculates the distance between each box and each barcode and prioritizes the pair with the smallest distance. Second, geometric inclusion and adjacency rules, which determine whether the position coordinates of a barcode fall completely inside the position coordinates of a box or are close to its specific edge (such as directly below or to the right). This is usually stronger evidence of association than simple distance. During execution, the first position coordinates of each boxed consumable are used as a reference to traverse the third position coordinates of each barcode in the barcode set. The aforementioned preset constraint rules are used to check and evaluate each box-barcode potential pairing. Only when the spatial relationship between a box and a barcode best matches or uniquely satisfies these constraints will they be established as a valid, high-confidence pairing. Through this rigorous screening and matching process, an identification barcode is finally found for each boxed consumable. All the corresponding identification barcodes are aggregated to generate an accurate UDI consumable identification result. This result clearly binds each physical entity (box) to its corresponding digital identity (UDI information).
[0099] Among them, the preset constraint rules refer to a set of predefined logical conditions used to determine whether a barcode belongs to a packaging box. These rules are usually based on spatial geometric relationships to ensure the accuracy of matching. The UDI consumable identification result is the final output of this method. It is a structured dataset that clearly lists the information of each identified packaging box and all the correctly associated UDI code segments, forming the data foundation for automated traceability and quality control.
[0100] Based on the above embodiments, as an optional embodiment, the preset constraint rules include: a first constraint rule: the barcode's horizontal coordinate is greater than or equal to the minimum horizontal coordinate of the first position coordinate and less than or equal to the maximum horizontal coordinate of the first position coordinate; a second constraint rule: the barcode's vertical coordinate is greater than or equal to the minimum vertical coordinate of the first position coordinate and less than or equal to the maximum vertical coordinate of the first position coordinate; a third constraint rule: the area intersection-union ratio (IU) of the target barcode's bounding box and the ROI's packaging box bounding box is greater than or equal to a first preset IU threshold and less than or equal to a second preset IU threshold.
[0101] These three pre-defined constraint rules together form a rigorous logical judgment system used to accurately determine whether a target barcode correctly belongs to a specific target boxed consumable within a unified spatial coordinate system. The first constraint rule filters in the horizontal dimension, checking whether the center point of the barcode falls between the left and right boundaries of the packaging box. The second constraint rule filters in the vertical dimension, checking whether the center point of the barcode falls between the top and bottom boundaries of the packaging box. For candidate pairs that pass the first two steps, the third constraint rule performs a final verification from the perspective of two-dimensional area overlap, ensuring that a sufficient portion of the barcode itself is indeed located on the packaging box it claims to belong to, while excluding anomalies such as complete embedding or excessively large areas. Only when a target barcode simultaneously satisfies all three rules does the system ultimately confirm its affiliation with the target boxed consumable.
[0102] The first constraint rule is a condition based on horizontal spatial inclusion, stating that the center point of a barcode must be within the left and right boundaries of the packaging box in the horizontal direction. The maximum x-coordinate of the first position coordinate refers to the rightmost x-axis coordinate of the target boxed consumable's bounding box in the global coordinate system. The second constraint rule is a condition based on vertical spatial inclusion, stating that the center point of a barcode must be within the top and bottom boundaries of the packaging box in the vertical direction. The maximum y-coordinate of the first position coordinate refers to the bottommost y-axis coordinate of the target boxed consumable's bounding box in the global coordinate system. The third constraint rule is a refined condition based on the degree of overlap of two-dimensional regions. The barcode bounding box of the target barcode is defined in the global coordinate system as the rectangular region tightly surrounding the target barcode. The packaging box bounding box of the ROI is the rectangular region tightly surrounding the target boxed consumable, defined by the first position coordinate. The area intersection-union ratio (IUU) is a measure of the degree of overlap between two bounding boxes, calculated as: intersection area / union area. The first preset intersection-union (IU) threshold is a lower limit (e.g., 0.05) to ensure that at least a portion of the barcode is on top of the packaging box, preventing mismatches of background barcodes far from the box. The second preset IU threshold is an upper limit (e.g., 1.0) to handle extreme cases, with a theoretical upper limit of 1, representing complete overlap.
[0103] against Figure 1 Step S106 shown can be implemented through steps S401-S403, which will be explained in detail below.
[0104] S401. Based on preset constraint rules, target barcodes that conform to the first constraint rule, the second constraint rule, and the third constraint rule are used as target code segments, and the target boxed consumables corresponding to the target barcodes include at least one target code segment.
[0105] Specifically, for each target boxed consumable, all target barcodes that meet the following conditions are selected: the x-coordinate of its center point satisfies the first constraint rule; the y-coordinate of its center point satisfies the second constraint rule; and the area intersection-union ratio (IUU) of its bounding box with the current consumable's bounding box satisfies the third constraint rule. All target barcodes that pass these three verifications and belong to the same target boxed consumable are uniformly defined as the target code segment for that consumable. The system ultimately maintains a target code segment list for each target boxed consumable, which contains at least one element, namely, the preliminary structure of a complete UDI consumable identification result at the code segment level.
[0106] The target code segment refers to the identity obtained when a target barcode is officially confirmed as a legitimate component (UDI information segment) of its target boxed consumable after passing all association verifications. It emphasizes that the barcode is no longer an isolated detection result, but a valid data unit that is bound to a specific packaging box. The target boxed consumable corresponding to the target barcode includes at least one target code segment, which describes the result status after a successful match: a successfully identified packaging box must have at least one (usually multiple) correctly matched UDI code segments associated with its attributes.
[0107] For example, for a certain consumable box (Box1), the barcode Code is found. A Code C It simultaneously satisfies the triple constraint rules, and the Code B Not satisfied. The system will then assign the code... A and Code C Identify the target code segment as Box1 and establish its attribution relationship: Box1→[Code] A Code C For consumable Box2, only the Code was found. B If all rules are met, an attribution relationship is established: Box2 → [Code] B At this point, each consumable has its own set of target code segments (Box1 has 2, Box2 has 1).
[0108] S402. For each target boxed consumable, aggregate all target code segments of the target boxed consumable to generate a set of UDI code segments corresponding to the target boxed consumable.
[0109] Specifically, for each identified target boxed consumable, the UDI data content represented by all code segments (e.g., the original string information obtained through barcode decoding) is extracted from the list of target code segments associated with it. These scattered UDI data fragments belonging to the same physical packaging box are then aggregated and organized into a logically unified data group. This data group, which integrates all the UDI information of a specific target boxed consumable, is defined as the UDI code segment set of that consumable.
[0110] The UDI code segment set corresponding to the target boxed consumable is the final output of this step. It contains all the UDI code segment information (such as DI, PI, etc.) of a specific target boxed consumable, which constitutes the complete digital identity of the consumable in the information system.
[0111] For example, continuing from the previous example, the system processes consumable Box1, which is already associated with two target code segments: Code A and Code C The system processes the code. A Decode the string to obtain DI123456; [the code is missing from the original text]. C Decoding yields the string PIBATCH0012025. The system then aggregates these two strings to form the UDI code segment set for Box1, which can be represented as {DI: DI123456, PI: PIBATCH0012025} or similar structured data. Similarly, the target code segment Code is aggregated for Box2. B Decoded data.
[0112] Through the above embodiments, a leap from geometric spatial association to information logical integration has been achieved. Multiple barcodes belonging to the same physical entity are identified through visual and spatial relationships and merged at the information level to generate a digital file representing the complete identity of the entity. This set of UDI code segments is the core business output of the entire automated identification process. It enables each consumable to have traceable and complete UDI information, directly meeting the mandatory requirements of medical device supervision that the unique identifier of a product must include multiple data such as DI and PI. It provides an accurate data foundation for subsequent warehousing management, circulation tracking, patient usage records and other links.
[0113] S403. Assemble the correspondence between each UDI code segment set and each target boxed consumable to obtain the UDI consumable identification result.
[0114] Specifically, the process iterates through all processed target boxed consumables. For each consumable, its own identification information (such as its location ID in the image or visual features) is bound to its complete UDI code segment set, forming a correspondence entry between consumable and complete UDI information. All correspondence entries for all target boxed consumables are then aggregated to construct a structured, global result list or data map. This complete result data volume clearly records each consumable identified in the current target scenario and its accurate UDI identity information, which is defined as the final UDI consumable identification result.
[0115] The correspondence refers to the one-to-one mapping between a target boxed consumable instance and its UDI code segment set, which clarifies which box corresponds to which set of UDI information.
[0116] For example, the system has generated UDI code segment sets for two consumables: the UDI code segment set for Box1 is {DI: DI123456, PI: PIBATCH0012025}, and the UDI code segment set for Box2 is {DI: DI789012, PI: PIBATCH0022026}. A list or JSON object will be created as the UDI consumable identification result.
[0117] The above embodiments ensure the absolute accuracy of the association between UDI data and physical entities, meeting the stringent requirements of industrial applications for data integrity and reliability, and marking the successful completion of the end-to-end conversion from raw images to actionable data.
[0118] Figure 3 This is a schematic diagram of the process for training the preset recognition model provided in the embodiments of this application.
[0119] Based on the above embodiments, as an optional embodiment, both the preset packaging box recognition model and the preset barcode recognition model are sub-models of the preset recognition model, see [link to relevant documentation]. Figure 3 ,against Figure 1 The preset identification model in the multi-target UDI consumable identification method shown can be used... Figure 3 The training is carried out in steps S501-S504, which are explained in detail below.
[0120] S501. Obtain the target training image and the benchmark recognition result of the target training image.
[0121] Specifically, in terms of system design, the preset packaging box recognition model and the preset barcode recognition model are not two completely independent models, but rather two components or specific functional output branches of a unified, larger preset recognition model. This preset recognition model is designed to learn the features of both packaging boxes and barcodes simultaneously during training. To train this unified model, a large number of target training images need to be collected. These images should cover various lighting conditions, different placement angles, multiple stacking densities, and different degrees of reflection. More importantly, each target training image must be equipped with an accurate baseline recognition result. This result includes the precise location bounding box coordinates of each packaging box and each barcode in the image, as well as their corresponding category labels (such as packaging box, DI code, PI code), providing the model with the standard answers necessary for learning.
[0122] The preset packaging box recognition model and the preset barcode recognition model are both sub-models of the preset recognition model. The preset recognition model is the parent model or overall architecture, while the preset packaging box recognition model and the preset barcode recognition model are two dedicated output heads or functional branches within it. They share the main feature extraction network of the model, but are separated in subsequent layers, respectively responsible for outputting the detection results of packaging boxes and barcodes. Target training images refer to a set of raw image data used to train and optimize the parameters of the preset recognition model. These images contain various real-world scenarios that the model needs to learn. Baseline recognition results, also often called labels or annotations, are standard data that are precisely pre-annotated manually and correspond one-to-one with each target training image. They clearly indicate the true bounding box positions and categories of all targets to be detected (packaging boxes, various types of barcodes) in the image, and are the absolute basis for the model to measure its prediction accuracy, calculate loss, and adjust parameters during the learning process.
[0123] S502. Identify the target training image through the initial recognition model to obtain the initial recognition result.
[0124] Specifically, a target training image is input into an initial recognition model that has not yet fully converged and whose parameters are yet to be optimized. Based on its current internal parameters (weights and biases), the model performs a series of complex nonlinear transformations such as convolution and pooling on the input image, performing feature extraction and inference calculations. Its output layer generates prediction information for the image, which includes the class probabilities of all targets (packaging boxes and barcodes) that the model believes exist in the image, as well as the coordinate data of the predicted bounding box corresponding to each target. This complete output, generated by the model's current capabilities and containing all predicted targets and their locations, is the initial recognition result.
[0125] The initial recognition model refers to a deep learning model whose network structure has been built but whose parameters have not yet been fully trained (or are in an intermediate training state). It represents the initial or intermediate form of the preset recognition model during the training process. The initial recognition result is the predicted output of the initial recognition model for the target training image, including the set of bounding boxes (predicted boxes) predicted by the model and the class confidence score for each box. This result usually differs significantly from the baseline recognition result (the true answer) in the early stages of training.
[0126] S503. Calculate the recognition loss value based on the initial recognition result and the bounding box regression loss function CIoU.
[0127] Specifically, the initial recognition result is paired with the baseline recognition result (i.e., the set of ground truth boxes) corresponding to the image. For each successfully paired predicted and ground truth boxes, the system calls the bounding box regression loss function CIoU for calculation, which combines three key geometric factors that measure the difference between two bounding boxes: Intersection over Union (IoU): This is the most basic metric, calculated by dividing the intersection area of the predicted box and the ground truth box by their union area, directly reflecting the degree of overlap between the two boxes. Center distance: CIoU additionally calculates the Euclidean distance between the center points of the two bounding boxes. This penalty term solves the problem that IoU alone cannot distinguish whether two non-overlapping boxes are far apart or about to touch, encouraging the center of the predicted box to move quickly towards the center of the ground truth box. Aspect Ratio Consistency: This is a unique advantage of CIoU compared to its predecessors (such as DIoU). It calculates the aspect ratio difference between two boxes and applies a penalty. This means that even if the center points of two boxes coincide and the overlap area is the same, if a predicted box is short and wide while the ground truth box is tall and narrow, CIoU will still consider it a prediction that needs improvement. During execution, the predicted bounding boxes and their corresponding ground truth bounding boxes from the initial recognition results are used as input. Then, according to the precise mathematical formula of the bounding box regression loss function CIoU, the differences in the above three dimensions are merged into a single, specific value. This final output value is the recognition loss value for this prediction. This value is not a simple right or wrong judgment, but a fine-grained quantitative score: a lower loss value means that the model's prediction is very close to the real target in terms of position, size, and shape; while a higher loss value precisely indicates that there is a large deviation in the prediction and implicitly provides a clear gradient direction for how the model should adjust its internal parameters to reduce this deviation.
[0128] The bounding box regression loss function CIoU is an advanced function used to measure the similarity between two bounding boxes. It considers not only the overlapping area but also the distance between the center points and the aspect ratio, making the regression more accurate and converging faster. The recognition loss value is a non-negative scalar value. The larger the value, the further the model's current prediction deviates from the reality; the smaller the value, the more accurate the prediction.
[0129] Through the above embodiments, the advanced loss function CIoU provides the model with richer and more accurate optimization guidance signals than simple IoU loss. It not only requires that the predicted box and the ground truth box have sufficient overlap area, but also forces the predicted box to move closer to the ground truth box in terms of center point position and aspect ratio. This enables the model to learn more accurately how to locate small-sized barcodes located in complex positions such as box edges and corners, effectively reducing the positioning drift problem caused by perspective distortion or film refraction.
[0130] Based on the above embodiments, as an optional embodiment, the benchmark identification result includes multiple benchmark bounding boxes, targeting... Figure 1 The step S503 shown can be implemented through steps S5031-S5033, which will be explained in detail below.
[0131] S5031. Obtain the reference coordinates of the target reference bounding box and the initial coordinates of the target initial bounding box corresponding to the target reference bounding box in the initial recognition result. The target reference bounding box can be any reference bounding box.
[0132] Specifically, a target baseline bounding box (i.e., a real target to be processed) is selected from the baseline recognition results. Then, a matching algorithm (such as based on category and location overlap) is used to find the prediction box that best corresponds to it in the initial recognition results, i.e., the initial target bounding box. Once found, the system extracts the baseline coordinates of the target baseline bounding box (i.e., the precise location data of the real box, such as [x]) from the baseline recognition results. gt_min y gt_min x gt_max y gt_max ]), and simultaneously extract the initial coordinates of the initial bounding box of the target that matches it from the initial recognition result (i.e., the position data predicted by the model corresponding to the real target, such as [x]). pred_min y pred_min x pred_max y pred_max This pair of reference coordinates and initial coordinates constitutes a data unit for calculating the loss of a single target.
[0133] The benchmark identification result includes multiple benchmark bounding boxes, indicating the data structure of the benchmark identification result. It is a set of bounding box annotations, each representing a real target to be detected. The target benchmark bounding box refers to one of the multiple bounding boxes arbitrarily selected from the benchmark identification result within the current processing cycle, used as the object for current evaluation and loss calculation. Benchmark coordinates refer to the coordinate data defining the position and extent of the target benchmark bounding box in the coordinate system of its associated target training image; these are the ground truth values used to measure prediction accuracy. The target initial bounding box corresponding to the target benchmark bounding box in the initial identification result refers to the predicted box determined from the numerous initial bounding boxes predicted by the model through a matching algorithm, representing the same object as the current target benchmark bounding box. Initial coordinates refer to the coordinate data defining the predicted position and extent of the target initial bounding box in the coordinate system of its associated target training image; these are the model's current output answer. The statement that the target benchmark bounding box can be any benchmark bounding box emphasizes that this operation is performed on a per-target basis for each real target in the benchmark identification result.
[0134] S5032. Based on the reference coordinates and the initial coordinates, calculate the intersection-union ratio of the training areas of the target reference bounding box and the target initial bounding box, as well as the Euclidean distance, and obtain the diagonal distance of the minimum closure region containing the target reference bounding box and the target initial bounding box.
[0135] Specifically, based on the coordinates of the two bounding boxes, their intersection and union regions are calculated. The intersection area / union area formula is used to calculate the training area intersection-union ratio, which measures the degree of overlap between the two boxes. Next, the coordinates of the center points of the two bounding boxes are calculated. Then, based on the coordinates of these two points, the Euclidean distance formula is used to calculate the Euclidean distance between them, which measures the deviation of the center point positions of the two boxes. Finally, a minimum rectangular region that can completely surround both the target reference bounding box and the target initial bounding box is determined, and the length of the diagonal of this rectangular region, i.e., the diagonal distance of the minimum closure region, is calculated. This value is used as a normalization factor to balance the distance measurement between targets of different sizes.
[0136] The training area intersection-over-union ratio (IoU) is a scalar value ranging from [0,1], representing the proportion of the overlapping area of two bounding boxes to their total area. In the training context, it specifically refers to the IoU calculated during the model training loop. The Euclidean distance is a scalar value representing the straight-line distance between the center point of the target baseline bounding box and the center point of the target initial bounding box. The minimum closure region containing both the target baseline bounding box and the target initial bounding box is the smallest bounding rectangle that can simultaneously enclose both the ground truth bounding box and the predicted bounding box. The diagonal distance is the diagonal length of the aforementioned minimum closure region.
[0137] S5033. Based on the preset weight parameters and preset size measurement parameters, determine the recognition loss value as follows: Recognition loss value = 1 - training area intersection-union ratio + (Euclidean distance / diagonal distance)2 + preset weight parameters × preset size measurement parameters.
[0138] Specifically, two preset hyperparameters are read from the system configuration: preset weight parameter (usually denoted as θ) and preset size measurement parameter (usually denoted as v). Then, the system strictly follows the mathematical formula of CIoU loss function to assemble and calculate: First, calculate 1 - training area intersection-union ratio, which measures the insufficiency of overlapping area; then calculate (Euclidean distance / diagonal distance)², which normalizes the center point distance so that it is independent of the target size; then calculate preset weight parameter × preset size measurement parameter, which introduces a penalty for aspect ratio consistency, which can be referred to formula (2). Finally, add these three results together, and the sum is the recognition loss value for the current target reference bounding box and the target initial bounding box.
[0139] L CIoU =1−IoU+(ρ / c) 2 +θ×v(2) Among them, L CIoU To identify the loss value, it is a non-negative scalar value, composed of the sum of the area overlap, the normalized squared distance between the center points, and the size ratio penalty term, comprehensively reflecting the overall difference between the predicted box and the ground truth box. IoU is the training intersection-union ratio. ρ is the Euclidean distance calculated in the previous steps, and c is the diagonal distance of the smallest closure region containing the predicted box and the ground truth box. θ is the preset weight parameter, a hyperparameter used to balance the weights of different components in the loss function, which controls the contribution of the preset size measurement parameter to the total loss. v is the preset size measurement parameter, a parameter specifically used to measure the aspect ratio consistency between the target baseline bounding box and the target initial bounding box, and its calculation formula can be found in formula (3).
[0140] v=(4 / π 2 )×[arctan(w gt / h gt )-arctan(w pred / h pred )] 2 (3) Among them, w gt h is the width of the target reference bounding box. gt The height of the target reference bounding box; w pred h is the width of the initial bounding box of the target; pred The height of the initial bounding box of the target.
[0141] Through the above embodiments, the system generates a loss signal containing richer optimization information than the traditional IoU loss through comprehensive calculation of the CIoU loss function. This loss value not only penalizes the insufficient overlapping area (1-IoU), but also penalizes the positioning deviation of the center point ((ρ / c)²), and further penalizes the mismatch of the aspect ratio (θ×v). This multi-objective optimization orientation forces the model to learn how to make the predicted box as close as possible to the real box in terms of position, size, and shape during the training process, thereby greatly improving the accuracy of bounding box regression. In particular, for small target barcodes located at the edges and corners in your solution, it can effectively reduce positioning drift and achieve pixel-level accurate positioning.
[0142] S504. Based on the recognition loss value, train the initial recognition model to obtain the preset recognition model.
[0143] Specifically, the model's performance is gradually improved by utilizing feedback from the loss signal. The gradient of the recognition loss value relative to the millions or even billions of trainable parameters (weights and biases) within the initial recognition model is calculated. These gradients indicate how each parameter should be fine-tuned (increased or decreased) to reduce the recognition loss value. Subsequently, the system uses an optimizer (such as Adam, SGD, etc.) to actually update all parameters of the initial recognition model based on the calculated gradients and a preset learning rate. This process (forward propagation to calculate loss → backpropagation to calculate gradients → optimizer to update parameters) iterates hundreds or thousands of times (epochs) on a dataset consisting of a large number of target training images. Finally, when the model's performance on the validation set stabilizes and reaches a preset standard, training stops. At this point, the fully optimized model with high-precision recognition capabilities is defined as the preset recognition model.
[0144] Through the above embodiments, by using the recognition loss value as a precise navigation signal, the model parameters are guided to search effectively in the huge solution space, and finally locate an optimal solution region that can accurately complete complex visual recognition tasks. The resulting preset recognition model is the core intelligent engine that enables the entire patented technology solution to be implemented. It encapsulates complex knowledge learned from massive training data on how to overcome reflection, distinguish scales, and accurately locate targets, ensuring high robustness and high accuracy when recognizing multi-target UDI consumables in practical applications.
[0145] The multi-target UDI consumable identification system in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of the structure of a multi-target UDI consumable identification system provided in an embodiment of this application.
[0146] It should be noted that, Figure 4The structure of the multi-target UDI consumable identification system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0147] like Figure 4 As shown, the multi-target UDI consumable identification system includes a central processing unit 601, which can perform various appropriate actions and processes based on a program stored in a read-only memory 602 or a program loaded from a storage section 608 into a random access memory 603, such as performing the methods described in the above embodiments. The random access memory 603 also stores various programs and data required for system operation. The central processing unit 601, read-only memory 602, and random access memory 603 are interconnected via a bus 604. An input / output interface 605 is also connected to the bus 604.
[0148] The following components are connected to the input / output interface 605: an input section 606 including audio input devices, push-button switches, etc.; an output section 607 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.
[0149] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit 601, it performs the various functions defined in the present invention. It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0151] Specifically, the multi-target UDI consumable identification system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the multi-target UDI consumable identification method provided in the above embodiment.
[0152] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the multi-target UDI consumable identification system described in the above embodiments; or it may exist independently and not assembled into the multi-target UDI consumable identification system. The storage medium carries one or more computer programs, which, when executed by a processor of the multi-target UDI consumable identification system, cause the multi-target UDI consumable identification system to implement the multi-target UDI consumable identification method provided in the above embodiments.
[0153] The above description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A multi-target UDI consumable identification method, characterized in that, include: Acquire a first polarization image of the target scene at a first polarization angle and a second polarization image at a second polarization angle, wherein the target scene includes multiple boxed consumables, and the first polarization angle and the second polarization angle are orthogonal; Based on the first polarization image and the second polarization image, a reflection-free fused image of the target scene is generated; Based on a preset packaging box recognition model, the target region of interest (ROI) corresponding to the target boxed consumable in the non-reflective fused image and the first position coordinates corresponding to the target ROI are identified. The target boxed consumable is any boxed consumable in the target scene. Based on a preset barcode recognition model, the barcode set in the target ROI is identified, and the second position coordinates of the target barcode in the target ROI are located, wherein the target barcode is any barcode in the barcode set; Based on the first position coordinates and the second position coordinates, the second position coordinates are transformed to obtain the third position coordinates of the target barcode in the non-reflective fused image; Based on the first location coordinates and the third location coordinates, the UDI consumable identification result of each barcode corresponding to each boxed consumable is determined by a preset constraint rule.
2. The method according to claim 1, characterized in that, The step of generating a reflection-free fused image of the target scene based on the first polarization image and the second polarization image includes: Obtain a first pixel intensity set of the first polarization image and a second pixel intensity set of the second polarization image. The first pixel intensity set includes multiple first pixel intensities corresponding to the first polarization image, and the second pixel intensity set includes multiple second pixel intensities corresponding to the second polarization image. For each pixel, the first pixel intensity of the target pixel is subtracted from the second pixel intensity of the target pixel to obtain the target pixel difference, where the target pixel is any of the aforementioned pixels; The sum of the first pixel intensity and the second pixel intensity of the target pixel is calculated to obtain the target pixel sum value; The linear polarization degree of the target pixel is obtained by calculating the ratio of the target pixel difference to the target pixel sum. Based on the linear polarization degree and the preset weight fusion function, the first polarization image and the second polarization image are fused to obtain the anti-reflection fused image.
3. The method according to claim 2, characterized in that, The step of fusing the first polarization image and the second polarization image based on the linear polarization degree and a preset weight fusion function to obtain the anti-reflective fused image includes: The difference between 1 and the linear polarization degree of the target pixel is used as the polarization compensation coefficient of the target pixel. Calculate the product of the first preset illumination compensation coefficient, the polarization compensation coefficient, and the first pixel intensity of the target pixel to obtain the first compensated pixel intensity of the target pixel after compensation; Calculate the product of the second preset illumination compensation coefficient and the second pixel intensity of the target pixel to obtain the second compensated pixel intensity of the target pixel; Based on the first compensated pixel intensity and the second compensated pixel intensity, the fused target pixel is obtained; Based on each of the fused target pixels, the anti-reflective fused image is obtained.
4. The method according to claim 1, characterized in that, Based on the first and second position coordinates, the second position coordinates are transformed to obtain the third position coordinates of the target barcode in the anti-reflective fused image, including: The sum of the minimum x-coordinate of the first position coordinate and the x-coordinate of the center point of the second position coordinate is calculated to obtain the x-coordinate of the target barcode in the non-reflective fused image. The sum of the minimum ordinate of the first position coordinate and the ordinate of the center point of the second position coordinate is calculated to obtain the ordinate of the target barcode in the non-reflective fused image. The horizontal and vertical coordinates of the barcode are used as the third position coordinates of the target barcode.
5. The method according to claim 4, characterized in that, The preset constraint rules include: First constraint rule: The horizontal coordinate of the barcode is greater than or equal to the minimum horizontal coordinate of the first position coordinate and less than or equal to the maximum horizontal coordinate of the first position coordinate; Second constraint rule: The barcode's ordinate is greater than or equal to the minimum ordinate of the first position coordinate and less than or equal to the maximum ordinate of the first position coordinate; Third constraint rule: The area intersection-union ratio of the target barcode bounding box and the packaging box bounding box of the ROI is greater than or equal to the first preset intersection-union ratio threshold and less than or equal to the second preset intersection-union ratio threshold; The step of determining the UDI consumable identification result of each barcode corresponding to each boxed consumable based on the first position coordinates and the third position coordinates, through preset constraint rules, includes: Based on the preset constraint rules, the target barcode that conforms to the first constraint rule, the second constraint rule and the third constraint rule is taken as the target code segment, and the target boxed consumable corresponding to the target barcode includes at least one of the target code segments; For each target boxed consumable, all target code segments of the target boxed consumable are aggregated to generate a UDI code segment set corresponding to the target boxed consumable; Assemble the correspondence between each set of UDI code segments and each of the target boxed consumables to obtain the UDI consumable identification result.
6. The method according to claim 1, characterized in that, The preset packaging box recognition model and the preset barcode recognition model are both sub-models of the preset recognition model, and the method further includes: Obtain the target training image and the baseline recognition result of the target training image; The target training image is identified using an initial recognition model to obtain an initial recognition result; Based on the initial recognition results and the bounding box regression loss function CIoU, the recognition loss value is calculated; Based on the recognition loss value, the initial recognition model is trained to obtain the preset recognition model.
7. The method according to claim 6, characterized in that, The baseline recognition result includes multiple baseline bounding boxes. The calculation of the recognition loss value based on the initial recognition result and the bounding box regression loss function CIoU includes: Obtain the reference coordinates of the target reference bounding box and the initial coordinates of the target initial bounding box corresponding to the target reference bounding box in the initial recognition result, wherein the target reference bounding box is any of the reference bounding boxes; Based on the reference coordinates and the initial coordinates, calculate the training area intersection-union ratio and Euclidean distance of the target reference bounding box and the target initial bounding box, and obtain the diagonal distance of the minimum closure region containing the target reference bounding box and the target initial bounding box; Based on preset weight parameters and preset size measurement parameters, the recognition loss value is determined in the following manner: The recognition loss value = 1 - the training area intersection-union ratio + (the Euclidean distance / the diagonal distance) 2 +The preset weight parameter ×The preset size measurement parameter.
8. A multi-target UDI consumable identification system, characterized in that, The multi-target UDI consumable identification system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the multi-target UDI consumable identification system to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the multi-target UDI consumable identification system, the multi-target UDI consumable identification system performs the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product is run on the multi-target UDI consumable identification system, the multi-target UDI consumable identification system performs the method as described in any one of claims 1-7.