A method and device for automatic label modeling, a storage medium, and an electronic device
Automated tag modeling using CNNs addresses inefficiencies and inconsistencies in manual tag modeling by filtering and refining edge frames on packaging images to create consistent tag templates for accurate placement.
Patent Information
- Application Number
- CN202210248818.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-14
AI Technical Summary
In the prior art, label modeling efficiency is low and the consistency of modeling cannot be guaranteed, making it difficult to control the quality of label pasting.
The target border map is automatically filtered by convolutional neural network and prior conditions, and a label template is created. The suspected border map is extracted from the packaging picture, and the candidate border map is obtained by filtering, and the target border map is filtered according to the prior conditions, which is used to position the paste position of the label information.
The automatic creation of label templates is realized, ensuring consistency and accuracy of modeling, improving modeling efficiency, and automatically labeling labels and product models to ensure the correctness of label pasting.
Smart Images

Figure CN114723936B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic modeling, and more particularly, to a method and device for automatic label modeling, a storage medium, and an electronic device. Background Art
[0002] In the related art, in order to control the label pasting quality on the packaging box of air conditioner products and avoid pasting the wrong label, a visual inspection device is required to determine whether the currently pasted label is correct according to the established label template diagram.
[0003] In the related art, the modeling form is to manually draw a frame for the label for modeling. However, since there are many product models and new models are generated every day, this brings a great workload to the modeling personnel. At the same time, manual modeling has subjective consciousness and cannot ensure the consistency of modeling.
[0004] In view of the above problems in the related art, no effective solution has been found yet. Summary of the Invention
[0005] In order to solve the problems of low efficiency of manual label modeling and inability to ensure modeling consistency in the prior art, the present invention provides a method and device for automatic label modeling, a storage medium, and an electronic device.
[0006] According to one aspect of the embodiments of the present application, a method for automatic label modeling is provided, including: obtaining a set of suspected border diagrams from a packaging picture of a target box; filtering the set of suspected border diagrams by using a convolutional neural network to obtain a set of candidate border diagrams; filtering the set of candidate border diagrams according to prior conditions to obtain a plurality of target border diagrams, wherein the target border diagrams include label information of the target box; creating a label template of the target box by using the plurality of target border diagrams and the packaging picture, wherein the label template is used to locate the pasting position of the label information on the target box.
[0007] Further, obtaining a set of suspected border diagrams from a packaging picture of a target box includes: obtaining a packaging picture of the outer surface of the target box; performing smoothing filtering and noise reduction processing on the packaging picture to obtain a first intermediate picture; performing edge detection on the first intermediate picture to extract the set of suspected border diagrams.
[0008] Further, perform edge detection on the first intermediate picture, and extract the suspected border picture set, including: detecting the edge straight line features of the first intermediate picture; finding multiple short straight lines in the first intermediate picture according to the edge straight line features, where the short straight line is a straight line with a length less than a preset length; for each first short straight line among the multiple short straight lines, find a second straight line segment in other short straight lines that is parallel to its trajectory and has the same length, and store the pairwise combination of the first short straight line and the second straight line segment into a parallel line group;
[0009] Perform parallelogram combination on the straight lines in the parallel line group to obtain a suspected border picture set.
[0010] Further, use a convolutional neural network to filter the suspected border picture set to obtain a candidate border picture set, including: for each suspected border picture in the candidate border pictures, convert the suspected border picture into a picture vector; use the picture vector to loop through the following steps until the last pooling layer: use the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer, and output a first feature map; use the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer, and output a second feature map with reduced dimensions; determine the second feature map as the input quantity for the next convolutional layer; input the target feature map output by the last pooling layer into a classification model, and output the classification information of the suspected border picture, where the classification information is used to indicate whether the corresponding suspected border picture is a candidate border picture.
[0011] Further, using the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer and output a first feature map includes: sequentially inputting the picture vector into M sliding windows of the first convolutional layer, and outputting M first feature maps, where the first convolutional layer includes M sliding windows with different weights, and the picture vector performs the following steps in each sliding window: multiply the weight value of the current sliding window by the gray value of the suspected border picture and sum to obtain a feature value, and M is an integer greater than 1.
[0012] Further, using the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer and output a second feature map with reduced dimensions includes: inputting the M first feature maps into M sliding windows of the first pooling layer, and outputting M second feature maps, where the first pooling layer includes M sliding windows with the same weights, and the M first feature maps perform the following steps in each sliding window: obtain the first feature value of the first feature map in the current sliding window, compare the first feature value with the second feature value of the second feature map in the previous sliding window, and determine the feature map with the largest feature value as the second feature map with reduced dimensions, and M is an integer greater than 1.
[0013] Further, filtering the candidate border map set according to prior conditions to obtain a plurality of target border maps, including: for each candidate border map in the candidate border map set, calculating the label parameters of the candidate border map; determining whether the label parameters meet the prior conditions, where the prior conditions include at least one of the following: the minimum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the minimum width-to-height ratio of the label, having characters within the label, and the hue difference ratio between the overall hue of the label and the hue of the box body; if the prior conditions are met, determining that the candidate border map is a target border map.
[0014] Further, calculating the label parameters of the candidate border map includes: calculating the spacing value between two parallel lines in the candidate border map, determining the horizontal spacing as the width value, determining the vertical spacing as the height value, and determining the ratio between the horizontal spacing and the vertical spacing as the width-to-height ratio, where the label parameters include: width value, height value, width-to-height ratio; determining whether the candidate border map contains characters through optical character recognition (OCR); calculating the first grayscale value of the candidate border map and calculating the second grayscale value of the packaging picture, and determining the ratio between the first grayscale value and the second grayscale value as the hue difference ratio.
[0015] According to another aspect of the embodiments of the present application, there is also provided a label automatic modeling device, including: an acquisition module, configured to acquire a set of suspected border maps from the packaging picture of the target box body; a first filtering module, configured to filter the set of suspected border maps by using a convolutional neural network to obtain a set of candidate border maps; a second filtering module, configured to filter the set of candidate border maps according to prior conditions to obtain a plurality of target border maps, where the target border maps include the label information of the target box body; a creation module, configured to create a label template of the target box body by using the plurality of target border maps and the packaging picture, where the label template is used to locate the pasting position of the label information on the target box body.
[0016] Further, the acquisition module includes: an acquisition unit, configured to acquire the packaging picture of the outer surface of the target box body; a processing unit, configured to perform smoothing filtering and noise reduction processing on the packaging picture to obtain a first intermediate picture; an extraction unit, configured to perform edge detection on the first intermediate picture to extract the set of suspected border maps.
[0017] Further, the extraction unit includes: a detection subunit, configured to detect the edge straight-line features of the first intermediate picture; a search subunit, configured to search for multiple short straight lines in the first intermediate picture according to the edge straight-line features, where the short straight line is a straight line with a length less than a preset length; a storage subunit, configured to, for each first short straight line among the multiple short straight lines, search for a second straight-line segment parallel to its trajectory and having the same length among other short straight lines, and store the pairwise combination of the first short straight line and the second straight-line segment into a parallel-line group; and a combination subunit, configured to perform a parallelogram combination on the straight lines in the parallel-line group to obtain a set of suspected border pictures.
[0018] Further, the first filtering module includes: a conversion unit, configured to convert each suspected border picture in the candidate border pictures into a picture vector; an execution unit, configured to repeatedly execute the following steps using the picture vector until the last pooling layer: use the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer, and output a first feature map; use the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer, and output a second feature map with reduced dimensions; determine the second feature map as the input quantity for the next convolutional layer; and an output unit, configured to input the target feature map output by the last pooling layer into a classification model and output classification information of the suspected border picture, where the classification information is used to indicate whether the corresponding suspected border picture is a candidate border picture.
[0019] Further, the execution unit includes: a first output subunit, configured to sequentially input the picture vector into M sliding windows of a first convolutional layer and output M first feature maps, where the first convolutional layer includes M sliding windows with different weights, and the picture vector performs the following steps in each sliding window: multiply the weight value of the current sliding window by the gray value of the suspected border picture and sum to obtain a feature value, and M is an integer greater than 1.
[0020] Further, the execution unit further includes: a second output subunit, configured to input the M first feature maps into M sliding windows of a first pooling layer and output M second feature maps, where the first pooling layer includes M sliding windows with the same weights, and the M first feature maps perform the following steps in each sliding window: obtain a first feature value of the first feature map in the current sliding window, compare the first feature value with a second feature value of the second feature map in the previous sliding window, and determine the feature map with the largest feature value as the second feature map with reduced dimensions, and M is an integer greater than 1.
[0021] Further, the second filtering module includes: a calculation unit configured to calculate label parameters of each candidate bounding box map in the candidate bounding box map set; a determination unit configured to determine whether the label parameters meet a prior condition, where the prior condition includes at least one of the following: the minimum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the maximum aspect ratio of the label, the minimum aspect ratio of the label, the presence of characters within the label, and the hue difference ratio between the overall hue of the label and the hue of the box body; a determination unit configured to, if the prior condition is met, determine that the candidate bounding box map is a target bounding box map.
[0022] Further, the calculation unit includes: a first calculation subunit configured to calculate the distance value between two parallel lines in the candidate bounding box map, determine the horizontal distance as the width value, determine the vertical distance as the height value, and determine the ratio between the horizontal distance and the vertical distance as the aspect ratio, where the label parameters include: the width value, the height value, and the aspect ratio; a judgment subunit configured to determine whether the candidate bounding box map contains characters through optical character recognition (OCR); a second calculation subunit configured to calculate a first grayscale value of the candidate bounding box map and calculate a second grayscale value of the packaging picture, and determine the ratio between the first grayscale value and the second grayscale value as the hue difference ratio.
[0023] According to another aspect of the embodiments of the present application, there is also provided a storage medium including a stored program, where the program, when running, executes the above method steps.
[0024] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory communicate with each other through the communication bus; where: the memory is configured to store a computer program; the processor is configured to execute the above method steps by running the program stored on the memory.
[0025] The embodiments of the present application also provide a computer program product including instructions that, when running on a computer, cause the computer to execute the steps in the above method.
[0026] Through the present invention, a set of suspected border images is obtained from the packaging image of the target box; a convolutional neural network is used to filter the set of suspected border images to obtain a set of candidate border images; the set of candidate border images is filtered according to prior conditions to obtain a plurality of target border images, wherein the target border images include the label information of the target box; the plurality of target border images and the packaging image are used to create a label template for the target box, wherein the label template is used to locate the pasting position of the label information on the target box. By using a convolutional neural network and prior conditions to screen target border images from the set of suspected border images of the packaging image and using the target border images to create the label template for the target box, the consistency and accuracy of the label template are ensured. At the same time, an automatic creation scheme for the label template is realized, solving the technical problems in the prior art that the efficiency of manually modeling the label is low and the modeling consistency cannot be guaranteed, increasing the modeling efficiency, ensuring the modeling consistency, and being able to automatically label the label, product model, and barcode. Description of the Drawings
[0027] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0028] Figure 1 is a hardware structure block diagram of a label automatic modeling device according to an embodiment of the present invention;
[0029] Figure 2 is a flowchart of a method for label automatic modeling according to an embodiment of the present invention;
[0030] Figure 3 is a schematic diagram of a scenario according to an embodiment of the present invention;
[0031] Figure 4 is a flowchart of the operation according to an embodiment of the present invention;
[0032] Figure 5 is a structure block diagram of a label automatic modeling device according to an embodiment of the present invention. Detailed Embodiments
[0033] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the scope of protection of this application. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0034] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Embodiment 1
[0036] The method embodiment provided in Embodiment 1 of this application can be executed on a computer, a server, a control device or a similar computing device. Taking running on a computer as an example, Figure 1 is a hardware structure block diagram of automatic label modeling according to an embodiment of the present invention. As Figure 1 shown, the automatic label modeling may include one or more ( Figure 1 only one is shown in Figure 1 ) processor 102 (processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned identification device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned automatic label modeling. For example, the automatic label modeling may further include more or fewer components than Figure 1 shown, or have a different configuration from
[0037] The memory 104 can be used to store programs for automatically modeling operation labels. For example, software programs and modules of application software, such as the automatic modeling program corresponding to a method for automatically modeling labels in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the automatic modeling program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the label automatic modeling device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the label automatic modeling device. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0039] In this embodiment, a method for automatically modeling labels is provided. Figure 2 It is a flowchart of a method for automatically modeling labels according to an embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:
[0040] Step S202, obtaining a set of suspected border graphs from the packaging picture of the target box;
[0041] In an implementation manner of this embodiment, a plurality of border graph patterns are attached to the outer packaging surface of the target box, such as bar code graphs, energy efficiency label graphs, product logos, caution icons, etc. All these border graphs are suspected border graphs. All suspected border graphs are obtained from the packaging picture of the target box and a set is formed. Optionally, the border graphs in this embodiment can be square, rectangular, rectangular, hexagonal, etc. In some special scenarios, they can also be trapezoidal, rhombic, triangular, pentagonal, etc.
[0042] Step S204, filtering the set of suspected border graphs by using a convolutional neural network to obtain a set of candidate border graphs;
[0043] In one implementation of this embodiment, a convolutional neural network binary classification model is used to filter the set of suspected border images. The binary classification model is first trained with training data, which consists of two types of data. One type is the local background images of cartons, and the other type is the images of labels. Finally, a model for distinguishing these two types of data is obtained, and the set of suspected border images is filtered through the model to obtain a set of candidate border images. The candidate border images are the partial label border images that meet the conditions after the set of suspected border images is classified and filtered by the binary classification model. The target border images can be obtained through further filtering.
[0044] Step S206: Filter the set of candidate border images according to prior conditions to obtain a number of target border images. Among them, the target border images include the label information of the target box body.
[0045] In one implementation of this embodiment, the prior conditions are used to judge the target border images that meet the conditions in the set of candidate border images. The prior conditions include parameters such as the size and ratio of the standard label box. The target border images include the label information of the target box body.
[0046] Step S208: Create a label template for the target box body using the number of target border images and the packaging image. The label template is used to locate the pasting position of the label information on the target box body.
[0047] In one implementation of this embodiment, the label template includes the coordinate position information of the target border image on the target box body. Saving the coordinate information and the corresponding packaging image to the database completes the automatic label modeling.
[0048] After obtaining the border position information, the target border images are accompanied by barcode images. By performing barcode recognition, the product barcode can be obtained. At the same time, because the product model is associated with the barcode, as long as the barcode content is scanned, the product model can be queried online. The above steps solve the problem of how to automatically label the product model, barcode, and label.
[0049] At the same time, because the target border images are accompanied by information data such as product models, barcodes, and label coordinate positions, this embodiment can also determine whether the label border is pasted correctly. Since the character content and size on the label have been provided to the packaging production line in advance, after pasting, it can be judged whether the label content is blurred or there are unclear areas such as signs and barcodes through the successfully modeled label border image template.
[0050] Through the above steps, a set of suspected border images is obtained from the packaging image of the target box body; a convolutional neural network is used to filter the set of suspected border images to obtain a set of candidate border images; the set of candidate border images is filtered according to prior conditions to obtain a plurality of target border images, wherein the target border images include the label information of the target box body; the plurality of target border images and the packaging image are used to create a label template for the target box body, wherein the label template is used to locate the pasting position of the label information on the target box body. By using a convolutional neural network and prior conditions to screen target border images from the set of suspected border images of the packaging image and using the target border images to create the label template for the target box body, the consistency and accuracy of the label template are ensured. At the same time, an automatic creation scheme for the label template is realized, solving the technical problems in the prior art that the efficiency of manually modeling the label is low and the modeling consistency cannot be guaranteed, increasing the modeling efficiency, ensuring the modeling consistency, and being able to automatically label the label, product model and barcode.
[0051] In this embodiment, obtaining a set of suspected border images from the packaging image of the target box body includes: obtaining a packaging image of the outer surface of the target box body; performing smoothing filtering and noise reduction processing on the packaging image to obtain a first intermediate image; performing edge detection on the first intermediate image to extract the set of suspected border images.
[0052] In the above steps, after obtaining the packaging image of the outer surface of the target box body through a camera, it is transmitted to a computer for processing. The packaging image is subjected to smoothing filtering and noise reduction processing, such as mean filtering and median filtering for noise reduction processing, to obtain a first intermediate image, and edge detection is performed to obtain each suspected border image on the packaging image and gather them together.
[0053] In this embodiment, performing edge detection on the first intermediate image to extract the set of suspected border images includes: detecting the edge straight line features of the first intermediate image; finding multiple short straight lines in the first intermediate image according to the edge straight line features, wherein the short straight line is a straight line with a length less than a preset length; for each first short straight line among the multiple short straight lines, finding a second straight line segment in other short straight lines that is parallel to its trajectory and has the same length, and storing the pairwise combination of the first short straight line and the second straight line segment into a parallel line group; performing parallelogram combination on the straight lines in the parallel line group to obtain the set of suspected border images.
[0054] In the above steps, each short straight line is found according to the straight line feature of the picture edge, and then the contour of each short straight line is calculated to find out the straight line segments parallel to each other in pairs as a parallel line group. Among them, the lengths of the straight line segments parallel to each other in pairs are the same, and the trajectories are parallel. Then, all the straight lines in the parallel line group are combined into parallelograms, and any number of adjacent and closable parallel line groups are selected and combined into a suspected border picture. Finally, a lot of suspected border pictures are obtained.
[0055] In this embodiment, a convolutional neural network is used to filter the set of suspected border pictures, and the set of candidate border pictures obtained includes: for each suspected border picture in the candidate border pictures, the suspected border picture is converted into a picture vector; the following steps are cyclically executed using the picture vector until the last pooling layer: the picture vector is used as an input quantity to perform feature extraction on the picture vector through a first convolutional layer, and a first feature map is output; the first feature map is used as an input quantity to perform downsampling on the first feature map through a first pooling layer, and a second feature map after dimensionality reduction is output; the second feature map is determined as the input quantity of the next convolutional layer; the target feature map output by the last pooling layer is input into a classification model, and the classification information of the suspected border picture is output, where the classification information is used to indicate whether the corresponding suspected border picture is a candidate border picture.
[0056] In the above steps, a convolutional neural network binary classification model is used to filter the suspected border pictures to obtain a set of candidate border pictures. The binary classification model has a total of 7 layers. Among them, the first layer is a 25*7*7 convolutional layer, the second layer is a 25*7*7 pooling layer, the third layer is a 75*9*9 convolutional layer, the fourth layer is a 75*9*9 pooling layer, the fifth layer is a 25*7*7 convolutional layer, the sixth layer is a 25*7*7 pooling layer, and the last layer is a classification layer.
[0057] During the training process of the convolutional neural network binary classification model, two types of data are used to train the model, and finally a model for distinguishing these two types of data is obtained. These two types of data are: several labeled pictures and local pictures of cartons. When the model is trained, the picture to be detected is input into the model, and the model will determine whether the picture is a labeled picture or a local picture of a carton. The internal processing of the model is that after the picture is input, it undergoes convolution and pooling to obtain a feature map with reduced dimensions, and then it is sent to the classification layer, such as a softmax classifier for classification, to determine which category the input picture belongs to, that is, whether it belongs to a labeled picture or a local picture of a carton.
[0058] In this embodiment, the role of the convolutional layer is to perform feature extraction on the input variable to obtain a feature map. The role of the pooling layer is to perform downsampling on the feature map obtained by the previous convolutional layer to obtain a feature map with reduced dimensions. The role of the classification layer is to classify the input feature map with reduced dimensions, such as using softmax for logistic regression.
[0059] In this embodiment, the picture vector is used as an input quantity, and the first convolutional layer is used to extract features from the picture vector. The output first feature map includes: the picture vector is sequentially input into M sliding windows of the first convolutional layer, and M first feature maps are output. Among them, the first convolutional layer includes M sliding windows with different weights. The picture vector performs the following steps in each sliding window: multiplying the weight value of the current sliding window by the gray value of the suspected border map and summing to obtain a feature value. M is an integer greater than 1.
[0060] In one example, M = 25, such as a 25*7*7 convolutional layer. 7*7 means sliding a 7*7 window with weights on the input picture, multiplying each weight value in the window by the corresponding picture gray value, and then summing them as the output. After sliding the window over the entire input picture, a feature map is obtained. 25 means 25 windows with different weights, and finally 25 different feature maps are obtained.
[0061] In this embodiment, the first feature map is used as an input quantity, and the first pooling layer is used to perform downsampling on the first feature map. The output second feature map after dimensionality reduction includes: the M first feature maps are input into M sliding windows of the first pooling layer, and M second feature maps are output. Among them, the first pooling layer includes M sliding windows with the same weights. The M first feature maps perform the following steps in each sliding window: obtaining the first feature value of the first feature map in the current sliding window, comparing the first feature value with the second feature value of the second feature map in the previous sliding window, and determining the feature map with the largest feature value as the second feature map after dimensionality reduction. M is an integer greater than 1.
[0062] In the above steps, for a 25*7*7 pooling layer, 7*7 means sliding a 7*7 window in the input feature map, taking the largest gray value in the window as the output, and obtaining a downsampled feature map. 25 means performing downsampling on 25 feature maps in the upper layer.
[0063] In this embodiment, filtering the candidate border map set according to prior conditions to obtain a number of target border maps includes: for each candidate border map in the candidate border map set, calculating the label parameter of the candidate border map; determining whether the label parameter meets the prior conditions, where the prior conditions include at least one of the following: the minimum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the maximum aspect ratio of the label, the minimum aspect ratio of the label, having characters within the label, the hue difference ratio between the overall hue of the label and the hue of the box; if the prior conditions are met, then determining the candidate border map as a target border map.
[0064] In the above steps, for the candidate border image, the label parameters are calculated to determine whether the candidate border image is a label border image, that is, it is judged by comparing the label parameters through prior conditions. When at least one of the prior conditions is met, it can be judged as a label border, that is, the target border image.
[0065] In an implementation manner of this embodiment, multiple sets of suspected border images are obtained from the packaging image of the target box body, including various logo images, barcode images, product precaution images, product logos, etc. After filtering through the convolutional neural network binary classification model, candidate border images are obtained, but there are still some non-label candidate border images that do not meet the requirements. The label border images that meet the requirements are further screened by judging the prior conditions. Different prior conditions can be set for different product packages. Although the labels are diverse, the label types of each product will be determined in advance before the packaging line production and can be directly applied to actual production.
[0066] In this embodiment, calculating the label parameters of the candidate border image includes: calculating the distance value between two parallel lines in the candidate border image, determining the horizontal distance as the width value, determining the vertical distance as the height value, and determining the ratio between the horizontal distance and the vertical distance as the width-to-height ratio. Among them, the label parameters include: width value, height value, width-to-height ratio; judging whether the candidate border image contains characters through Optical Character Recognition (OCR); calculating the first gray value of the candidate border image, and calculating the second gray value of the packaging image, and determining the ratio between the first gray value and the second gray value as the hue difference ratio.
[0067] In the above steps, the width and height of the label can be obtained by calculating the distance between two parallel lines of the candidate border image. Compare this with the minimum width and height, maximum width and height, maximum width-to-height ratio, and minimum width-to-height ratio of the label under the prior conditions, and then use OCR to identify whether there are characters in the candidate border image. At this time, the target box body has a gray value, and there is also a gray value in the candidate border image. By comparing the two gray values, the one with a large difference, that is, a large hue difference ratio, will delete the candidate border image if the compared gray difference is large.
[0068] Figure 3 This is a schematic diagram of the scenario of the embodiment of the present invention. The suspected border image of the box body is obtained through a camera, including the label border image, which is uploaded to a computer for processing and finally displayed on the monitor interface, and automatic modeling is completed.
[0069] Figure 4It is the flowchart of the embodiment of the present invention. After taking pictures to obtain images, suspected quadrilateral frames are obtained. Then, through a deep learning model, namely a convolutional neural network binary classification model, non-labeled quadrilateral frames are filtered out. Next, according to the prior conditions of the label, non-conforming quadrilateral frames are filtered out. Finally, the quadrilateral frame diagram of the label is obtained.
[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus necessary general mechanical equipment. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, can be embodied in the form of software controlling mechanical equipment. The software is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a mechanical equipment (such as a label automatic modeling device) to execute the methods described in various embodiments of the present invention.
[0071] Embodiment 2
[0072] In this embodiment, a label automatic modeling device is also provided to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0073] Figure 5 It is the structural block diagram of a label automatic modeling device according to an embodiment of the present invention. As Figure 5 shown, the device includes: an acquisition module 50, a first filtering module 52, a second filtering module 54, and a creation module 56. Among them,
[0074] The acquisition module 50 is used to obtain a set of suspected border diagrams from the packaging pictures of the target box body;
[0075] The first filtering module 52 is used to filter the set of suspected border diagrams by using a convolutional neural network to obtain a set of candidate border diagrams;
[0076] The second filtering module 54 is used to filter the set of candidate border diagrams according to prior conditions to obtain several target border diagrams, where the target border diagrams include the label information of the target box body;
[0077] The creation module 56 is used to create a label template of the target box body by using the several target border diagrams and the packaging pictures, where the label template is used to locate the pasting position of the label information on the target box body.
[0078] Optionally, the obtaining module includes: an obtaining unit configured to obtain a packaging picture of the outer surface of the target box body; a processing unit configured to perform smoothing filtering and noise reduction processing on the packaging picture to obtain a first intermediate picture; and an extraction unit configured to perform edge detection on the first intermediate picture and extract a set of suspected border pictures.
[0079] Optionally, the extraction unit includes: a detection subunit configured to detect edge straight line features of the first intermediate picture; a search subunit configured to search for multiple short straight lines in the first intermediate picture according to the edge straight line features, where the short straight line is a straight line with a length less than a preset length; a storage subunit configured to, for each first short straight line among the multiple short straight lines, search for a second straight line segment in other short straight lines that is parallel to its trajectory and has the same length, and store the combination of the first short straight line and the second straight line segment into a parallel line group in pairs; and a combination subunit configured to perform parallelogram combination on the straight lines in the parallel line group to obtain a set of suspected border pictures.
[0080] Optionally, the first filtering module includes: a conversion unit configured to convert each suspected border picture in the candidate border pictures into a picture vector; an execution unit configured to repeatedly execute the following steps using the picture vector until the last pooling layer: use the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer and output a first feature map; use the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer and output a second feature map with reduced dimensions; determine the second feature map as the input quantity of the next convolutional layer; and an output unit configured to input the target feature map output by the last pooling layer into a classification model and output classification information of the suspected border picture, where the classification information is used to indicate whether the corresponding suspected border picture is a candidate border picture.
[0081] Optionally, the execution unit includes: a first output subunit configured to sequentially input the picture vector into M sliding windows of the first convolutional layer and output M first feature maps, where the first convolutional layer includes M sliding windows with different weights, and the picture vector performs the following steps in each sliding window: multiply the weight value of the current sliding window by the gray value of the suspected border picture and sum to obtain a feature value, and M is an integer greater than 1.
[0082] Optionally, the execution unit further includes: a second output subunit, configured to input the M first feature maps into M sliding windows of a first pooling layer, and output M second feature maps, where the first pooling layer includes M sliding windows with the same weight, and the M first feature maps perform the following steps in each sliding window: obtaining a first feature value of the first feature map in the current sliding window, comparing the first feature value with a second feature value of the second feature map in the previous sliding window, and determining the feature map with the largest feature value as the second feature map after dimensionality reduction, where M is an integer greater than 1.
[0083] Optionally, the second filtering module includes: a calculation unit, configured to calculate a label parameter of each candidate bounding box map in the candidate bounding box map set; a judgment unit, configured to judge whether the label parameter meets a prior condition, where the prior condition includes at least one of the following: the minimum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the maximum aspect ratio of the label, the minimum aspect ratio of the label, having characters in the label, and the hue difference ratio between the overall hue of the label and the hue of the box body; a determination unit, configured to, if the prior condition is met, judge that the candidate bounding box map is a target bounding box map.
[0084] Optionally, the calculation unit includes: a first calculation subunit, configured to calculate the spacing value between two parallel lines in the candidate bounding box map, determine the horizontal spacing as the width value, determine the vertical spacing as the height value, and determine the ratio between the horizontal spacing and the vertical spacing as the aspect ratio, where the label parameter includes: the width value, the height value, and the aspect ratio; a judgment subunit, configured to judge whether the candidate bounding box map contains characters through optical character recognition (OCR); a second calculation subunit, configured to calculate a first gray value of the candidate bounding box map and calculate a second gray value of the packaging picture, and determine the ratio between the first gray value and the second gray value as the hue difference ratio.
[0085] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0086] Embodiment 3
[0087] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0088] Optionally, in this embodiment, the above storage medium can be configured to store a computer program for executing the following steps:
[0089] S1. Obtain a set of suspected border images from the packaging image of the target box;
[0090] S2. Use a convolutional neural network to filter the set of suspected border images to obtain a set of candidate border images;
[0091] S3. Filter the set of candidate border images according to prior conditions to obtain a number of target border images, where the target border images include the label information of the target box;
[0092] S4. Create a label template for the target box using the number of target border images and the packaging image, where the label template is used to locate the pasting position of the label information on the target box.
[0093] Optionally, in this embodiment, the above storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, etc., various media that can store computer programs.
[0094] An embodiment of the present invention also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0095] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0096] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0097] S1. Obtain a set of suspected border images from the packaging image of the target box;
[0098] S2. Use a convolutional neural network to filter the set of suspected border images to obtain a set of candidate border images;
[0099] S3. Filter the set of candidate border images according to prior conditions to obtain a number of target border images, where the target border images include the label information of the target box;
[0100] S4. Create a label template for the target box using the number of target border images and the packaging image, where the label template is used to locate the pasting position of the label information on the target box.
[0101] Optionally, for the specific examples in this embodiment, reference may be made to the examples described in the above embodiments and alternative embodiments, and details thereof will not be repeated here.
[0102] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0103] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0104] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0105] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0107] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0108] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for automatic label modeling, characterized in that, Including: Obtaining a set of suspected border images from the packaging image of the target box body; Filtering the set of suspected border images by using a convolutional neural network to obtain a set of candidate border images; Filtering the set of candidate border images according to prior conditions to obtain a plurality of target border images, wherein the target border images include label information of the target box body, and the prior conditions are used to judge the target border images that meet the conditions in the set of candidate border images; Creating a label template of the target box body by using the plurality of target border images and the packaging image, wherein the label template is used to locate the pasting position of the label information on the target box body.
2. The method according to claim 1, wherein Obtaining a set of suspected border images from the packaging image of the target box body includes: Obtaining a packaging image of the outer surface of the target box body; Performing smoothing filtering and noise reduction processing on the packaging image to obtain a first intermediate image; Performing edge detection on the first intermediate image to extract a set of suspected border images.
3. The method according to claim 2, wherein Performing edge detection on the first intermediate image to extract a set of suspected border images includes: Detecting the edge straight line features of the first intermediate image; Finding multiple short straight lines in the first intermediate image according to the edge straight line features, wherein the short straight lines are straight lines with lengths less than a preset length; For each first short straight line in the multiple short straight lines, finding a second straight line segment in other short straight lines that is parallel to its trajectory and has the same length, and storing the pairwise combination of the first short straight line and the second straight line segment into a parallel line group; Performing parallelogram combination on the straight lines in the parallel line group to obtain a set of suspected border images.
4. The method according to claim 1, characterized in that, Filtering the set of suspected border images by using a convolutional neural network to obtain a set of candidate border images includes: For each suspected border image in the candidate border images, converting the suspected border image into a picture vector; Using the picture vector to repeatedly execute the following steps until the last pooling layer: using the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer, and outputting a first feature map; using the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer, and outputting a second feature map after dimensionality reduction; determining the second feature map as the input quantity of the next convolutional layer; Inputting the target feature map output by the last pooling layer into a classification model, and outputting classification information of the suspected border image, wherein the classification information is used to indicate whether the corresponding suspected border image is a candidate border image.
5. The method according to claim 4, characterized in that, Using the picture vector as an input quantity to perform feature extraction on the picture vector through a first convolutional layer and outputting a first feature map includes: Sequentially inputting the picture vector into M sliding windows of the first convolutional layer, and outputting M first feature maps, wherein the first convolutional layer includes M sliding windows with different weights, and the picture vector performs the following steps in each sliding window: multiplying the weight value of the current sliding window by the gray value of the suspected border image and summing to obtain a feature value, and M is an integer greater than 1.
6. The method according to claim 5, wherein Using the first feature map as an input quantity to perform downsampling on the first feature map through a first pooling layer and outputting a second feature map after dimensionality reduction includes: Input the M first feature maps into M sliding windows of the first pooling layer to output M second feature maps. Herein, the first pooling layer includes M sliding windows with the same weights. The M first feature maps perform the following steps within each sliding window: obtain the first feature value of the first feature map within the current sliding window, compare the first feature value with the second feature value of the second feature map within the previous sliding window, and determine the feature map with the largest feature value as the second feature map after dimensionality reduction. M is an integer greater than 1.
7. The method according to claim 1, characterized in that, Filter the set of candidate bounding box maps according to prior conditions to obtain several target bounding box maps, including: For each candidate bounding box map in the set of candidate bounding box maps, calculate the label parameters of the candidate bounding box map; Determine whether the label parameters meet the prior conditions, where the prior conditions include at least one of the following: the minimum width-to-height ratio of the label, the maximum width-to-height ratio of the label, the maximum aspect ratio of the label, the minimum aspect ratio of the label, the presence of characters within the label, the hue difference ratio between the overall hue of the label and the hue of the box body; If the prior conditions are met, determine that the candidate bounding box map is a target bounding box map.
8. The method according to claim 7, wherein Calculating the label parameters of the candidate bounding box map includes at least one of the following: Calculate the spacing value between two parallel lines in the candidate bounding box map, determine the horizontal spacing as the width value, the vertical spacing as the height value, and the ratio between the horizontal spacing and the vertical spacing as the aspect ratio. Herein, the label parameters include: width value, height value, aspect ratio; Determine whether the candidate bounding box map contains characters through optical character recognition (OCR); Calculate the first grayscale value of the candidate bounding box map and calculate the second grayscale value of the packaging picture, and determine the ratio between the first grayscale value and the second grayscale value as the hue difference ratio.
9. An apparatus for automatic label modeling, characterized in that Including: An acquisition module, configured to acquire a set of suspected bounding box maps from the packaging picture of the target box body; A first filtering module, configured to filter the set of suspected bounding box maps by using a convolutional neural network to obtain a set of candidate bounding box maps; A second filtering module, configured to filter the set of candidate bounding box maps according to prior conditions to obtain several target bounding box maps. Herein, the target bounding box maps include the label information of the target box body, and the prior conditions are used to determine the target bounding box maps that meet the conditions in the set of candidate bounding box maps; A creation module, configured to create a label template of the target box body by using the several target bounding box maps and the packaging picture, where the label template is used to locate the pasting position of the label information on the target box body.
10. A storage medium, characterized in that, The storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 7 above.
11. An electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein, A processor, a communication interface, and a memory complete mutual communication through a communication bus; wherein: The memory is used to store a computer program; The processor is configured to execute the method according to any one of claims 1 to 7 by running the program stored on the memory.
Citation Information
Patent Citations
Target recognition device and method of industrial sorting robot
CN108846415A
Data joint modeling method, device, computer device and storage medium
CN109409922A