Label identification method and device, electronic equipment and storage medium
Through artificial intelligence technology, hierarchical feature extraction, attention relationship recognition and feature fusion are solved, and the problem of insufficient accuracy of manual tag recognition is realized, automated tag recognition is realized and the safety of logistics and transportation is improved.
Patent Information
- Application Number
- CN202311868867.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-30
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the accuracy of label identification by manual means is insufficient, resulting in the safety of parcels in logistics and transportation cannot be effectively guaranteed.
Using artificial intelligence technology, we obtain the target object image for hierarchical feature extraction, attention relationship recognition, feature fusion and label detection, obtain estimated tag information, and perform image cropping and category recognition to improve the accuracy of tag recognition.
Automatic label identification is realized, reducing the situation of misidentifying other labels as hazardous goods labels, and improving the safety of logistics and transportation and the speed of transportation management.
Smart Images

Figure CN120236111A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a label recognition method and device, an electronic device, and a storage medium. Background Art
[0002] Currently, in scenarios such as logistics transportation, it is necessary to paste labels on the transported packages to distinguish the categories of the objects inside the packages. For example, dangerous goods labels can be pasted on packages containing dangerous substances such as inflammable and explosive materials. When a certain package is identified as having a dangerous goods label pasted on it, the transportation method of such packages needs to be restricted. For example, such packages can only be transported by land and cannot be transported by air, so as to ensure the safety of package transportation.
[0003] In the related art, label recognition is performed manually or in other ways, and this recognition method will affect the accuracy of label recognition. Based on this, how to provide a label recognition method to improve the accuracy of label recognition has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a label recognition method and device, an electronic device, and a storage medium, aiming to improve the accuracy of label recognition.
[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a label recognition method, and the method includes:
[0006] Obtain a target object image of a target object;
[0007] Extract hierarchical features from the target object image to obtain initial image features corresponding to different levels;
[0008] Perform attention relationship recognition based on the initial image features of different levels to obtain attention features;
[0009] Perform feature fusion based on the attention features of different levels to obtain a target fusion feature;
[0010] Perform label detection based on the target fusion feature to obtain predicted label information;
[0011] Perform image cropping on the target object image according to the predicted label information to obtain a target label image;
[0012] Perform category recognition on the target label image to obtain a target label category.
[0013] In some embodiments, the performing feature fusion based on the attention features of different levels to obtain a target fusion feature includes:
[0014] Perform feature fusion based on the attention features of adjacent levels to obtain initial fusion features;
[0015] Perform feature fusion based on the initial fusion features to obtain the target fusion features.
[0016] In some embodiments, the performing category recognition based on the target label image to obtain a target label category includes:
[0017] Perform category recognition on the target label image according to a preset category recognition model to obtain the target label category;
[0018] Before performing category recognition on the target label image according to the preset category recognition model, the method further includes training the category recognition model, including:
[0019] Obtain a sample label image of a sample object and a sample label category of the sample label image;
[0020] Perform category recognition on the sample label image based on the category recognition model to obtain a predicted label category;
[0021] Calculate a similarity angle based on the sample label category and the predicted label category;
[0022] Train the category recognition model according to the similarity angle and a preset angle interval parameter.
[0023] In some embodiments, the estimated label information includes label position information and label confidence. The performing image cropping on the target object image according to the estimated label information to obtain a target label image includes:
[0024] If the label confidence is greater than or equal to a preset threshold, perform image cropping on the target object image based on the label position information corresponding to the label confidence to obtain an original label image;
[0025] Perform binarization processing on the original label image to obtain the target label image.
[0026] In some embodiments, the performing binarization processing on the original label image to obtain the target label image includes:
[0027] Perform multiple downsampling processes on the original label image to obtain multi-level label features;
[0028] Perform feature fusion on the multi-level label features to obtain the target label image.
[0029] In some embodiments, the target object image includes a first image and a second image. Obtaining the target object image of the target object includes:
[0030] Obtaining a first image of the target object;
[0031] Performing a flipping process on the target object based on a preset angle, and obtaining a second image of the target object.
[0032] In some embodiments, the method further includes:
[0033] Comparing the target label category with a preset comparison label category to obtain a comparison result;
[0034] Performing transportation management on the target object according to the comparison result and a preset transportation method.
[0035] To achieve the above object, a second aspect of the embodiments of the present application provides a label recognition device, the device includes:
[0036] An image acquisition unit, configured to acquire a target object image of a target object;
[0037] A feature extraction unit, configured to perform hierarchical feature extraction according to the target object image to obtain initial image features corresponding to different levels;
[0038] An attention recognition unit, configured to perform attention relationship recognition according to the initial image features of different levels to obtain attention features;
[0039] A feature fusion unit, configured to perform feature fusion according to the attention features of different levels to obtain a target fusion feature;
[0040] A label detection unit, configured to perform label detection according to the target fusion feature to obtain estimated label information;
[0041] An image cropping unit, configured to perform image cropping according to the target object image to obtain a target label image;
[0042] A category recognition unit, configured to perform category recognition according to the target label image to obtain a target label category.
[0043] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0044] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0045] The label recognition method and device, electronic device and storage medium proposed in the present application perform hierarchical feature extraction, attention relationship recognition, feature fusion and label detection operations based on the target object image, and can obtain the estimated label information of the target object image, and crop the target object image based on the estimated label information to obtain the target label image. According to the target label image, category recognition can be performed to obtain the corresponding target label category. It can be seen that the embodiment of the present application can automatically recognize the label of the target object image, and the hierarchical feature extraction, attention relationship recognition, feature fusion and label detection operations can mine the potential label information in the target object image. Therefore, compared with the manual recognition method in the related art, the embodiment of the present application can improve the accuracy of label recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of a tag identification method provided in an embodiment of the present application;
[0047] Figure 2 yes Figure 1 Flow chart of step S101 in FIG.
[0048] Figure 3 It is a structural diagram of a tag detection model provided in an embodiment of the present application;
[0049] Figure 4 yes Figure 1 Flow chart of step S105 in FIG.
[0050] Figure 5 yes Figure 1 Flow chart of step S107 in FIG.
[0051] Figure 6 yes Figure 5 Flow chart of step S503 in FIG.
[0052] Figure 7 is a flowchart of another embodiment of the tag identification method provided in an embodiment of the present application;
[0053] Figure 8A and 8B Schematic diagram of a label provided in an embodiment of the present application;
[0054] Figure 9 is a flowchart of another embodiment of the tag identification method provided in an embodiment of the present application;
[0055] Figure 10 It is a schematic structural diagram of the label recognition device provided by an embodiment of the present application;
[0056] Figure 11 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0058] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0060] First, several terms involved in the present application are analyzed:
[0061] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.
[0062] At present, in scenarios such as logistics transportation, it is necessary to paste labels on the transported packages to distinguish the categories of objects inside the packages. For example, a dangerous goods label can be pasted on a package containing dangerous substances such as inflammable and explosive substances. When a package is identified as having a dangerous goods label pasted on it, it is necessary to restrict the transportation method of such packages. For example, such packages can only be transported by land and cannot be transported by air, so as to ensure the safety of package transportation.
[0063] In the related art, label recognition is performed manually or in other ways, and this recognition method will affect the accuracy of label recognition.
[0064] Based on this, the embodiments of the present application provide a label recognition method, device, electronic device, and storage medium, aiming to improve the accuracy of label recognition.
[0065] The label recognition method, device, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the label recognition method in the embodiments of the present application is described.
[0066] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application system.
[0067] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0068] The label recognition method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The label recognition method provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the label recognition method, etc., but is not limited to the above forms.
[0069] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0070] Figure 1 is an alternative flowchart of the tag recognition method provided by an embodiment of this application, Figure 1 The method in may include but is not limited to steps S101 to S107.
[0071] Step S101, obtain a target object image of the target object;
[0072] Step S102, perform hierarchical feature extraction based on the target object image to obtain initial image features corresponding to different levels;
[0073] Step S103, perform attention relationship recognition based on the initial image features of different levels to obtain attention features;
[0074] Step S104, perform feature fusion based on the attention features of different levels to obtain a target fusion feature;
[0075] Step S105, perform tag detection based on the target fusion feature to obtain predicted tag information;
[0076] Step S106, perform image cropping based on the target object image to obtain a target tag image;
[0077] Step S107, perform category recognition based on the target tag image to obtain a target tag category.
[0078] In steps S101 to S107 shown in the embodiment of the present application, hierarchical feature extraction, attention relationship recognition, feature fusion and label detection operations are performed according to the target object image, and estimated label information of the target object image can be obtained, and the target object image is cropped based on the estimated label information to obtain the target label image. According to the target label image, category recognition can be performed to obtain the corresponding target label category. It can be seen that the embodiment of the present application can automatically identify the label of the target object image, and the hierarchical feature extraction, attention relationship recognition, feature fusion and label detection operations can mine the potential label information in the target object image. Therefore, compared with the manual recognition method in the related art, the embodiment of the present application can improve the accuracy of label recognition.
[0079] In step S101 of some embodiments, the target object may refer to an object to be tagged, for example, in a logistics transportation scenario, the target object may refer to a package. The target object image may refer to an image obtained by photographing the surface of the target object by a camera device. The camera device may be provided on a package sorting device.
[0080] Reference Figure 2 In some embodiments, the target object image includes a first image and a second image, and step S101 includes but is not limited to steps S201 to S202.
[0081] Step S201, acquiring a first image of a target object;
[0082] Step S202: Flip the target object based on a preset angle and acquire a second image of the target object.
[0083] In step S201 of some embodiments, the first image may refer to an image obtained by photographing any surface of the target object using a camera.
[0084] In step S202 of some embodiments, the target object may be flipped based on a preset flipping device, and the preset angle may refer to the angle at which the flipping device controls the rotation of the target object. The specific value of the preset angle may be adaptively set according to actual conditions, and this embodiment of the present application does not specifically limit this. The second image may refer to an image obtained by photographing the target object based on a camera device after the target object is flipped.
[0085] In some embodiments, step S202 can be repeatedly executed multiple times, so that multiple second images can be obtained. Taking the target object as a cuboid as an example, the first image can be the front image of the target object, and the multiple second images can include the left side image, the right side image, the bottom image, the top image, and the back image of the target object. Alternatively, the preset angle can be a relatively small value, so that images corresponding to the same label on the target object under the influence of different environmental factors (such as light factors) can be obtained.
[0086] The advantages of steps S201 to S202 are that they can improve the comprehensiveness of the target object image, reduce the situation of missing label recognition, and improve the accuracy of label recognition.
[0087] In addition, steps S201 to S202 describe a scenario with a single camera device. In some embodiments, multiple camera devices can also be set, and different camera devices are used to obtain images of different surfaces of the target object. The method of setting multiple camera devices can reduce the flipping of the multi-target object and reduce the damage to the target object.
[0088] In some embodiments, the initial image features can be processed based on a pre-set label detection model or other machine learning methods to obtain estimated label information. The label detection model can refer to a pre-set model with label detection capabilities. According to different target recognition labels, the label detection model can have different label detection capabilities. For example, when a dangerous goods label is used as the target recognition label, the label detection model can detect all labels suspected of being dangerous goods labels in the target object image. It can be understood that the label detection model can be based on the YOLOv5 model (the YOLOv5 model can use convolutional layers and feature pyramids to extract multi-scale features, enabling the model to capture the details and context information of the target and improving the detection accuracy) or other models (such as the PP-YOLOE model, etc. The PP-YOLOE model is an object detection model developed based on a deep learning framework. The PP-YOLOE model enhances the model's detection ability for small targets by introducing an attention mechanism). Specifically, as Figure 3 shown, the label detection model can include a feature extraction module, an attention module, a feature fusion module, and a label detection module.
[0089] In step S102 of some embodiments, the target object image is used as the input data of the feature extraction module, and hierarchical feature extraction is performed on the target object image based on the feature extraction module. It can be understood that hierarchical feature extraction can refer to performing feature extraction on the target object image in layers to obtain features of different scales. Therefore, the feature extraction module can output initial image features corresponding to different levels. For example, referring to Figure 3, the initial image features corresponding to level C2, the initial image features corresponding to level C3, the initial image features corresponding to level C4, and the initial image features corresponding to level C5 can be obtained. It can be understood that the scales of the above four initial image features are different.
[0090] In step S103 of some embodiments, the attention module may include multiple multi-head attention sub-modules. One multi-head attention sub-module is used to identify the attention relationship of the initial image features of one level to obtain the corresponding attention features. The attention features can be used to represent the relationship between the initial image features of the current level and the initial image features of other levels. For example, the attention module includes a multi-head attention sub-module corresponding to level C2, a multi-head attention sub-module corresponding to level C3, a multi-head attention sub-module corresponding to level C4, and a multi-head attention sub-module corresponding to level C5. Taking the multi-head attention sub-module corresponding to level C2 as an example, this attention sub-module is used to explore the relationship between the initial image features output by level C2 and the initial image features output by other levels, and output more distinguishable attention features.
[0091] In step S104 of some embodiments, multiple attention features are used as the input data of the feature fusion module to perform feature fusion on different attention features based on the feature fusion module to obtain the target fusion features.
[0092] Refer to Figure 4 , in some embodiments, the feature fusion module includes a first fusion module and a second fusion module, and step S105 includes but is not limited to steps S401 to S402.
[0093] Step S401, perform feature fusion according to the attention features of adjacent levels to obtain the initial fusion features;
[0094] Step S402, perform feature fusion according to the initial fusion features to obtain the target fusion features.
[0095] In step S401 of some embodiments, the first fusion module may be a feature pyramid structure, that is, the first fusion module may include a first structure for performing feature fusion from top to bottom and a second structure for performing feature fusion from bottom to top. For example, refer to Figure 3, corresponding to the four levels of the feature extraction module, the first structure may include four fusion sub-modules P5, P4, P3, and P2, and the second structure may include four fusion sub-modules H2, H3, H4, and H5. Each fusion sub-module is used to fuse the features of the corresponding level and the adjacent levels. Taking the fusion sub-module P4 as an example, this fusion sub-module can fuse the attention features corresponding to level C4 and the attention features corresponding to level C5 to achieve top-down fusion. The features obtained by fusing the fusion sub-module P4 are also input to the fusion sub-module H4, and the fusion sub-module H4 is used to fuse the features obtained by fusing the fusion sub-module H3 and the features obtained by fusing the fusion sub-module P4 to achieve bottom-up fusion. The features obtained by fusing each fusion sub-module in the second structure are used as the initial fusion features. Therefore, corresponding to the four fusion sub-modules H2, H3, H4, and H5, four initial image features can be obtained. The advantage of setting the first fusion module as a feature pyramid structure is that it can process features of different scales.
[0096] In step S402 of some embodiments, as Figure 3 shown, the feature scale continuously increases after passing through the first structure, and the feature scale continuously decreases after passing through the second structure. It can be seen from this that the multiple initial image features have different scales. To solve the problem of inconsistent learning objectives between the initial image features of different scales, the embodiment of the present application also sets a second fusion module in the feature fusion module. Specifically, the second fusion module can be constructed based on an adaptively spatial feature fusion (ASFF) module, and the second fusion module is used to learn a multi-scale feature fusion method, thereby reducing the problem of inconsistent learning of features of different scales. The second fusion module may include four fusion sub-modules ASFF1, ASFF2, ASFF3, and ASFF4. Each fusion sub-module can obtain multiple initial image features and perform feature fusion processing on the multiple initial image features. Based on the second fusion module, the final fusion feature, that is, the target fusion feature, can be obtained.
[0097] In step S105 of some embodiments, the target fusion feature is used as the input data of the label detection module to perform label detection on the target fusion feature based on the label detection module to obtain the estimated label information.
[0098] From the descriptions of steps S102 to S105, it can be seen that the embodiment of the present application sets the attention module as an independent model structure between the feature extraction module and the feature fusion module. Compared with the model structure of "C2 → attention module → C3 → attention module → C4 → attention module → C5", the embodiment of the present application can reduce the influence of the attention module on the features extracted from different levels, that is, the attention module only processes the initial image features extracted from the corresponding level.
[0099] In step S106 of some embodiments, the area of the region where the label is located is relatively small compared to the overall area of the target object image. In order to improve the accuracy of label recognition, the target object image can be cropped based on the estimated label information to obtain a target label image containing only label information, or a target label image containing label information and less noise.
[0100] Reference Figure 5 In some embodiments, the estimated tag information may include tag location information and tag confidence, and step S107 includes but is not limited to step S501 and step S502.
[0101] Step S501: if the label confidence is greater than or equal to a preset threshold, the target object image is cropped based on the label position information corresponding to the label confidence to obtain an original label image;
[0102] Step S502, binarization is performed on the original label image to obtain a target label image.
[0103] In step S501 of some embodiments, the label position information may refer to the estimated coordinates of each label in the target object image, and the label confidence may refer to the confidence value that the corresponding label is a target identification label, such as the confidence value that the corresponding label is a hazardous goods label. When the label confidence is less than a preset threshold, it indicates that the probability that the corresponding label is a target identification label is small. At this time, the label can be discarded, that is, the label will not be subsequently identified. When the label confidence is greater than or equal to the preset threshold, it indicates that the probability that the corresponding label is a target identification label is high. At this time, the target object image can be cropped based on the label position information of the label to obtain the image corresponding to the label, that is, the original label image.
[0104] In step S502 of some embodiments, the target object image is usually relatively dark, with basically no available color information. Therefore, in the embodiment of the present application, the image noise in the original label image is removed based on denoising processing, and a clearer image is obtained based on binarization processing. Specifically, the original label image can be denoised based on a filter-based method, a model-based method, a deep learning-based method, etc., which is not specifically limited in the embodiment of the present application. It is understandable that the filter method includes a Gaussian filter, and the model-based method includes a Markov random field model. After the original label image is denoised, the initial label image can be obtained.
[0105] By performing binarization on the initial label image, a clearer binarized image, i.e., the target label image, can be obtained. Among them, the methods of binarization can include canny edge detection (canny edge detection is an edge detection algorithm including grayscale processing, Gaussian filtering, non-maximum suppression, double-threshold processing, and edge connection), sobel edge detection (sobel edge detection is a method based on image gradient, and sobel edge detection detects edges by calculating the gradient intensity of each pixel in the image), and HED edge detection (HED edge detection is an algorithm that outputs edge detection results through an end-to-end neural network), etc. The embodiments of the present application do not make specific limitations on this.
[0106] Referring to Figure 6 , in some embodiments, step S503 includes but is not limited to steps S601 to S602.
[0107] Step S601, perform multiple downsampling processes on the initial label image to obtain multi-layer label features;
[0108] Step S602, perform feature fusion on the multi-layer label features to obtain the target label image.
[0109] In steps S601 to S602 of some embodiments, perform denoising on the original label image to obtain the initial label image, perform multiple downsampling processes on the initial label image, such as performing five downsamplings on the initial label image to obtain label features corresponding to five levels. Then, fuse the label features of these five levels and finally output a binarized image with the same size as the original image, i.e., the target label image. This processing method can remove irrelevant information in the initial label image and only retain the pattern information and edge information related to the label.
[0110] Referring to Figure 7 , in some embodiments, before step S108, the method provided by the embodiments of the present application may further include training the category recognition module, specifically including but not limited to steps S701 to S704.
[0111] Step S701, obtain the sample label image of the sample object and the sample label category of the sample label image;
[0112] Step S702, perform category recognition on the sample label image based on the category recognition model to obtain the predicted label category;
[0113] Step S703, calculate the similarity angle according to the sample label category and the predicted label category;
[0114] Step S704, train the category recognition model according to the similarity angle and the preset angle interval parameter.
[0115] In step S701 of some embodiments, the sample object may refer to the object to be identified for the label category. For example, in the logistics transportation scenario, the sample object may refer to a package. The sample label image may refer to the image of the label pasted on the surface of the sample object, and the sample label image may represent any type of label. The sample label category may refer to the true category of the corresponding label.
[0116] It can be understood that the sample label image may be an image obtained based on a label detection model, that is, the surface image of the sample object is detected based on the label detection model to obtain the corresponding label coordinates and confidence values. The surface image is cropped based on the label coordinates and confidence values to obtain the sample label image. The advantage of obtaining the sample label image through the label detection model is that it can realize the joint training of the sample detection model and the category recognition model, improve the adaptability of the sample detection model and the category recognition model, and thus improve the accuracy of label recognition.
[0117] In step S702 of some embodiments, the category recognition model may be a model with the ability to recognize label categories. The category recognition model may use MoCoVit as the backbone network, and MoCoVit is a model constructed based on a lightweight transformer (the transformer model is a neural network model based on the self-attention mechanism). The sample label image is used as the input data of the category recognition model, and the category recognition model is used to perform category recognition on the sample label image to obtain the predicted label category corresponding to the sample label image.
[0118] In steps S703 to S704 of some embodiments, the category recognition loss of the category recognition model may be determined based on the predicted label category and the sample label category. Based on this category recognition loss, the model parameters of the category recognition model may be adjusted, thereby realizing the training of the category recognition model. Specifically, the category recognition loss L may be calculated according to the following formula:
[0119]
[0120] where, y i represents the sample label category. θ yi represents the angle between the feature vector corresponding to the normalized sample label category and the reference vector in the feature space, that is, θ yi represents the similarity angle, and θ yi can be calculated based on the predicted label category and the sample label category. m represents a preset angle interval parameter, which can usually be set to 0.5. m can be used to increase the angle spacing between the same categories, thereby increasing the loss value, and further prompting the predicted label categories output by the category recognition model to have the smallest intra-class distance and the largest inter-class distance. s represents a feature scaling constant, which can usually be set to 60.
[0121] The advantage of training the category recognition model based on the above formula in the embodiments of the present application is that in an actual scenario, the label category can be determined based on the color and content of the label. For example, in an actual scenario, Figure 8A the label representing the "flammable liquid" category is red, Figure 8B the label representing the "flammable solid" category is blue. Therefore, these two types of labels can be distinguished based on the color. However, since the sample label image is a binarized image and has only black and white colors, only the Figure 8A number "3" in Figure 8B and the number "4" in
[0122] In step S107 of some embodiments, the category of the target label image can be recognized based on a preset category recognition model or other machine learning methods to obtain the target label category. For example, the target label image can be used as the input data of the category recognition model to recognize the category of the target label image based on the category recognition model, so as to obtain the target label category. It can be understood that the target label category can be the label category closest to the target label image confirmed from the preset label categories. The number of preset label categories can be adaptively set according to the actual target recognition label, and the embodiments of the present application do not make specific limitations in this regard. Taking the target recognition label as a dangerous goods label as an example, the categories corresponding to the dangerous goods label include nine categories: "explosives", "gases", "flammable liquids", "flammable solids", "oxidizing substances and organic peroxides", "toxic substances and infectious substances", "radioactive substances", "corrosives", and "miscellaneous dangerous substances and articles". Therefore, these nine categories and the "non-dangerous goods label" category can be used as the preset label categories.
[0123] It can be understood that in some embodiments, when the target object image is a set of a first image and multiple second images, and the preset angle is a relatively small angle, the first image and the multiple second images can be processed according to multiple threads. Each thread recognizes the image in the manner described in steps S102 to S107, so as to obtain multiple target label categories. In order to improve the accuracy of label recognition, the category with the highest repetition rate among the multiple target label categories can be used as the final label category, thereby reducing the influence of environmental factors on the accuracy of label recognition.
[0124] Refer to Figure 9, in some embodiments, the label recognition method provided by the embodiments of the present application may further include, but is not limited to, steps S901 to S902.
[0125] Step S901, comparing the target label category with a preset comparison label category to obtain a comparison result;
[0126] Step S902, performing transportation management on the target object according to the comparison result and a preset transportation method.
[0127] In step S901 of some embodiments, the comparison label category may be any one or more categories determined from preset label categories according to actual needs. For example, when it is necessary to intercept an object of "gas", the preset label category of "gas" may be used as the comparison label category. When it is necessary to intercept an object of "explosive" and an object of "corrosive", the preset label category of "explosive" and the preset label category of "corrosive" may be used as the comparison label categories. Comparing the target label category with the comparison label category to obtain a comparison result that the target label category is consistent with the comparison label category, or obtaining a comparison result that the target label category is inconsistent with the comparison label category. It can be understood that when the comparison label category includes multiple categories such as "explosive" and "corrosive", as long as the target label category is the same as any one of them, a comparison result of category consistency can be obtained.
[0128] In step S902 of some embodiments, multiple transportation methods may be preset, such as including land transportation and air transportation, etc. Different transportation methods may be mapped to different comparison results. According to the comparison result and the mapping relationship, the final transportation method of the target object can be confirmed, thereby improving the transportation safety of the target object.
[0129] It can be understood that, in some embodiments, when the target object image is a set of six-face images of a cuboid, the label recognition operation described in the above embodiments may be performed on the six-face images in sequence. When the target label category of the label of any one face image recognized is the category corresponding to the dangerous goods label, the corresponding target object is intercepted, thereby improving the transportation safety of the target object.
[0130] The label recognition method provided by the embodiments of the present application has the following beneficial effects:
[0131] (1) The method for label recognition based on a label detection model and a category recognition model in the embodiments of the present application can reduce the situation of misidentifying other labels (such as "fragile" label, "keep dry" label, etc.) as dangerous goods labels compared with the method of simultaneously realizing label positioning and category recognition based on a single detection model in the related art;
[0132] (2) It can achieve the recognition of multiple label categories. For example, when setting ten preset label categories, the recognition of ten label categories can be achieved;
[0133] (3) Based on the lightweight structure of the category recognition model, the embodiments of the present application can improve the speed of label recognition, thereby improving the speed of transportation management of target objects;
[0134] (4) Based on the training method of the category recognition model, the embodiments of the present application can achieve the recognition of subtle differences in images, thereby improving the accuracy of label recognition.
[0135] Referring to Figure 10 , the embodiments of the present application also provide a label recognition device, which can implement the above label recognition method. The device includes:
[0136] An image acquisition unit 1001, configured to acquire a target object image of a target object;
[0137] A feature extraction unit 1002, configured to perform hierarchical feature extraction according to the target object image to obtain initial image features corresponding to different levels;
[0138] An attention recognition unit 1003, configured to perform attention relationship recognition according to the initial image features of different levels to obtain attention features;
[0139] A feature fusion unit 1004, configured to perform feature fusion according to the attention features of different levels to obtain a target fusion feature;
[0140] A label detection unit 1005, configured to perform label detection according to the target fusion feature to obtain estimated label information;
[0141] An image cropping unit 1006, configured to crop the image according to the target object image to obtain a target label image;
[0142] A category recognition unit 1007, configured to perform category recognition according to the target label image to obtain a target label category.
[0143] In some embodiments, the attention recognition unit 1003 is further configured to:
[0144] Perform feature fusion according to the attention features of adjacent levels to obtain an initial fusion feature;
[0145] Perform feature fusion according to the initial fusion feature to obtain a target fusion feature.
[0146] In some embodiments, the category recognition unit 1007 is further configured to:
[0147] Obtain a sample label image of a sample object and a sample label category of the sample label image;
[0148] Perform category recognition on the sample label image based on the category recognition model to obtain the predicted label category;
[0149] Calculate the similarity angle based on the sample label category and the predicted label category;
[0150] Train the category recognition model according to the similarity angle and the preset angle interval parameter.
[0151] In some embodiments, the image cropping unit 1006 is further configured to:
[0152] If the label confidence is greater than or equal to the preset threshold, perform image cropping on the target object image based on the label position information corresponding to the label confidence to obtain the original label image;
[0153] Perform binarization processing on the original label image to obtain the target label image.
[0154] In some embodiments, the image cropping unit 1006 is further configured to:
[0155] Perform multiple downsampling processes on the original label image to obtain multi-layer label features;
[0156] Perform feature fusion on the multi-layer label features to obtain the target label image.
[0157] In some embodiments, the image acquisition unit 1001 is further configured to:
[0158] Acquire the first image of the target object;
[0159] Perform flipping processing on the target object based on a preset angle and acquire the second image of the target object.
[0160] In some embodiments, the label recognition device further includes a transportation unit, configured to:
[0161] Compare the target label category with a preset comparison label category to obtain a comparison result;
[0162] Perform transportation management on the target object according to the comparison result and the preset transportation method.
[0163] The specific implementation manner of this label recognition device is basically the same as the specific embodiments of the above label recognition method, and will not be elaborated here.
[0164] An embodiment of this application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above label recognition method. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0165] Refer toFigure 11 , Figure 11 schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0166] a processor 1101, which can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0167] a memory 1102, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102 and are called by the processor 1101 to execute the label recognition method of the embodiments of the present application;
[0168] an input / output interface 1103, which is used to implement information input and output;
[0169] a communication interface 1104, which is used to implement communication interaction between this device and other devices, and can communicate through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0170] a bus 1105, which transmits information between various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);
[0171] wherein the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 are communicatively connected to each other inside the device through the bus 1105.
[0172] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned label recognition method is implemented.
[0173] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0174] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0175] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0178] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above figures are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0179] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0180] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0181] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0183] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0184] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of rights of the embodiments of this application.
Claims
1. A label recognition method, characterized in that, The method includes: Obtaining a target object image of a target object; Performing hierarchical feature extraction based on the target object image to obtain initial image features corresponding to different levels; Performing attention relationship recognition based on the initial image features of different levels to obtain attention features; Performing feature fusion based on the attention features of different levels to obtain a target fusion feature; Performing label detection based on the target fusion feature to obtain predicted label information; Performing image cropping on the target object image according to the predicted label information to obtain a target label image; Performing category recognition based on the target label image to obtain a target label category.
2. The method according to claim 1, wherein The performing feature fusion based on the attention features of different levels to obtain a target fusion feature includes: Performing feature fusion based on the attention features of adjacent levels to obtain an initial fusion feature; Performing feature fusion based on the initial fusion feature to obtain the target fusion feature.
3. The method according to claim 1, wherein the performing category recognition based on the target label image to obtain a target label category includes: Performing category recognition on the target label image according to a preset category recognition model to obtain the target label category; Before performing category recognition on the target label image according to the preset category recognition model, the method further includes training the category recognition model, including: Obtaining a sample label image of a sample object and a sample label category of the sample label image; Performing category recognition on the sample label image based on the category recognition model to obtain a predicted label category; Calculating a similarity angle according to the sample label category and the predicted label category; Training the category recognition model according to the similarity angle and a preset angle interval parameter.
4. The method according to claim 1, characterized in that The predicted label information includes label position information and label confidence. The performing image cropping on the target object image according to the predicted label information to obtain a target label image includes: If the label confidence is greater than or equal to a preset threshold, performing image cropping on the target object image based on the label position information corresponding to the label confidence to obtain an original label image; Performing binarization processing on the original label image to obtain the target label image.
5. The method according to claim 4, characterized in that The performing binarization processing on the original label image to obtain the target label image includes: Performing multiple downsampling processes on the original label image to obtain multi-layer label features; Performing feature fusion on the multi-layer label features to obtain the target label image.
6. The method according to any one of claims 1 to 5, characterized in that, The target object image includes a first image and a second image. The obtaining a target object image of a target object includes: Obtaining a first image of a target object; Performing a flipping process on the target object based on a preset angle and obtaining a second image of the target object.
7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Comparing the target label category with a preset comparison label category to obtain a comparison result; Performing transportation management on the target object according to the comparison result and a preset transportation method.
8. A label recognition device, characterized in that, The device includes: An image acquisition unit for obtaining a target object image of a target object; A feature extraction unit, configured to perform hierarchical feature extraction based on the target object image to obtain initial image features corresponding to different levels; An attention recognition unit, configured to perform attention relationship recognition based on the initial image features at different levels to obtain attention features; A feature fusion unit, configured to perform feature fusion based on the attention features at different levels to obtain target fusion features; A label detection unit, configured to perform label detection based on the target fusion features to obtain estimated label information; An image cropping unit, configured to perform image cropping based on the target object image to obtain a target label image; A category recognition unit, configured to perform category recognition based on the target label image to obtain a target label category.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.