Object attribute extraction method, device, storage medium and electronic device

Through an extraction model, multiple attributes are extracted from the initial features of the target object, and the attention mechanism and convolutional neural network are used to divide the region and integrate the structure, content and color attributes. This solves the problems of high data processing complexity and waste of storage resources in the existing technology, and realizes efficient and accurate attribute extraction.

CN114863414BActive Publication Date: 2025-09-30SHANGHAI SHANMA INTELLIGENT TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210496186.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-09-30
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

In the existing technology, when extracting attributes of a target object with a single-layer or double-layer structure, it is necessary to first classify it into a single-layer or double-layer structure, and then use different algorithm models for identification, which leads to high data processing complexity and waste of storage resources, and does not fully utilize the correlation between content attributes and structural attributes.

Method used

An extraction model is used to extract multiple attributes from the initial features of the target object. The target object is divided into multiple regions using the attention mechanism and convolutional neural network. The structure, content and color attributes are fused to obtain the overall attributes by multiplying and concatenating the domain value and content data.

Benefits of technology

The algorithm process is simplified, the efficiency of attribute extraction is improved, the number of models is reduced, and the accuracy of attribute extraction is improved by utilizing the correlation between content and structure attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863414B_ABST
    Figure CN114863414B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for extracting object attributes, relating to the field of data processing. Initial features of a target object are obtained, and multiple attributes of the target object are extracted from the initial features using an extraction model. When the structural attributes of the target object differ, the extraction model employs a unique algorithm, and the multiple attributes of the target object are integrated to obtain the overall attributes of the target object. This method addresses the problem of complex methods for extracting target object attributes, simplifies the process for extracting target object attributes, and improves the efficiency of target object attribute extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing, and in particular, to a method, device, storage medium and electronic device for extracting object attributes in the field of transportation. Background Art

[0002] In existing technologies, one approach to extracting attributes from single-layer or double-layered target objects is to first classify the objects as single-layer or double-layer, and then use algorithmic models corresponding to these structures to identify their content. This method extracts the target objects' content and structural attributes separately. Because the content and structural attributes of target objects are somewhat correlated, extracting them separately is not only inconvenient, but also requires pre-determining the structural attributes, increasing data processing complexity and wasting storage resources.

[0003] Most existing technologies for identifying double-layered objects pre-separate the object into single-layer and double-layer structures, then employ targeted techniques to identify each separately. A small number of techniques involve combining single-layer and double-layer structures for recognition, but these techniques require complex pre- and post-processing methods to unify the structure of the object, followed by individual recognition of its contents.

[0004] The above method for extracting target object attributes is typically used in text and image data processing scenarios, especially in the processing of odd and even license plate attributes.

[0005] In short, in the prior art, when a target object has multiple structures, the algorithm for extracting the attribute is relatively complex and has a high degree of difficulty. Summary of the Invention

[0006] The main advantage of the present invention is to provide an object attribute extraction method, device, storage medium and electronic device, which can simultaneously extract multiple attributes of a target object, simplify the target object attribute extraction process, and improve the efficiency of target object attribute extraction;

[0007] The main advantage of the present invention is to provide an object attribute extraction method, device, storage medium and electronic device, which use one extraction model to extract multiple attributes of the target object, reducing the number of models and simplifying the algorithm process;

[0008] Another advantage of the present invention is that it provides an object attribute extraction method, device, storage medium and electronic device, which continuously train the extraction model by utilizing the correlation between the attributes of the target object, so that the information extracted by the extraction model is more accurate during the process of extracting the target object attributes.

[0009] According to an embodiment of the present invention, a method for extracting object attributes is provided, comprising:

[0010] Obtain the initial features of the target object;

[0011] Extracting a plurality of attributes of the target object from the initial features of the target object using an extraction model, wherein when the structural attributes of the target object are different, the algorithm of the extraction model is unique;

[0012] The multiple attributes of the target object are fused to obtain the overall attribute of the target object.

[0013] According to an embodiment of the present invention, extracting the multiple attributes of the target object from the initial features of the target object using the extraction model includes at least one of the following:

[0014] extracting structural attributes of the target object from the initial features of the target object using an extraction model;

[0015] extracting content attributes of the target object from the initial features of the target object using an extraction model;

[0016] Extracting the color attribute of the target object from the initial features of the target object using an extraction model;

[0017] According to an embodiment of the present invention, extracting the content attributes of the target object from the initial features of the target object using an extraction model includes:

[0018] Dividing the target object into multiple regions using an attention mechanism;

[0019] Values ​​are assigned to the multiple regions respectively to obtain corresponding domain values, wherein the domain values ​​are used to confirm the content attribute of the target object.

[0020] According to an embodiment of the present invention, extracting the content attributes of the target object from the initial features of the target object using an extraction model includes:

[0021] Multiplying the content data of each region by the corresponding domain value to obtain a content representation of the corresponding region;

[0022] splicing the content representations of each of the regions to obtain an overall content representation of the target object;

[0023] The overall content representation of the target object is parsed to obtain the content attributes of the target object.

[0024] According to one embodiment of the present invention, when the target content information does not exist in the area, the content representation obtained by multiplying the content data of the area by the corresponding domain value is used as the first threshold; when the target content information exists in the area, the content representation obtained by multiplying the content data of the area by the corresponding domain value is still the content data of the area.

[0025] According to another embodiment of the present invention, there is provided an object attribute extraction device, comprising:

[0026] An acquisition module is used to obtain the initial features of the target object;

[0027] an extraction module, configured to extract a plurality of attributes of the target object from the initial features of the target object; wherein when the structural attributes of the target object are different, the algorithm of the extraction model is unique;

[0028] The fusion module is used to fuse the multiple attributes of the target object to obtain the overall attribute of the target object.

[0029] According to another embodiment of the present invention, the extraction module further comprises:

[0030] A region division module, wherein the region division module is configured to divide each target object into multiple regions using the attention mechanism.

[0031] A region assignment module, wherein the region assignment module is configured to assign values ​​to the regions of the target object based on the initial features of the target object using a convolutional neural network to obtain corresponding domain values, and to assign values ​​to the multiple regions to obtain corresponding domain values, wherein the domain values ​​are used to confirm the content attributes of the target object.

[0032] According to another embodiment of the present invention, the extraction module further includes a content processing module, wherein the content processing module includes:

[0033] a calculation unit, wherein the calculation unit is configured to multiply the content data of the region by the domain value to obtain a content representation of the region;

[0034] a splicing unit, configured to splice the content representations of all the regions to obtain an overall content representation of the target object; and

[0035] A parsing unit is used to parse the overall content representation of the target object to obtain content attributes of the target object.

[0036] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.

[0037] According to another embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a hardware structure block diagram of a mobile terminal according to an object attribute extraction method of an embodiment of the present invention;

[0039] Figure 2 is a flow chart of a method for extracting object attributes according to an embodiment of the present invention;

[0040] Figure 3 is a flow chart of a method for extracting object attributes according to an embodiment of the present invention;

[0041] Figure 4 4 is a structural block diagram of an object attribute extraction device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.

[0043] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0044] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG is a hardware structure block diagram of a mobile terminal of an object attribute extraction method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1More or fewer components than shown, or with Figure 1 Different configurations shown.

[0045] Memory 104 can be used to store computer programs, such as software programs and modules of application software, such as a computer program corresponding to an object attribute extraction method in an embodiment of the present invention. Processor 102 executes the computer programs stored in memory 104 to execute various functional applications and data processing, thereby implementing the aforementioned method. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory remotely located relative to processor 102, and such remote memory may be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0046] Transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the mobile terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, configured to communicate with the Internet wirelessly.

[0047] Existing technology shows that when processing target objects with both single-layer and double-layer structures, two different methods are often used for attribute recognition. Therefore, when target objects exhibit multiple structural types, a more concise algorithm is needed to uniformly process and extract attributes from target objects of different structures. Specifically, an algorithm can simultaneously extract both the content and structural attributes of the target object and fuse them together to obtain the overall attributes of the target object.

[0048] According to a preferred embodiment of the present invention, Figure 2 As shown, the present invention provides an object attribute extraction method, comprising the following steps:

[0049] Obtain the initial features of the target object;

[0050] Extracting a plurality of attributes of the target object from the initial features of the target object using an extraction model, wherein when the structural attributes of the target object are different, the algorithm of the extraction model is unique;

[0051] fusing the plurality of attributes of the target object to obtain an overall attribute of the target object;

[0052] The initial features of the target object are obtained as follows:

[0053] Obtain the target object;

[0054] performing data encoding on the target object;

[0055] More specifically, a convolutional neural network (CNN) is used to perform a convolution operation on the target object to obtain initial features of the target object.

[0056] The initial features of the target object are encoded data obtained after the target object is subjected to a convolution operation. Preferably, the encoded data is original encoded data containing all attributes of the target object. Based on the encoded data, a plurality of specific attribute features of the target object are obtained through an extraction model, such as content attributes, structural attributes, and color attributes. The attribute features of the target object include but are not limited to content attributes, structural attributes, and color attributes.

[0057] According to some embodiments of the present invention, the target object is a picture or video containing characters.

[0058] According to a preferred embodiment of the present invention, the object attribute extraction method, apparatus, storage medium, and electronic device provided by the present invention are preferably applied in the field of transportation, particularly in scenarios involving the extraction of vehicle license plate information. In other embodiments of the present invention, the present invention can also be applied in scenarios involving text recognition.

[0059] According to an embodiment of the present invention, extracting multiple attributes of the target object from the initial features of the target object using an extraction model includes one of the following:

[0060] Extracting structural attributes of the target object from the initial features of the target object through an attention mechanism;

[0061] Extracting content attributes of the target object from the initial features of the target object through an attention mechanism;

[0062] Extracting the color attribute of the target object from the initial features of the target object through an attention mechanism;

[0063] Preferably, the structural attributes, content attributes, and color attributes of the target object are extracted simultaneously by the extraction model. It is understandable that, in some embodiments, the structural attributes, content attributes, and color attributes of the target object may also be extracted in different orders.

[0064] According to some embodiments of the present invention, when the target object is a single-layer or double-layer license plate, the structural attribute of the target object is single-layer or double-layer. It is worth mentioning that the structural attribute of the target object includes but is not limited to single-layer or double-layer, and can also have three layers and four layers. The number of layers of this structural attribute is not limited, and the extraction model requires a preset algorithm model of the corresponding structure.

[0065] According to a preferred embodiment of the present invention, the content attribute of the target object is the specific content contained in the target object. Generally, the specific content is presented in the form of characters. For example, when the target object is a text image, the specific content is text. When the target object is a license plate image, the specific content is a combination of text, numbers, letters, or symbols.

[0066] According to a preferred embodiment of the present invention, the color attribute of the target object is the color specifically presented by the target object. For example, when the target object is a license plate image, the color attribute is the color presented by the license plate.

[0067] According to a preferred embodiment of the present invention, the extraction model may be pre-set with algorithmic models corresponding to multiple attributes of the target object to be extracted. For example, in this embodiment, to extract the content attribute, structure attribute, and color attribute of the target object, the extraction model needs to be pre-set with algorithmic models corresponding to the content attribute, structure attribute, and color attribute of the target object to be extracted.

[0068] According to an embodiment of the present invention, extracting the color attribute of the target object from the initial features of the target object through an attention mechanism includes:

[0069] Extracting the color content representation of the target object from the initial features of the target object through an attention mechanism;

[0070] parsing the color content representation to obtain the color attribute;

[0071] According to an embodiment of the present invention, extracting the structural attributes of the target object from the initial features of the target object through an attention mechanism includes:

[0072] Extracting the structural content representation of the target object from the initial features of the target object through an attention mechanism;

[0073] parsing the structural content representation to obtain the structural attributes;

[0074] According to a preferred embodiment of the present invention, extracting the content attributes of the target object from the initial features of the target object through an attention mechanism includes:

[0075] The target object is divided into multiple regions using the attention mechanism.

[0076] Specifically, according to the structural attributes of the target object corresponding to the algorithm model of the preset attention mechanism, the attention mechanism can automatically divide each target object into a first area and a second area. The first area and the second area can be located at different positions of the target object respectively. For example, the first area is located in the upper area of ​​the target object, and the second area is located in the lower area of ​​the target object; in some other embodiments, the first area is located in the lower area of ​​the target object, and the second area is located in the upper area of ​​the target object. It is worth mentioning that in some embodiments, the first area and the second area can also be located in the left area and the right area of ​​the target object. The positions of the first area and the second area of ​​the target object may vary in different embodiments, which will not be repeated here.

[0077] According to a preferred embodiment of the present invention, the first area and the second area include content data of the target object in the area, i.e., first area content data and second area content data, wherein the content data is derived from the initial features of the target object, i.e., the content data of the first area and the second area are presented in the form of encoded data.

[0078] According to a preferred embodiment of the present invention, extracting the content attributes of the target object from the initial features of the target object using an extraction model includes:

[0079] A convolutional neural network is used to assign values ​​to the multiple regions respectively to obtain corresponding domain values.

[0080] The domain value is used to determine the content attribute of the target object.

[0081] Specifically, based on the initial features of the target object, a convolutional neural network (CNN) is used to assign values ​​to the first region and the second region of the target object to obtain a first domain value and a second domain value.

[0082] The assigning of values ​​to the first region and the second region of the target object using a convolutional neural network (CNN) includes:

[0083] Extracting a mask of the target object using a convolutional neural network;

[0084] By using a convolutional neural network to determine whether the mask contains target content information, the first area and the second area are assigned values ​​to obtain a first domain value and a second domain value.

[0085] Because the initial features of the target object are raw encoded data containing all of the target object's attributes, the convolutional neural network can extract the target object's mask based on the initial features, i.e., the raw encoded data. Therefore, when the target object is divided into a first region and a second region by the attention mechanism, the first region's corresponding mask is extracted, and the second region's corresponding mask is extracted.

[0086] Furthermore, a convolutional neural network (CNN) is used to detect whether the target object mask contains target information. When the target object mask does not detect target information, the area corresponding to the mask is assigned a value of 0. When the target object mask detects target information, the area corresponding to the mask is assigned a value of 1. That is, when the first area mask or the second area mask of the target object does not detect any target information, the first area or the second area is assigned a value of 0, that is, the first domain value or the second domain value is 0; when the first area mask or the second area mask of the target object detects target information, the first area or the second area is assigned a value of 1, that is, the first domain value or the second domain value is 1.

[0087] It is worth mentioning that because the algorithm model preset in the extraction model, that is, the algorithm model of the attention mechanism, always divides the target object into the first region and the second region, the first region mask and the second region mask of the target object always exist. In other words, in this embodiment, no matter how many regions the target object is divided into, each region of the target object will correspond to a layer of mask. Therefore, in this embodiment, the first region mask and the second region mask of the target object always exist.

[0088] In other words, in this embodiment, regardless of whether the target object has a single-layer or double-layer structure, the target object is divided into the first region and the second region. Correspondingly, regardless of whether the target object has a single-layer or double-layer structure, the first region mask and the second region mask are extracted from the target object. This means that when the structural attributes of the target object vary, for example, whether a license plate has a single-layer or double-layer structure, the extraction model algorithm remains unique.

[0089] According to a preferred embodiment of the present invention, Figure 3As shown, extracting the content attributes of the target object from the initial features of the target object using the extraction model also includes:

[0090] Multiplying the first region content data by the first threshold value, and multiplying the second region content data by the second threshold value, respectively, to obtain a first region content representation and a second region content representation;

[0091] splicing the first region content representation and the second region content representation to obtain an overall content representation of the target object;

[0092] Parsing the overall content representation of the target object to obtain content attributes of the target object;

[0093] According to a preferred embodiment of the present invention, the first region content representation and the second region content representation may be spliced ​​together in a horizontal manner or in other splicing manners, which are not limited herein. The splicing manner may be preset in the extraction model.

[0094] According to an embodiment of the present invention, the first region content data is multiplied by the first threshold value, and the second region content data is multiplied by the second threshold value, respectively, to obtain a first region content representation and a second region content representation, specifically including:

[0095] When the first domain value or the second domain value is 0, it means that no target information is detected by the first area mask or the second area mask, which means that there is no content information in the first area or the second area of ​​the target object, and the corresponding first area content data or the second area content data is also 0.

[0096] When the first domain value or the second domain value is 1, it means that the first area mask or the second area mask has detected target information, which means that there is valid target content information in the first area or the second area of ​​the target object, and the corresponding first area content data or the second area content data is the encoded data of the corresponding target content information.

[0097] Therefore, when the first threshold value or the second threshold value is 0, the first region content representation obtained by multiplying the first region content data by the first threshold value is 0, i.e., the first threshold value, or the second region content representation obtained by multiplying the second region content data by the second threshold value is 0, i.e., the first threshold value;

[0098] Therefore, when the first domain value or the second domain value is 1, the first region content representation obtained by multiplying the first region content data by the first domain value is still the first region content data, or the second region content representation obtained by multiplying the second region content data by the second domain value is still the second region content data.

[0099] According to an embodiment of the present invention, when the target object is a single-layer yellow license plate with the number plate number being Shanghai AB0010, at this time, the license plate is divided into a first area and a second area. Assuming that the first area is the upper area and the second area is the lower area, then the upper area of this single-layer license plate, that is, the first area, actually has no content data. That is, at this time, the content data of the first area does not exist and is 0, and the first threshold value of this single-layer license plate is 0. And the content data of the second area of the lower area of this single-layer license plate, that is, the second area, is the encoded data of Shanghai AB0010. When there is valid target content data in the second area of this single-layer license plate, the second threshold value is 1. Therefore, at this time, the content representation of the first area obtained by multiplying the content data of the first area by the first threshold value is 0, and the content representation of the second area obtained by multiplying the content data of the second area by the second threshold value is still the content data of the second area, that is, the encoded data of Shanghai AB0010.

[0100] Further, splicing the content representation of the first area being 0 and the content representation of the second area being the encoded data of Shanghai AB0010, the overall content representation of the target object, that is, the single-layer yellow license plate with the number plate number being Shanghai AB0010, is obtained, that is, the encoded data of Shanghai AB0010. Analyzing the overall content representation of the target object, that is, the encoded data of Shanghai AB0010, the overall content attribute of the target object, that is, Shanghai AB0010, is obtained.

[0101] According to an embodiment of the present invention, when the target object is a double-layer yellow license plate with the number plate number being Shanghai DB1310, as shown in the figure, at this time, the license plate is divided into a first area and a second area. Assuming that the first area is the upper area and the second area is the lower area, then the content data of the first area of the upper area of this double-layer license plate, that is, the first area, is the encoded data of Shanghai D. When there is valid target content data in the first area of this double-layer license plate, the first threshold value is 1. And the content data of the second area of the lower area of this double-layer license plate, that is, the second area, is the encoded data of B1310. When there is valid target content data in the second area of this double-layer license plate, the second threshold value is 1. Therefore, at this time, the content representation of the first area obtained by multiplying the content data of the first area by the first threshold value is still the content data of the first area, that is, the encoded data of Shanghai D. And the content representation of the second area obtained by multiplying the content data of the second area by the second threshold value is still the content data of the second area, that is, the encoded data of B1310.

[0102] Further, splice the encoded data representing the license plate number "Hu D" in the first region and the encoded data representing "B1310" in the second region to obtain the overall content representation of the target object, i.e., the double-layer yellow license plate "Hu D B1310", which is the encoded data of "Hu D B1310". Analyze the overall content representation of the target object, i.e., the encoded data of "Hu D B1310", to obtain the overall content attribute of the target object, which is "Hu D B1310".

[0103] As can be seen from the above embodiments, regardless of whether the target object is a single-layer structure or a double-layer structure, that is, in the embodiments, regardless of whether the target object is a single-layer license plate or a double-layer license plate, the target object is divided into a first region and a second region, and then the content representations finally obtained for each region are spliced, which unifies the algorithms when the target object has different structures. In the prior art, different algorithm models need to be adopted for target objects with different structures. There are more algorithm models and the algorithm processes are more complex. The object attribute extraction method provided by the present invention reduces the algorithm models and simplifies the algorithm processes.

[0104] According to a preferred embodiment of the present invention, fusing multiple attributes of the target object to obtain the overall attribute of the target object includes:

[0105] Fusing the content attribute, the structure attribute and the color attribute of the target object to obtain the overall attribute of the target object.

[0106] Specifically, the overall attribute of the target object is composed of the structure attribute, the color attribute and the content attribute. For example, when the structure attribute is double-layer, the color attribute is yellow, and the content attribute is "Hu A B0010", the overall attribute of the target object is obtained as: double-layer yellow license plate "Hu A B0010". The word order of the overall attribute in this embodiment is not restricted and can be set according to specific requirements.

[0107] It is worth mentioning that the overall attributes of the target object, that is, the license plate, such as the double-layer yellow license plate "Hu AB0010", can obtain the vehicle type information as a truck or a large vehicle through common knowledge in the art. Information users (such as traffic managers) can perform highway toll collection based on the overall license plate information. When the overall attribute of the obtained license plate is a single-layer blue license plate "Hu AB0010" and the image of the vehicle itself shows a truck, information users (such as traffic managers) can judge that there is a problem with the license plate based on the overall license plate information, and then further take measures to judge whether there is an act of license plate cloning or an act of evading tolls by changing the vehicle type. Therefore, by integrating the content attribute, the structural attribute, and the color attribute of the target object to obtain the overall attribute of the target object, the overall attribute will have significant beneficial effects in the scenario of highway toll collection according to the vehicle type. Thus, the usefulness of the information is further improved.

[0108] According to a preferred embodiment of the present invention, the extraction model is a trained convolutional neural network model for extracting the attributes of the target object. The extraction model provided in this embodiment is trained by a large number of pictures of sample objects and the attribute data of corresponding samples, so as to extract multiple attribute information from the target object.

[0109] According to an embodiment of the present invention, the extraction model can be continuously updated and learned during the process of extracting the attributes of the target object. Specifically, by using the correlation between multiple attributes of the target image as additional supervision information to confirm whether the relevant attributes extracted by the extraction model are correct. If the additional supervision information is inconsistent with the relevant attributes extracted by the extraction model, the target loss is calculated, and then the extraction model is trained using the target loss.

[0110] According to the embodiment of the present invention, when the threshold value of the first region mask of the target object is 0, there is no content data in the first region; when the threshold value of the first region mask of the target object is 1, there is target content data in the first region. That is, when the target object is a single-layer license plate or a double-layer license plate, if it is detected through the corresponding mask that the threshold value of the first region mask is 0, it can be judged that the target object is a single-layer license plate; if it is detected through the corresponding mask that the threshold value of the first region mask is 1, it can be judged that the target object is a double-layer license plate. Therefore, according to the threshold value result, it can be supervised whether the structural attributes extracted by the extraction model are correct, thereby on the one hand supervising the final output result of the extraction model, and on the other hand training the extraction model to make the extraction model more accurate in extracting information during the process of extracting the attributes of the target object.

[0111] As can be seen from the above embodiments, the object attribute extraction method provided by the present invention utilizes the correlation between the content attributes and structural attributes of the target object to make the extracted attribute information of the target object more accurate. In the prior art, to extract attributes from target objects with single-layer or double-layer structures, it is necessary to first classify the objects as single-layer or double-layer, and then use algorithmic models corresponding to different structures to perform content recognition. This method extracts the content attributes and structural attributes of the target object separately. This method does not fully utilize the correlation between the content attributes and structural attributes of the target object, resulting in relatively high data processing complexity and a waste of storage resources.

[0112] On the other hand, the structural attribute information extracted by the extraction model for the target object can promote the content information extracted by the extraction model to be more accurate. Because when the extraction model extracts the structural attribute of the target object, different regions corresponding to the target object will have different domain values. The structural attribute and domain value of the target object correspond to each other. That is, when the extraction model extracts that the structural attribute of the target object is single-layer, the domain value of the first region of the target object is 0, and the domain value of the second region is 1. When the extraction model extracts that the structural attribute of the target object is double-layer, the domain value of the first region of the target object is 1, and the domain value of the second region is 1. Therefore, based on the structural attribute information of the target object, it is possible to supervise whether the domain value determined by extracting the mask is correct, and the correctness of the domain value can promote the final content attribute of the target object to be more accurate.

[0113] As can be seen from the above embodiments, the object attribute extraction method provided by the present invention utilizes the correlation between the content attributes and structural attributes of the target object to make the extracted content information of the target object more accurate. In contrast, in the prior art, when extracting attributes from target objects with single-layer or double-layer structures, the objects are first classified as single-layer or double-layer, and then content recognition is performed using algorithmic models corresponding to the different structures. This method extracts the content attributes and structural attributes of the target object separately. This method does not fully utilize the correlation between the content attributes and structural attributes of the target object, resulting in relatively high data processing complexity and a waste of storage resources.

[0114] In short, it can be seen from the above embodiments that the object attribute extraction method provided by the present invention utilizes the correlation between the content attributes and the structure attributes of the target object to make the extracted attribute information of the target object more accurate.

[0115] In this embodiment, a device for extracting object attributes is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments. The details already described will not be repeated here. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0116] Figure 4 FIG. 1 is a structural block diagram of an apparatus for extracting object attributes according to an embodiment of the present invention. Figure 4 As shown, the device includes:

[0117] An acquisition module 10 is used to acquire initial features of a target object;

[0118] An extraction module 20, configured to extract multiple attributes of the target object from the initial features of the target object, wherein when the structural attributes of the target object are different, the extraction algorithm of the extraction model is unique;

[0119] A fusion module 30 is configured to fuse the multiple attributes of the target object to obtain an overall attribute of the target object;

[0120] The acquisition module further includes a target object acquisition module 11 and a coding acquisition module 12 , wherein the target object acquisition module 11 is used to acquire the target object, and the coding acquisition module 12 is used to acquire the initial features of the target object.

[0121] The initial features of the target object are encoded data obtained after the target object is subjected to a convolution operation. Preferably, the encoded data is original encoded data containing all attributes of the target object. Based on the encoded data, a plurality of specific attributes of the target object, such as content attributes, structural attributes, and color attributes, are obtained by an extraction module. Here, the attributes of the target object include but are not limited to content attributes, structural attributes, and color attributes.

[0122] According to an embodiment of the present invention, the extraction module 20 is configured to extract the structural attributes, content attributes and color attributes of the target object using an attention mechanism.

[0123] Preferably, the structural attributes, content attributes, and color attributes of the target object are extracted simultaneously by the extraction module. It is understandable that, in some embodiments, the structural attributes, content attributes, and color attributes of the target object may also be extracted in different orders.

[0124] According to one embodiment of the present invention, the extraction module 20 further includes a region division module 21, wherein the region division module 21 is configured to automatically divide each target object into a first region and a second region using the attention mechanism. The first region and the second region may be located at different positions of the target object. The algorithm model of the attention mechanism may be preset, and the preset algorithm model corresponds to the region division of the target object.

[0125] According to one embodiment of the present invention, the extraction module also includes a region assignment module 22, wherein the region assignment module 22 is configured to assign values ​​to the first region and the second region of the target object based on the initial features of the target object using a convolutional neural network (CNN) to obtain a first domain value and a second domain value.

[0126] According to one embodiment of the present invention, the region assignment module 22 includes a mask extraction unit 221 and a region assignment unit 222, wherein the mask extraction unit 221 is configured to use a convolutional neural network to extract the mask of the target object, and the region assignment unit 222 is configured to use a convolutional neural network to determine whether the mask contains target content information, and assign values ​​to the first region and the second region.

[0127] Because the initial features of the target object are original coded data containing all attributes of the target object, the mask extraction unit 221 can extract the mask of the target object based on the initial features, i.e., the original coded data. Therefore, when the target object is divided into a first region and a second region by the region division module 21, the mask extraction unit 221 extracts a first region mask for the first region, and extracts a second region mask for the second region.

[0128] The region assignment unit 222 of the region assignment module 22 further includes a detection module 2221. The detection module 2221 is configured to use a convolutional neural network (CNN) to detect whether the target object mask contains target information to obtain a domain value for the corresponding region, wherein the domain value is used to confirm the content attribute of the target object. When no target information is detected in the target object mask, the region corresponding to the mask is assigned a value of 0; when target information is detected in the target object mask, the region corresponding to the mask is assigned a value of 1. That is, when no target information is detected in the first region mask or the second region mask of the target object, the first region or the second region is assigned a value of 0, i.e., the first domain value or the second domain value is 0; when target information is detected in the first region mask or the second region mask of the target object, the first region or the second region is assigned a value of 1, i.e., the first domain value or the second domain value is 1.

[0129] According to a preferred embodiment of the present invention, the extraction module 20 further includes a content processing module 23, wherein the content processing module includes a calculation unit 231, a splicing unit 232 and a parsing unit 233, wherein the calculation unit 231 is used to respectively multiply the first region content data by the first domain value and the second region content data by the second domain value to obtain the first region content representation and the second region content representation respectively, the splicing unit 232 is used to splice the first region content representation and the second region content representation to obtain the overall content representation of the target object, and the parsing unit 233 is used to parse the overall content representation of the target object to obtain the content attributes of the target object.

[0130] According to a preferred embodiment of the present invention, the splicing unit 232 may splice the first region content representation and the second region content representation in a horizontal manner or other splicing manners, which are not limited herein. The splicing manner may be preset in the extraction module 20.

[0131] The parsing unit 233 is further configured to parse the color content representation and the structure content representation into the structure attribute and the color attribute of the target object, respectively.

[0132] According to an embodiment of the present invention, the fusion module 30 fuses the structural attributes, the color attributes, and the content attributes of the target object to obtain the overall attributes of the target object.

[0133] As can be seen from the above embodiments, the object attribute extraction device provided by the present invention utilizes the correlation between the content attributes and structural attributes of the target object to make the extracted content information of the target object more accurate. In the prior art, when extracting attributes from target objects with single-layer or double-layer structures, the objects are first classified as single-layer or double-layer, and then content recognition is performed using algorithmic models corresponding to the different structures. This method extracts the content attributes and structural attributes of the target object separately. This method does not fully utilize the correlation between the content attributes and structural attributes of the target object, resulting in relatively high data processing complexity and a waste of storage resources.

[0134] Through the description of the above embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0135] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0136] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.

[0137] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0138] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0139] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0140] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0141] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0142] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for extracting object attributes, characterized in that: Applicable to vehicle license plate information extraction scenarios, including: Acquire initial features of a target object, wherein the target object is a license plate image; Extracting the structural attributes, content attributes, and color attributes of the target object from the initial features of the target object using an extraction model, including: dividing the target object into multiple regions using an attention mechanism; assigning values ​​to the multiple regions to obtain corresponding domain values, wherein the domain values ​​are used to confirm the content attributes of the target object; multiplying the content data of each region by the corresponding domain value to obtain a content representation of the corresponding region; concatenating the content representations of each region to obtain an overall content representation of the target object; parsing the overall content representation of the target object to obtain the content attributes of the target object, wherein, when the target object is a single-layer structure or a double-layer structure, an algorithm model preset in the extraction model divides the target object into a first region and a second region, extracting a first region mask and a second region mask of the target object, so that when the structural attributes of the target object are different, the algorithm of the extraction model is unique, and the structural attributes are used to indicate whether the target object is a single-layer structure or a double-layer structure; The structural attribute, the content attribute, and the color attribute of the target object are fused to obtain an overall attribute of the target object, and the type information of the vehicle is determined according to the overall attribute.

2. The method according to claim 1, characterized in that Extracting the structural attribute, content attribute, and color attribute of the target object from the initial features of the target object using the extraction model includes at least one of the following: extracting the structural attributes of the target object from the initial features of the target object using the extraction model; extracting the content attributes of the target object from the initial features of the target object using the extraction model; The color attribute of the target object is extracted from the initial features of the target object using the extraction model.

3. The method according to claim 1, characterized in that When target content information does not exist in the area, the content representation obtained by multiplying the content data of the area by the corresponding domain value is used as the first threshold; when target content information exists in the area, the content representation obtained by multiplying the content data of the area by the corresponding domain value is still the content data of the area.

4. An object attribute extraction device, characterized in that: Applicable to vehicle license plate information extraction scenarios, including: An acquisition module, configured to acquire initial features of a target object, wherein the target object is a license plate image; An extraction module is configured to extract the structural attributes, content attributes, and color attributes of the target object from the initial features of the target object, including: dividing the target object into multiple regions using an attention mechanism; assigning values ​​to the multiple regions to obtain corresponding domain values, wherein the domain values ​​are used to confirm the content attributes of the target object; multiplying the content data of each region by the corresponding domain value to obtain a content representation of the corresponding region; concatenating the content representations of each region to obtain an overall content representation of the target object; and parsing the overall content representation of the target object to obtain the content attributes of the target object. When the target object has a single-layer structure or a double-layer structure, an algorithm model preset in the extraction module divides the target object into a first region and a second region, extracting a first region mask and a second region mask of the target object, so that when the structural attributes of the target object are different, the algorithm of the extraction module is unique, and the structural attributes are used to indicate whether the target object has a single-layer structure or a double-layer structure. A fusion module is used to fuse the structural attributes, the content attributes and the color attributes of the target object to obtain the overall attributes of the target object, and determine the type information of the vehicle according to the overall attributes.

5. The device according to claim 4, characterized in that The extraction module also includes: a region division module, wherein the region division module is configured to divide each target object into a plurality of regions using an attention mechanism; A region assignment module, wherein the region assignment module is configured to assign values ​​to the regions of the target object based on the initial features of the target object using a convolutional neural network to obtain corresponding domain values, and to assign values ​​to the multiple regions to obtain corresponding domain values, wherein the domain values ​​are used to confirm the content attributes of the target object.

6. The device according to claim 5, characterized in that The extraction module further includes a content processing module, wherein the content processing module includes: a calculation unit, wherein the calculation unit is configured to multiply the content data of the region by the domain value to obtain a content representation of the region; a splicing unit, configured to splice the content representations of all the regions to obtain an overall content representation of the target object; and A parsing unit is configured to parse the overall content representation of the target object to obtain the content attributes of the target object.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 3 when executed.

8. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • License plate recognition method and training method and device of license plate recognition model

    CN111832568A

  • Trajectory data processing method, device and storage medium

    WO2021017675A1