Image processing method, device, system, storage medium and computer equipment
By superimposing a mask signal on the license plate image and combining it with a text recognition model, the problem of low license plate recognition accuracy is solved, and efficient and accurate license plate recognition is achieved. It is suitable for the recognition of various license plate types, improving recognition efficiency and user experience.
Patent Information
- Application Number
- CN202110304027.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-22
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-03-22
AI Technical Summary
The existing technology has low license plate recognition accuracy, making it difficult to effectively improve license plate recognition capabilities and efficiency.
By superimposing a mask signal on the license plate image, a signal generator is used to generate a mask signal and superimpose it on the license plate image, and recognition is performed in combination with a text recognition model to improve the distinctiveness of the license plate text content, and the trained text recognition model is used for recognition.
It achieves improved accuracy of license plate recognition, improves efficiency and accuracy of license plate recognition, is applicable to the recognition of single-layer and double-layer license plates, and enhances the versatility of the recognition method and user experience.
Smart Images

Figure CN115116043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of images, and in particular to an image processing method, apparatus, system, storage medium and computer equipment. Background Art
[0002] In the city brain, the ability to recognize various types of license plates (motor vehicles and non-motor vehicles) is an extremely important and basic capability, because the license plate is the unique identity of various types of motor vehicles and non-motor vehicles. The accuracy and speed of its recognition are related to many applications, such as vehicle trajectory query and mining, vehicle search and comparison, etc.
[0003] In related technologies, there are already many types of license plates. For example, double-layer license plates of various specifications are an extremely important category. Moreover, license plate recognition also has a variety of application scenarios in various applications. For example, the recognition of double-layer license plates is also a common demand scenario, such as the rear license plates of various large vehicles and various trailer license plates, which are all key focuses in application and implementation. Therefore, how to improve license plate recognition capabilities and efficiency is currently a problem that needs to be solved.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present invention provide an image processing method, apparatus, system, storage medium and computer equipment to at least solve the technical problem of low license plate recognition accuracy in the related art.
[0006] According to one aspect of an embodiment of the present invention, there is provided an image processing method, comprising: acquiring a first license plate image; generating a mask signal using a signal generator; superimposing the mask signal on the first license plate image to obtain a second license plate image; and recognizing the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image.
[0007] Optionally, when the mask signal includes a first mask signal and a second mask signal, superimposing the mask signal on the first license plate image to obtain the second license plate image includes: superimposing the first mask signal on the first image area of the first license plate image; superimposing the second mask signal on the second image area of the first license plate image; wherein the first license plate image includes the first image area and the second image area.
[0008] Optionally, the first image area includes an upper half of the first license plate image, and the second image area includes a lower half of the first license plate image.
[0009] Optionally, the first image area includes an upper third of the first license plate image, and the second image area includes a lower two-thirds of the first license plate image.
[0010] Optionally, the first mask signal includes a sine mask signal, and the second mask signal includes a cosine mask signal; or the first mask signal includes a cosine mask signal, and the second mask signal includes a sine mask signal.
[0011] Optionally, an original license plate image is obtained; and text region detection is performed on the original license plate image to obtain the first license plate image including the text region.
[0012] Optionally, before using the text recognition model to identify the second license plate image and obtaining the text recognition result on the first license plate image, it also includes: obtaining a sample training set, wherein the training data in the sample training set includes: a license plate image and a license plate in the license plate image, and the license plate image is an image obtained after superimposing a mask signal; using the sample training set to train the initial model to obtain the text recognition model.
[0013] Optionally, the first license plate image includes: an image of a double-layer license plate.
[0014] Optionally, the text recognition model includes at least one of the following: a text recognition model based on a convolutional recurrent neural network, and a text recognition model based on an attention mechanism.
[0015] Optionally, the vehicle corresponding to the license plate marked by the text recognition result is located; the speed of the vehicle corresponding to the license plate marked by the text recognition result is measured; and the vehicle corresponding to the license plate marked by the text recognition result is monitored.
[0016] According to another aspect of an embodiment of the present invention, an image processing method is also provided, including: receiving a first license plate image in an input area of a display interface; receiving a license plate recognition instruction on the display interface; and displaying a text recognition result of the first license plate image on the display interface in response to the license plate recognition instruction, wherein the text recognition result is obtained by recognizing a second license plate image using a text recognition model, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image.
[0017] According to another aspect of an embodiment of the present invention, an image processing method is also provided, including: a client device sends a first license plate image to a server device; the server device uses a signal generator to generate a mask signal; the server device superimposes the mask signal on the first license plate image to obtain a second license plate image; and uses a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image; the client device receives the text recognition result returned by the server device.
[0018] According to one aspect of an embodiment of the present invention, an image processing device is provided, comprising: a first acquisition module for acquiring a first license plate image; a first generation module for generating a mask signal using a signal generator; an overlay module for superimposing the mask signal on the first license plate image to obtain a second license plate image; and a first recognition module for recognizing the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image.
[0019] According to another aspect of an embodiment of the present invention, an image processing device is also provided, including: a first receiving module for receiving a first license plate image in an input area of a display interface; a second receiving module for receiving a license plate recognition instruction on the display interface; a display module for responding to the license plate recognition instruction and displaying a text recognition result of the first license plate image on the display interface, wherein the text recognition result is obtained by recognizing a second license plate image using a text recognition model, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image.
[0020] According to another aspect of an embodiment of the present invention, an image processing system is also provided, including a client device and a server device, wherein the client device is used to send a first license plate image to the server device; the server device is used to generate a mask signal using a signal generator; the server device is also used to superimpose the mask signal on the first license plate image to obtain a second license plate image; and use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image; the client device is also used to receive the text recognition result returned by the server device.
[0021] According to one aspect of an embodiment of the present invention, there is provided an image processing method, comprising: acquiring a first license plate image of a vehicle; using a text recognition model to recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; acquiring a load of the vehicle based on the text recognition result; and issuing a freight instruction for the vehicle based on the load.
[0022] According to another aspect of an embodiment of the present invention, an image processing method is provided, comprising: acquiring a first license plate image of a vehicle and audio data of the vehicle passing by; using a text recognition model to recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and determining the driving status of the vehicle based on the text recognition result and the audio data.
[0023] According to another aspect of an embodiment of the present invention, there is provided an image processing method, comprising: acquiring a first license plate image of a vehicle and positioning data of the vehicle; using a text recognition model to recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and generating a navigation route for the vehicle based on the text recognition result and the positioning data.
[0024] According to one aspect of an embodiment of the present invention, there is provided an image processing device, comprising: a second acquisition module for acquiring a first license plate image of a vehicle; a second recognition module for recognizing the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; a third acquisition module for acquiring the load of the vehicle based on the text recognition result; and an issuing module for issuing a freight instruction for the vehicle based on the load.
[0025] According to another aspect of an embodiment of the present invention, an image processing device is provided, including: a fourth acquisition module, used to acquire a first license plate image of a vehicle and audio data when the vehicle passes by; a third recognition module, used to use a text recognition model to recognize the second license plate image, and obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and a determination module, used to determine the driving status of the vehicle based on the text recognition result and the audio data.
[0026] According to another aspect of an embodiment of the present invention, an image processing device is provided, including: a fifth acquisition module, used to acquire a first license plate image of a vehicle and positioning data of the vehicle; a fourth recognition module, used to use a text recognition model to recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and a second generation module, used to generate a navigation route for the vehicle based on the text recognition result and the positioning data.
[0027] In an embodiment of the present invention, a mask signal is superimposed on the license plate image. Since the mask signal is superimposed on the license plate image, the text content on the license plate image can be more distinguished from other content. When the trained text recognition model is used for recognition, the purpose of improving recognition accuracy is achieved, and the technical effect of efficiently recognizing license plates is achieved, thereby solving the technical problem of low license plate recognition accuracy in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0029] Figure 1 is a block diagram of the hardware structure of a computer terminal for implementing an image processing method according to an embodiment of the present invention;
[0030] Figure 2 is a flowchart of an image processing method 1 according to embodiment 1 of the present invention;
[0031] Figure 3 is a flowchart of a second image processing method according to embodiment 1 of the present invention;
[0032] Figure 4 is a flowchart of the third image processing method according to embodiment 1 of the present invention;
[0033] Figure 5 is a flowchart of an image processing method 4 according to embodiment 1 of the present invention;
[0034] Figure 6 is a flowchart of the fifth image processing method according to embodiment 1 of the present invention;
[0035] Figure 7 is a flowchart of an image processing method 6 according to embodiment 1 of the present invention;
[0036] Figure 8 is a schematic diagram of a double-layer license plate recognition method based on single- and double-layer classification according to Example 1 of the present invention;
[0037] Figure 9 Schematic diagram of misclassification of a double-layer license plate according to embodiment 1 of the present invention;
[0038] Figure 10 is a schematic diagram of a double-layer license plate recognition method based on feature layer splicing according to Example 1 of the present invention;
[0039] Figure 11 is a schematic diagram of a double-layer license plate recognition method based on signal mask according to an optional embodiment of the present invention;
[0040] Figure 12 is a structural block diagram of an image processing device according to embodiment 2 of the present invention;
[0041] Figure 13 is a structural block diagram of an image processing device 2 according to embodiment 2 of the present invention;
[0042] Figure 14 is a structural block diagram of an image processing device 3 according to embodiment 2 of the present invention;
[0043] Figure 15 is a structural block diagram of an image processing apparatus according to Embodiment 2 of the present invention;
[0044] Figure 16 is a structural block diagram of an image processing device 5 according to embodiment 2 of the present invention;
[0045] Figure 17 is a structural block diagram of an image processing apparatus 6 according to embodiment 2 of the present invention;
[0046] Figure 18 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0048] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0049] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0050] Optical Character Recognition (OCR) refers to the process by which an electronic device (such as a scanner or digital camera) examines characters printed on paper, determines their shape by detecting dark and light patterns, and then uses character recognition methods to translate the shape into computer text. In other words, for printed characters, the text in a paper document is optically converted into a black and white dot matrix image file, and recognition software is used to convert the text in the image into text format for further editing and processing by word processing software.
[0051] Convolutional Neural Networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. Convolutional neural networks possess representation learning capabilities and can perform shift-invariant classification of input information based on their hierarchical structure. Hence, they are also known as shift-invariant artificial neural networks (SIANNs).
[0052] A recurrent neural network (RNN) is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and has all nodes (recurrent units) connected in a chain-like manner.
[0053] Long Short-Term Memory (LSTM) neural networks are a type of recurrent neural network designed to address the long-term dependency issues inherent in typical RNNs (recurrent neural networks). All RNNs have a chain-like structure of repeating neural network modules. In a standard RNN, this repeating module consists of a very simple structure, such as a single tanh layer.
[0054] In a fully connected network (FC), each node is connected to all nodes in the previous layer, integrating previously extracted features. Due to its fully connected nature, FCs typically have the most parameters.
[0055] Encoder-Decoder Network, Encoder-Decoder Network, Encoding is the process of converting information from one form or format to another. Decoding is the inverse process of encoding.
[0056] The Connectionist Temporal Classification (CTC) loss function can automatically align unaligned data and is primarily used for training sequential data that has not been pre-aligned, such as speech recognition and optical character recognition.
[0057] The attention mechanism is a resource allocation scheme that is the primary means of addressing information overload when computing power is limited. It allocates computing resources to more important tasks. Attention is generally divided into two types: one is top-down, conscious attention, called focused attention. Focused attention is attention that is intentional, task-dependent, and actively and consciously focused on a specific object. The other is bottom-up, unconscious attention, called saliency-based attention. Saliency-based attention is attention driven by external stimuli, does not require active intervention, and is unrelated to the task.
[0058] Example 1
[0059] According to an embodiment of the present invention, an image processing method embodiment is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0060] The method embodiment provided in Example 1 of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (shown as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0061] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0062] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image processing method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the image processing method of the application described above. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0063] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0064] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0065] Under the above operating environment, this application provides Figure 2 The image processing method shown. Figure 2 is a flow chart of an image processing method 1 according to embodiment 1 of the present invention. Figure 2 As shown, the method includes the following steps:
[0066] Step S202, obtaining a first license plate image;
[0067] Step S204, using a signal generator to generate a mask signal;
[0068] Step S206, superimposing the mask signal on the first license plate image to obtain a second license plate image;
[0069] Step S208: Using a text recognition model, recognize the second license plate image to obtain a text recognition result on the first license plate image.
[0070] In an embodiment of the present invention, a mask signal is superimposed on the license plate image. Since the mask signal is superimposed on the license plate image, the text content on the license plate image can be more distinguished from other content. When the trained text recognition model is used for recognition, the purpose of improving recognition accuracy is achieved, and the technical effect of efficiently recognizing license plates is achieved, thereby solving the technical problem of low license plate recognition accuracy in related technologies.
[0071] As an optional embodiment, the first license plate image can be an image of various types of license plates, and can be an image of a single-layer license plate or an image of a double-layer license plate. When performing double-layer license plate recognition, images of double-layer license plates of various different types of vehicles can be recognized, such as: large cars (trucks, semi-trailer tractors, trams), trailers, motorcycles, large and medium-sized buses, and so on. Therefore, this optional embodiment can be applied not only to the recognition of single-layer license plates, but also to double-layer or even multi-layer license plates, or other license plates of different designs, effectively solving the current technical problem of low accuracy in the recognition of double-layer or even multi-layer license plates. That is, this optional embodiment does not have any license plate type limitation requirements, effectively realizing the versatility of the license plate recognition method, improving the convenience of users using the method in terms of consistency, and enhancing the user experience.
[0072] As an optional embodiment, a first license plate image can be acquired in a variety of ways. For example, the original license plate image can be first acquired, and text region detection can be performed on the original license plate image to obtain a first license plate image including a text region. The above processing ensures that the acquired first license plate image includes a text region, effectively avoiding image misrecognition (e.g., recognizing an image without a license plate). Furthermore, it can ensure, to a certain extent, that subsequent recognition is performed on the text region, while other regions without text (i.e., regions not containing the license plate) can be omitted, effectively improving recognition efficiency. Furthermore, a variety of implementation methods can be used to acquire the original license plate image, such as road monitoring photography, high-speed speed measurement photography, camera photography, etc. After acquiring the original license plate image and before performing text region detection on the license plate image, the acquired original license plate image can also be preprocessed. For example, this can include at least one of the following: adjusting the image hue, such as brightness and contrast, etc.; adjusting the image size, such as zooming in and out, adjusting the length, width, and height, etc. For example, when monitoring vehicle speed on a highway, an image of the front of a large car is captured. In addition to the license plate image, the image also includes the front of the car and the environmental background image. At this time, it is necessary to perform text area detection on the original license plate image and extract the first license plate image in the text area; if the weather is cloudy, the original license plate image taken is darker, and the image brightness can be increased to better detect and identify the license plate image; if the license plate is located at a low position on the front of the car, the image taken by the high camera will be distorted, and the problem of low license plate height may occur. At this time, the height of the original license plate image can be adjusted, and so on.
[0073] As an optional embodiment, when superimposing a mask signal on the first license plate image to obtain the second license plate image, different superimposition methods correspond to different mask signal representations. For example, when the first license plate image is divided into a first image region and a second image region, the first image region corresponds to the first mask signal, and the second image region corresponds to the second mask signal. That is, when the mask signal includes the first mask signal and the second mask signal, superimposing the mask signal on the first license plate image to obtain the second license plate image can be performed as follows: superimposing the first mask signal on the first image region of the first license plate image; and superimposing the second mask signal on the second image region of the first license plate image; wherein the first license plate image includes the first image region and the second image region. It should be noted that dividing the first license plate image into two image regions: the first image region and the second image region, is merely an example of this optional embodiment. For example, the first license plate image can also be divided into more regions, such as three, four, and so on. When the first license plate image is divided into more regions, mask signals need to be generated for each of these more regions.
[0074] As an optional embodiment, when the first license plate image is divided into two image areas, there can be a variety of division methods, for example, the method of evenly dividing the upper and lower areas is adopted, the first image area includes the upper half of the first license plate image, and the second image area includes the lower half of the first license plate image. For another example, the method of unevenly dividing the upper and lower areas can be adopted, the first image area includes the upper third of the first license plate image, and the second image area includes the lower two-thirds of the first license plate image. For example, for a double-layer license plate, the upper half is not the upper license plate and the lower half is not the lower license plate. For the rear license plate of a large vehicle specified in the national standard for license plates, the rear license plate (size 440 length and 220 width), the size range of the upper half area is 15-75, the text size of the lower half area is 90-210, the width of the upper half is one-third of the overall width, and the width of the lower half is two-thirds of the overall width. Avoid dividing the upper and lower areas equally during the license plate recognition process, which results in that although the text recognition of the lower area is basically reasonable in the text recognition model, the text of the upper area may often be ignored, resulting in recognition errors. A variety of methods are adopted to divide the image area to be recognized, which improves the accuracy of the text recognition model and thus improves the accuracy of double-layer license plate recognition. It should be pointed out that the above-mentioned upper and lower equal division and non-equal division methods are both examples. If there are other ways to design the license plate in the future, then the method of the present application is used to superimpose mask signals on each divided area to recognize the license plate, which can achieve the effect of accurate and efficient license plate recognition and is part of the present application. As an optional embodiment, superimposing a mask signal on the first license plate image is used to make the text part of the license plate more clearly distinguishable from the non-text part. For example, when a mask signal is superimposed, it performs no operation where text is present, while casting a shadow where text is absent. This allows the license plate image with the mask signal to more clearly identify areas with text and more clearly and accurately identify the specific content of the text. Furthermore, when the first license plate image is divided into different image regions, different mask signals can be superimposed on each image region to distinguish the different image regions. For example, the first mask signal can include a sine mask signal, and the second mask signal can include a cosine mask signal; or, the first mask signal can include a cosine mask signal, and the second mask signal can include a sine mask signal. It can be seen that the first mask signal and the second mask signal are each a sine mask signal and a cosine mask signal. This is merely a specific implementation. The signal generator is abstractly denoted by T, where I represents the i-th image. The signal generator for the upper half is denoted by Tup, and the signal generator for the lower half is denoted by Tbottom. The purpose of the signal generator is to generate a mask signal that is superimposed on the input image through another signal channel. Its abstract form is as follows:
[0075]
[0076] By superimposing different mask signals on different regions of the double-layer license plate image, different mask signals can identify the license plate text information in their respective regions, avoiding feature confusion, text corruption, or omission. Furthermore, the superimposed image only has one additional channel of sine and cosine signals compared to the original image, resulting in a negligible increase in computational effort, resulting in no loss in license plate recognition and actually an improvement.
[0077] As an optional embodiment, before using a text recognition model to recognize the second license plate image and obtaining the text recognition result on the first license plate image, the method further includes: obtaining a sample training set, wherein the training data in the sample training set includes: a license plate image and the license plate in the license plate image, wherein the license plate image is an image obtained by superimposing a mask signal; and using the sample training set to train the initial model to obtain a text recognition model. The initial model is trained using the license plate image with the superimposed mask signal to obtain a text recognition model. Since the training samples are all masked, the subsequent use of the trained text recognition model to recognize the license plate image can specifically recognize the license plate image, effectively avoiding problems such as mismatched, inaccurate, and missed recognition. This allows the license plate area in the license plate image to be better learned, and the distinctive features including the text are learned. With continuous sample training, the recognition accuracy is gradually improved.
[0078] As an optional embodiment, the text recognition model may include multiple types, for example, it may include at least one of the following: a text recognition model based on a convolutional recurrent neural network, a text recognition model based on an attention mechanism. It should be noted that the above-mentioned text recognition model based on a convolutional recurrent neural network and the text recognition model based on an attention mechanism are merely examples, and other text recognition models not listed one by one may also be applied to this application. The text recognition model based on a convolutional recurrent neural network is similar to the text recognition model based on the attention mechanism, and is trained to recognize license plates. The text recognition model can be based on different mechanisms and can be selected according to different needs, providing a variety of methods to choose from, making it more flexible and convenient to use, and greatly improving the applicability of license plate recognition. In addition, in the text recognition model based on a convolutional recurrent neural network, the model layer of a bidirectional recurrent neural network can be selected based on the trade-off between recognition accuracy and recognition efficiency. Therefore, it can be applied to more scenarios and double-layer license plate recognition with different needs.
[0079] As an optional embodiment, the result of the text recognition marking can also be used for at least one of the following: locating the vehicle corresponding to the license plate marked by the text recognition result; measuring the speed of the vehicle corresponding to the license plate marked by the text recognition result; monitoring the vehicle corresponding to the license plate marked by the text recognition result. The license plate marked by the text recognition result can be used in a variety of scenarios, such as: City OCR (text detection and recognition in street scenes), etc. Moreover, it can be used for a variety of purposes, such as: vehicle trajectory query and mining, vehicle speeding measurement, vehicle search and comparison, vehicle violation records, etc. The license plate is the unique identity of various types of motor vehicles and non-motor vehicles. The license plate recognition capability is an extremely important and basic capability. The accuracy of license plate recognition is related to many subsequent applications. Marking the license plate through text recognition can enhance the functional application in different scenarios and can achieve a variety of purposes more flexibly and accurately.
[0080] Figure 3 is a flow chart of the second image processing method according to embodiment 1 of the present invention. Figure 3 As shown, the method includes the following steps:
[0081] Step S302, receiving a first license plate image in an input area of the display interface;
[0082] Step S304, receiving a license plate recognition instruction on the display interface;
[0083] Step S306, responding to the license plate recognition instruction, displaying the text recognition result of the first license plate image on the display interface, wherein the text recognition result is obtained by recognizing the second license plate image using the text recognition model, and the second license plate image is an image obtained by superimposing the mask signal on the first license plate image.
[0084] Through the above steps, a display interface is used to receive a license plate image, and the text recognition result of the license plate image is displayed on the display interface. A mask signal is superimposed on the license plate image. Since the mask signal is superimposed on the license plate image, the text content on the license plate image can be more distinguished from other content. When the trained text recognition model is used for recognition, the purpose of improving recognition accuracy and facilitating human-computer interaction is achieved, and the technical effect of efficiently recognizing license plates is achieved, thereby solving the technical problem of low license plate recognition accuracy in related technologies.
[0085] Figure 4 is a flowchart of the image processing method 3 according to embodiment 1 of the present invention. Figure 4 As shown, the method includes the following steps:
[0086] Step S402: The client device sends the first license plate image to the server device;
[0087] Step S404: the server device generates a mask signal using a signal generator;
[0088] In step S406, the server device superimposes the mask signal on the first license plate image to obtain a second license plate image; and uses a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image;
[0089] Step S408: The client device receives the text recognition result returned by the server device.
[0090] Through the above steps, the client device sends the license plate image, the server device receives the image for image recognition, and then returns the recognition result to the client device. The method of superimposing a mask signal on the license plate image is adopted. Since the mask signal is superimposed on the license plate image, the text content on the license plate image can be more distinguished from other content. When the trained text recognition model is used for recognition, the purpose of improving recognition accuracy and software as a service (SaaS) is achieved, and the technical effect of efficiently recognizing license plates is achieved, thereby solving the technical problem of low license plate recognition accuracy in related technologies.
[0091] Figure 5 : is a flowchart of the image processing method 4 according to embodiment 1 of the present invention. Figure 5 As shown, the method includes the following steps:
[0092] Step S502, obtaining a first license plate image of the vehicle;
[0093] Step S504: Using a text recognition model, recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image;
[0094] Step S506, obtaining the vehicle's load capacity based on the text recognition result;
[0095] Step S508: issuing a freight instruction to the vehicle based on the load.
[0096] Through the above steps, a mask signal is superimposed on the license plate image. This allows the text content within the license plate image to be more distinctly identified than other content, thereby improving recognition accuracy and achieving the technical effect of efficient license plate recognition. Furthermore, based on the text recognition results and vehicle weight, freight instructions corresponding to the vehicle can be determined and issued, thereby improving freight efficiency.
[0097] As an optional embodiment, various methods can be used to obtain a vehicle's load capacity based on text recognition results. For example, license plates generally classify vehicles by type. A vehicle type can be a type of vehicle, distinguished by common features, intended use, and function. For example, sedans, trucks, buses, trailers, incomplete vehicles, and motorcycles are all separate types. In addition to the above classifications, vehicle types can also be further classified, for example, by load capacity, such as 50 tons for Class A and 30 tons for Class B. Regardless of the classification method, each type may correspond to a load capacity. Load capacity represents the maximum weight a vehicle can bear and is a parameter that characterizes its capabilities. Therefore, after obtaining the text recognition results for a vehicle, the vehicle type can be determined based on the text recognition results, and the vehicle's load capacity can be found based on the correspondence between license plate type and load capacity. It should be noted that the correspondence between license plate type and load capacity can be expressed in the form of a table or a formula.
[0098] As an optional embodiment, when issuing a freight instruction to a vehicle based on the load, the vehicle's current load capacity can be first obtained, and the vehicle's status can be detected based on the load capacity and load. Freight instructions can then be issued based on different statuses. For example, if the vehicle is detected to be empty based on the load capacity and load, the vehicle can be instructed to drive to the loading area to wait for loading; if the vehicle is detected to be underloaded based on the load capacity and load, the vehicle can be instructed to drive to the replenishment area to replenish; if the vehicle is detected to be overloaded based on the load capacity and load, the vehicle can be instructed to drive to the unloading area to replenish, and so on.
[0099] Figure 6 is a flowchart of the image processing method 5 according to embodiment 1 of the present invention. Figure 6 As shown, the method includes the following steps:
[0100] Step S602, obtaining a first license plate image of a vehicle and audio data of the vehicle passing by;
[0101] Step S604: Using a text recognition model, recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image;
[0102] Step S606: Determine the driving status of the vehicle based on the text recognition result and the audio data.
[0103] Through the above steps, a mask signal is superimposed on the license plate image. This allows the text content on the license plate image to be more distinctly identified than other content, improving recognition accuracy and achieving the technical effect of efficient license plate recognition. Furthermore, based on the text recognition results and audio data, the vehicle's driving status can be determined, thereby achieving the effect of vehicle monitoring.
[0104] As an optional embodiment, when determining the driving status of a vehicle based on text recognition results and audio data, the audio data can be obtained through a microphone array, wherein the microphone array can be deployed at a camera that obtains license plate images. For example, when a camera is installed on an overpass, a microphone array can be deployed at the camera, so that the license plate image can be obtained through the camera, and audio data of the vehicle passing by can be collected through the microphone array. The audio data of the vehicle passing by can reflect the speed of the vehicle to a certain extent. Therefore, the license plate of the vehicle can be determined based on the text recognition results, and the speed of the vehicle can be determined based on the audio data of the vehicle passing by. In this way, the speed of the vehicle with the license plate can be monitored, the speed of the vehicle can be measured, and a speeding reminder can be issued to the vehicle.
[0105] Figure 7 : is a flowchart of the image processing method 6 according to embodiment 1 of the present invention. Figure 7 As shown, the method includes the following steps:
[0106] Step S702, obtaining a first license plate image of the vehicle and positioning data of the vehicle;
[0107] Step S704: Using a text recognition model, recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image;
[0108] Step S706: Generate a navigation route for the vehicle based on the text recognition result and the positioning data.
[0109] Through the above steps, a mask signal is superimposed on the license plate image. This allows the text content on the license plate image to be more distinctly identified than other content, thereby improving recognition accuracy and achieving the technical effect of efficient license plate recognition. Furthermore, based on the text recognition results and positioning data, a navigation route is generated for the vehicle, thereby achieving vehicle navigation.
[0110] As an optional embodiment, when a navigation route is generated for a vehicle based on text recognition results and positioning data, the positioning data of the vehicle can be provided by a Global Positioning System (GPS). For example, after determining the license plate of the vehicle based on the text recognition results of the vehicle, the historical driving data of the vehicle is obtained based on the license plate. For the same starting point and end point, a navigation route in the historical driving data is selected, and the selected navigation route is recommended to the vehicle. With the above processing, when there are multiple navigation routes with the same starting point and end point, a familiar navigation route is selected for the user, and the user does not need to choose from multiple routes, thereby achieving the purpose of fast and efficient navigation for the user.
[0111] In the related art, the recognition of double-layer license plates can be done using the following two methods.
[0112] The first one is a two-layer license plate recognition method based on single and double-layer classification.
[0113] Figure 8 Schematic diagram of a double-layer license plate recognition method based on single and double-layer classification according to Example 1 of the present invention, Figure 8 As shown, the method includes the following processing: double-layer license plates are identified in advance (for example, a classification module for single-layer and double-layer license plates is added to the license plate recognition model). If the license plate recognition model finds that it is a double-layer license plate, the upper and lower layers are manually segmented before the license plate is input into the license plate recognition model for recognition. After segmentation, the upper and lower layer images are spliced together to form an image of a single-row license plate, which is then input into the license plate recognition model for license plate recognition. There are also methods that input the segmented images into the license plate recognition model twice for recognition, and perform string splicing on the final recognition results, but this method belongs to the same category as the former.
[0114] The problem with the first method is that it is affected by the accuracy of the single-layer and double-layer classification models. When the single-layer and double-layer classifications are wrong, the recognition will also be affected, such as classifying a double-layer license plate as a single-layer license plate. Figure 9 Schematic diagram of double-layer license plate misclassification according to embodiment 1 of the present invention. Figure 9 As shown, the model recognition will report an error. At the same time, the first method is a hard cut image, which will damage the text features of the image, such as Figure 8 As shown in , it will damage the lower half of the image and mislead the recognition model.
[0115] The second is a double-layer license plate recognition method based on feature layer splicing.
[0116] Figure 10 Schematic diagram of a double-layer license plate recognition method based on feature layer splicing according to Example 1 of the present invention, Figure 10As shown, this method performs "cutting and splicing" at the feature level, including the following processing: This method does not need to distinguish between single-layer and double-layer license plates, thus avoiding the recognition error problem caused by misclassification of single-layer and double-layer license plates. For the license plates obtained by license plate detection, whether single-layer or double-layer, they are directly input into the license plate recognition model after preprocessing such as size scaling. The image preprocessing stage here is the same as that of the ordinary text recognition model. The difference is that when the second method is scaled, the height of the image is usually 1.5 times or 2 times that of the previous single layer. For example, the height of the single-layer license plate recognition is usually H single (usually 32), then the height of the image setting in the second method is recorded as H multi (usually 48 or 64). The purpose of this setting is to ensure that the text information of the license plate can be retained as much as possible. Such input is feature encoded through the Encoder encoding network, and the output encoded features have a dimension that is 2 instead of 1, which is different from the dimension output by ordinary methods (for example, the first method mentioned above). This method model regards the features of the first layer as the representation features of the upper half of the license plate image (first layer), and the features of the second layer as the representation features of the lower half of the license plate image (second layer). Before inputting the encoded features into the decoding network for decoding, the method model splices the first layer features with the second layer features, and the resulting feature dimension is batch×1×(2*w′)×c, where batch represents the total number of input images, w′ represents the feature width, and c represents the number of feature representations. By splicing at the feature level, sequential decoding of the upper and lower layer image texts is achieved.
[0117] The problem with the second method is that it solves the problem of hard segmentation in the first method by soft feature segmentation and splicing, but this model also brings the problem of feature confusion. That is, for the input image I, after being resized to a fixed size: w×h (for example, 128×64 is actually used in the model), after passing through the feature encoding network, the output features are recorded as in, Represents the features of the upper half of the input image, Represents the features of the lower half of the input image. But in fact, due to the characteristics of the convolutional neural network model, The middle part does not only have the characteristics of the upper part. The lower half of the image isn't the only feature. Specifically, for large vehicle rear license plates (440mm long and 220mm wide) as specified in the national license plate standard, the upper half ranges in size from 15-75mm, while the text in the lower half ranges from 90-210mm. The width of the upper half is one-third of the overall width. This makes the model's predictions about the text in the lower half generally reasonable, but it may often overlook the text in the upper half, leading to recognition errors.
[0118] Both of the above methods cannot accurately identify the text in the image, and there are also problems with low classification accuracy and destruction of text area information. Because the above two methods are based on fixed distance to achieve image region splitting, naturally they cannot accurately identify the positions of different text layers in the image, and destroy the text area information.
[0119] In view of this, in this optional embodiment, based on the text recognition model (CRNN), an additional signal mask superposition module is introduced to address the problem of text area information destruction encountered in double-layer license plate recognition, so that the upper and lower license plate areas of the double-layer license plate can be better learned, and distinctive features can be learned to avoid feature confusion, thereby achieving improved recognition accuracy and further improving the performance of the double-layer license plate recognition model. The increased computational effort is almost negligible, and at the same time, there is no loss in the recognition of single-layer license plates and there is also improvement.
[0120] Figure 11 This is a schematic diagram of a double-layer license plate recognition method based on signal masking according to an optional embodiment of the present invention, taking the classic Convolutional Recurrent Neural Network (CRNN) text recognition model as an example (the same is true for the text recognition model based on the Attention Mechanism). Figure 11 As shown, the method includes:
[0121] S1, when the image is input, a signal mask is generated by the signal generator and superimposed on the input image;
[0122] Specifically, for the upper half of the image, the cosine function is used to generate the signal, and for the lower half of the image, the sine function is used to generate the signal. The expression for the generated signal is as follows:
[0123]
[0124] Among them, normalized width is the normalized image width. Note that the above generates a one-dimensional signal. For the upper half of the image (experienced as If h is set to 64, then the image 0-31 is considered to be the upper half area, and the image 32-63 is considered to be the lower half area. It can also be set to image height), this signal needs to be repeated times, and for the lower half, y bottom repeat The signal is completely superimposed on the input image. At the same time, after the superposition of the signal mask, the input is only expanded from a 3-channel image to a 4-channel image input, and the increase in computational complexity is very small and can be ignored.
[0125] S2, the text recognition model recognizes the input image.
[0126] (1) After the specific text area is proposed according to the text detection model, the text area is cut out, recorded as image I, and resized to a fixed size: w×h.
[0127] (2) Generate sine signal and cosine mask signal through signal generator respectively, and superimpose them on the upper half area and the lower half area of the image. The obtained image is still recorded as I. The image after superposition is as follows Figure 11 shown.
[0128] (3) Image I is passed through a convolutional neural network, called the backbone CNN, to perform feature encoding and obtain high-level image features. After the convolution and pooling layers of the CNN, the feature size is recorded as w′×2. The image height is reduced from h to the dimension of 2, and the image width is reduced from w to w′.
[0129] (4) The obtained image features are sliced and concatenated to form features with a width of (w′×2) and a height of 1.
[0130] (5) Because text is usually a sequence, the features extracted by CNN are then passed through a bidirectional cyclic neural network (Bi-LSTM) to extract higher-level sequence features, which are recorded as Its dimension is batch×2w′×c. (Note that this model layer is optional, depending on the trade-off between recognition accuracy and recognition efficiency)
[0131] (6) The above features obtained by the image encoding module are mapped through a fully connected layer to generate the words corresponding to each position in the width:
[0132] Y pred =Softmax(Wc)
[0133] Where W is the transformation matrix that maps the latent vector to the vocabulary. predThe dimension is batch×2w′×C num , C num To predict the number of characters in the dictionary. During the training phase, the CTC loss function is used for training. At the same time, batch processing is usually performed, and N images are sent to the network for training. At this time, the loss of this batch of images is calculated as follows:
[0134]
[0135] in, is the probability (logits) of the predicted output of the i-th image, The ground truth is the real label of the input image in the training phase. In the test phase, since there is no need to calculate the loss, it is directly passed to Y pred Get the predicted characters, and then connect the generated characters to form words, sentences, etc.
[0136] Through the above optional implementation, at least the following beneficial effects can be achieved:
[0137] 1. No need to classify single-layer and double-layer license plates, avoiding the problems of slow classification and low accuracy in manual classification;
[0138] 2. Feature resolution of different license plate layers and sample training to avoid feature confusion and text area information destruction;
[0139] 3. Double-layer license plate recognition accuracy is improved;
[0140] 4. In addition, it should be noted that the optimization strategy proposed in this optional implementation can be extended to perform text recognition on any multi-layer image.
[0141] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0143] Example 2
[0144] According to an embodiment of the present invention, a device for implementing the above-mentioned image processing method 1 is also provided. Figure 12 : is a structural block diagram of an image processing device according to embodiment 2 of the present invention. Figure 12 As shown, the device includes: a first acquisition module 1202, a first generation module 1204, a superposition module 1206 and a first identification module 1208. The device is described in detail below:
[0145] The first acquisition module 1202 is used to acquire a first license plate image; the first generation module 1204 is connected to the above-mentioned first acquisition module 1202 and is used to generate a mask signal using a signal generator; the superposition module 1206 is connected to the above-mentioned first generation module 1204 and is used to superimpose the mask signal on the first license plate image to obtain a second license plate image; the first recognition module 1208 is connected to the above-mentioned superposition module 1206 and is used to use a text recognition model to recognize the second license plate image and obtain a text recognition result on the first license plate image.
[0146] It should be noted that the first acquisition module 1202, the first generation module 1204, the superposition module 1206, and the first identification module 1208 described above correspond to steps S202 to S208 in Example 1. The examples and application scenarios implemented by these modules and the corresponding steps are the same, but are not limited to those disclosed in Example 1. It should be noted that the above modules, as part of the apparatus, can be run in the computer terminal 10 provided in Example 1.
[0147] According to an embodiment of the present invention, a device for implementing the above-mentioned second image processing method is also provided. Figure 13 is a structural block diagram of an image processing apparatus 2 according to embodiment 2 of the present invention. Figure 13 As shown, the device includes: a first receiving module 1302, a second receiving module 1304 and a display module 1306. The device is described in detail below:
[0148] The first receiving module 1302 is used to receive the first license plate image in the input area of the display interface; the second receiving module 1304 is connected to the above-mentioned first receiving module 1302, and is used to receive the license plate recognition instruction on the display interface; the display module 1306 is connected to the above-mentioned second receiving module 1304, and is used to respond to the license plate recognition instruction and display the text recognition result of the first license plate image on the display interface, wherein the text recognition result is obtained by using a text recognition model to recognize the second license plate image, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image.
[0149] It should be noted that the first receiving module 1302, the second receiving module 1304, and the display module 1306 correspond to steps S302 to S306 in Example 1. The examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0150] According to an embodiment of the present invention, a system for implementing the above-mentioned image processing method 3 is also provided. Figure 14 : is a structural block diagram of the image processing device 3 according to embodiment 2 of the present invention, as shown in FIG. Figure 14 As shown, the apparatus includes: a client device 1402 and a server device 1404. The apparatus is described in detail below:
[0151] The client device 1402 is used to send the first license plate image to the server device; it is also used to receive the text recognition result returned by the server device; the server device 1404 is connected to the above-mentioned client device 1402, and is used to generate a mask signal using a signal generator; it is also used to superimpose the mask signal on the first license plate image to obtain a second license plate image; and use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image.
[0152] It should be noted that the client device 1402 and the server device 1404 correspond to steps S402 to S404 in Example 1. The instances and application scenarios implemented by the two devices and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the apparatus, can be run in the computer terminal 10 provided in Example 1.
[0153] According to an embodiment of the present invention, a device for implementing the above-mentioned fourth image processing method is also provided. Figure 15 : is a structural block diagram of an image processing apparatus according to Embodiment 2 of the present invention. Figure 15As shown, the device includes: a second acquisition module 1502, a second identification module 1504, a third acquisition module 1506 and an issuing module 1508. The device is described below.
[0154] The second acquisition module 1502 is used to acquire the first license plate image of the vehicle; the second recognition module 1504 is connected to the above-mentioned second acquisition module 1502, and is used to use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; the third acquisition module 1506 is connected to the above-mentioned second recognition module 1504, and is used to acquire the vehicle's load according to the text recognition result; the issuing module 1508 is connected to the above-mentioned third acquisition module 1506, and is used to issue freight instructions for the vehicle according to the load.
[0155] It should be noted that the second acquisition module 1502, second identification module 1504, third acquisition module 1506, and issuance module 1508 described above correspond to steps S502 to S508 in Example 1. The examples and application scenarios implemented by these modules and corresponding steps are the same, but are not limited to those disclosed in Example 1. It should be noted that the above modules, as part of the apparatus, can be run in the computer terminal 10 provided in Example 1.
[0156] According to an embodiment of the present invention, a device for implementing the fifth image processing method is also provided. Figure 16 : is a structural block diagram of an image processing device 5 according to embodiment 2 of the present invention. Figure 16 As shown, the device includes: a fourth acquisition module 1602, a third identification module 1604 and a determination module 1606. The device is described below.
[0157] The fourth acquisition module 1602 is used to obtain the first license plate image of the vehicle and the audio data when the vehicle passes by; the third recognition module 1604 is connected to the above-mentioned fourth acquisition module 1602, and is used to use the text recognition model to recognize the second license plate image to obtain the text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing the mask signal generated by the signal generator on the first license plate image; the determination module 1606 is connected to the above-mentioned third recognition module 1604, and is used to determine the driving status of the vehicle based on the text recognition result and the audio data.
[0158] It should be noted that the fourth acquisition module 1602, the third identification module 1604, and the determination module 1606 correspond to steps S602 to S606 in Example 1. The examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0159] According to an embodiment of the present invention, a device for implementing the sixth image processing method is also provided. Figure 17 : is a structural block diagram of an image processing apparatus 6 according to embodiment 2 of the present invention. Figure 17 As shown, the device includes: a fifth acquisition module 1702, a fourth identification module 1704 and a second generation module 1706. The device is described below.
[0160] The fifth acquisition module 1702 is used to obtain the first license plate image of the vehicle and the positioning data of the vehicle; the fourth recognition module 1704 is connected to the above-mentioned fifth acquisition module 1702, and is used to use the text recognition model to recognize the second license plate image to obtain the text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing the mask signal generated by the signal generator on the first license plate image; the second generation module 1706 is connected to the above-mentioned fourth recognition module 1704, and is used to generate a navigation route for the vehicle based on the text recognition result and positioning data.
[0161] It should be noted that the fifth acquisition module 1702, the fourth identification module 1704, and the second generation module 1706 correspond to steps S702 to S706 in Example 1. The examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0162] Example 3
[0163] The embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0164] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0165] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the image processing method of the application: obtaining a first license plate image; using a signal generator to generate a mask signal; superimposing the mask signal on the first license plate image to obtain a second license plate image; using a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image.
[0166] Optionally, Figure 18 1 is a block diagram of a computer terminal according to an embodiment of the present invention. Figure 18 As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 1802, memory 1804, etc.
[0167] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned image processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranet, local area network, mobile communication network and combinations thereof.
[0168] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first license plate image; use a signal generator to generate a mask signal; superimpose the mask signal on the first license plate image to obtain a second license plate image; use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image.
[0169] Optionally, the above-mentioned processor can also execute the program code of the following steps: when the mask signal includes a first mask signal and a second mask signal, superimposing the mask signal on the first license plate image to obtain a second license plate image, including: superimposing the first mask signal on the first image area of the first license plate image; superimposing the second mask signal on the second image area of the first license plate image; wherein the first license plate image includes the first image area and the second image area.
[0170] Optionally, the processor may further execute program code of the following steps: the first image area includes an upper half of the first license plate image, and the second image area includes a lower half of the first license plate image.
[0171] Optionally, the processor may further execute program code of the following steps: the first image area includes an upper third of the first license plate image, and the second image area includes a lower two-thirds of the first license plate image.
[0172] Optionally, the processor may further execute program code of the following steps: the first mask signal includes a sine mask signal, and the second mask signal includes a cosine mask signal; or the first mask signal includes a cosine mask signal, and the second mask signal includes a sine mask signal.
[0173] Optionally, the processor may further execute program codes of the following steps: obtaining an original license plate image; performing text region detection on the original license plate image to obtain a first license plate image including a text region.
[0174] Optionally, the above-mentioned processor can also execute the program code of the following steps: before using the text recognition model to identify the second license plate image and obtain the text recognition result on the first license plate image, it also includes: obtaining a sample training set, wherein the training data in the sample training set includes: a license plate image and a license plate in the license plate image, and the license plate image is an image obtained after superimposing a mask signal; using the sample training set to train the initial model to obtain a text recognition model.
[0175] Optionally, the processor may further execute program code of the following steps: the first license plate image includes: an image of a double-layer license plate.
[0176] Optionally, the processor may further execute program code of the following steps: the text recognition model includes at least one of the following: a text recognition model based on a convolutional recurrent neural network, and a text recognition model based on an attention mechanism.
[0177] Optionally, the processor may also execute program codes for the following steps: locating the vehicle corresponding to the license plate marked by the text recognition result; measuring the speed of the vehicle corresponding to the license plate marked by the text recognition result; and monitoring the vehicle corresponding to the license plate marked by the text recognition result.
[0178] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving a first license plate image in the input area of the display interface; receiving a license plate recognition instruction on the display interface; in response to the license plate recognition instruction, displaying the text recognition result of the first license plate image on the display interface, wherein the text recognition result is obtained by using a text recognition model to recognize the second license plate image, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image.
[0179] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: the client device sends the first license plate image to the server device; the server device uses a signal generator to generate a mask signal; the server device superimposes the mask signal on the first license plate image to obtain a second license plate image; and uses a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image; the client device receives the text recognition result returned by the server device.
[0180] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first license plate image of the vehicle; use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; obtain the vehicle's load based on the text recognition result; and issue a freight instruction for the vehicle based on the load.
[0181] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first license plate image of the vehicle and audio data when the vehicle passes by; use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; determine the driving status of the vehicle based on the text recognition result and audio data.
[0182] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a first license plate image of the vehicle and the positioning data of the vehicle; use a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; generate a navigation route for the vehicle based on the text recognition result and the positioning data.
[0183] It can be understood by those skilled in the art that Figure 18 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 18 It does not limit the structure of the above electronic device. For example, the computer terminal may also include Figure 18 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 18 Different configurations shown.
[0184] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0185] Example 4
[0186] The embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the image processing method provided in the first embodiment.
[0187] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0188] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining a first license plate image; generating a mask signal using a signal generator; superimposing the mask signal on the first license plate image to obtain a second license plate image; and recognizing the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image.
[0189] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: when the mask signal includes a first mask signal and a second mask signal, superimposing the mask signal on the first license plate image to obtain a second license plate image, including: superimposing the first mask signal on the first image area of the first license plate image; superimposing the second mask signal on the second image area of the first license plate image; wherein the first license plate image includes the first image area and the second image area.
[0190] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: the first image area includes an upper half of the first license plate image, and the second image area includes a lower half of the first license plate image.
[0191] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: the first image area includes an upper third of the first license plate image, and the second image area includes a lower two-thirds of the first license plate image.
[0192] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: the first mask signal includes: a sine mask signal, and the second mask signal includes: a cosine mask signal; or, the first mask signal includes: a cosine mask signal, and the second mask signal includes: a sine mask signal.
[0193] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: acquiring an original license plate image; performing text region detection on the original license plate image to obtain a first license plate image including a text region.
[0194] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: before using the text recognition model to identify the second license plate image and obtain the text recognition result on the first license plate image, it also includes: obtaining a sample training set, wherein the training data in the sample training set includes: a license plate image and a license plate in the license plate image, and the license plate image is an image obtained after superimposing a mask signal; using the sample training set to train the initial model to obtain a text recognition model.
[0195] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: the first license plate image includes: an image of a double-layer license plate.
[0196] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: the text recognition model includes at least one of the following: a text recognition model based on a convolutional recurrent neural network, a text recognition model based on an attention mechanism.
[0197] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: the method also includes at least one of the following: locating the vehicle corresponding to the license plate marked by the text recognition result; measuring the speed of the vehicle corresponding to the license plate marked by the text recognition result; monitoring the vehicle corresponding to the license plate marked by the text recognition result.
[0198] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving a first license plate image in an input area of a display interface; receiving a license plate recognition instruction on the display interface; and displaying a text recognition result of the first license plate image on the display interface in response to the license plate recognition instruction, wherein the text recognition result is obtained by recognizing a second license plate image using a text recognition model, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image.
[0199] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: the client device sends the first license plate image to the server device; the server device uses a signal generator to generate a mask signal; the server device superimposes the mask signal on the first license plate image to obtain a second license plate image; and uses a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image; the client device receives the text recognition result returned by the server device.
[0200] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining a first license plate image of the vehicle; using a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; obtaining the vehicle's load based on the text recognition result; and issuing freight instructions for the vehicle based on the load.
[0201] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining a first license plate image of the vehicle and audio data when the vehicle passes by; using a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and determining the driving status of the vehicle based on the text recognition result and the audio data.
[0202] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining a first license plate image of the vehicle and the positioning data of the vehicle; using a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image; and generating a navigation route for the vehicle based on the text recognition result and the positioning data.
[0203] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0204] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0205] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0206] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0207] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0209] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that: include: Obtaining a first license plate image; A signal generator is used to generate a mask signal; superimposing the mask signal on the first license plate image to obtain a second license plate image, wherein the mask signal is used to clearly distinguish the text portion of the first license plate image from the non-text portion, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; The second license plate image is recognized using a text recognition model to obtain a text recognition result on the first license plate image.
2. The method according to claim 1, characterized in that In a case where the mask signal includes a first mask signal and a second mask signal, superimposing the mask signal on the first license plate image to obtain the second license plate image includes: superimposing the first mask signal on a first image region of the first license plate image; superimposing the second mask signal on the second image area of the first license plate image; The first license plate image includes the first image area and the second image area.
3. The method according to claim 2, characterized in that The first image area includes an upper half of the first license plate image, and the second image area includes a lower half of the first license plate image.
4. The method according to claim 2, characterized in that The first image area includes an upper third of the first license plate image, and the second image area includes a lower two-thirds of the first license plate image.
5. The method according to claim 2, characterized in that The first mask signal includes: a sine mask signal, and the second mask signal includes: a cosine mask signal; or, The first mask signal includes a cosine mask signal, and the second mask signal includes a sine mask signal.
6. The method according to claim 1, characterized in that Acquiring the first license plate image includes: Get the original license plate image; Perform text region detection on the original license plate image to obtain the first license plate image including the text region.
7. The method according to claim 1, characterized in that Before using the text recognition model to recognize the second license plate image to obtain the text recognition result on the first license plate image, the method further includes: Obtaining a sample training set, wherein the training data in the sample training set includes: a license plate image and a license plate in the license plate image, wherein the license plate image is an image obtained by superimposing a mask signal; The sample training set is used to train the initial model to obtain the text recognition model.
8. The method according to any one of claims 1 to 7, characterized in that The first license plate image includes: an image of a double-layer license plate.
9. The method according to claim 8, characterized in that The text recognition model includes at least one of the following: a text recognition model based on a convolutional recurrent neural network, and a text recognition model based on an attention mechanism.
10. The method according to claim 9, characterized in that The method further comprises at least one of the following: Locating the vehicle corresponding to the license plate marked by the text recognition result; Measuring the speed of the vehicle corresponding to the license plate marked by the text recognition result; The vehicle corresponding to the license plate marked by the text recognition result is monitored.
11. An image processing method, characterized in that: include: Receiving a first license plate image in an input area of the display interface; receiving a license plate recognition instruction on the display interface; In response to the license plate recognition instruction, the text recognition result of the first license plate image is displayed on the display interface, wherein the text recognition result is obtained by recognizing the second license plate image using a text recognition model, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image, wherein the mask signal is used to make the text part of the first license plate image clearly distinguishable from the non-text part, and when the first license plate image is divided into different image areas, different mask signals are superimposed on different image areas.
12. An image processing method, characterized in that: include: The client device sends the first license plate image to the server device; The server device uses a signal generator to generate a mask signal, wherein the mask signal is used to clearly distinguish the text portion of the first license plate image from the non-text portion, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; The server device superimposes the mask signal on the first license plate image to obtain a second license plate image; and uses a text recognition model to recognize the second license plate image to obtain a text recognition result on the first license plate image; The client device receives the text recognition result returned by the server device.
13. An image processing device, characterized in that: include: A first acquisition module, configured to acquire a first license plate image; A first generating module, configured to generate a mask signal using a signal generator; a superposition module, configured to superimpose the mask signal on the first license plate image to obtain a second license plate image, wherein the mask signal is used to clearly distinguish the text portion of the first license plate image from the non-text portion, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; The first recognition module is used to recognize the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image.
14. An image processing device, characterized in that: include: A first receiving module is configured to receive a first license plate image in an input area of a display interface; A second receiving module is used to receive a license plate recognition instruction on the display interface; A display module is used to respond to the license plate recognition instruction and display the text recognition result of the first license plate image on the display interface, wherein the text recognition result is obtained by recognizing the second license plate image using a text recognition model, and the second license plate image is an image obtained by superimposing a mask signal on the first license plate image, wherein the mask signal is used to make the text part of the first license plate image clearly distinguishable from the non-text part, and when the first license plate image is divided into different image areas, different mask signals are superimposed on different image areas.
15. An image processing system, characterized in that: It includes a client device and a server device, wherein the client device is used to send a first license plate image to the server device; The server device is configured to generate a mask signal using a signal generator; The server device is further configured to superimpose the mask signal on the first license plate image to obtain a second license plate image; and to recognize the second license plate image using a text recognition model to obtain a text recognition result on the first license plate image, wherein the mask signal is configured to clearly distinguish the text portion of the first license plate image from the non-text portion, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; The client device is further configured to receive the text recognition result returned by the server device.
16. A storage medium, characterized in that The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the image processing method according to any one of claims 1 to 12.
17. A computer device, characterized in that: include: memory and processor, The memory stores a computer program; The processor is configured to execute a computer program stored in the memory, and when the computer program is executed, the processor is enabled to execute the image processing method according to any one of claims 1 to 12.
18. An image processing method, characterized in that: include: Acquire a first license plate image of the vehicle; Using a text recognition model, recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image, wherein the mask signal is used to clearly distinguish a text portion from a non-text portion of the first license plate image, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; Obtaining the load of the vehicle according to the text recognition result; A freight instruction is issued for the vehicle according to the load.
19. An image processing method, characterized in that: include: Acquire a first license plate image of a vehicle and audio data of the vehicle passing by; Using a text recognition model, recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image, wherein the mask signal is used to clearly distinguish a text portion from a non-text portion of the first license plate image, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; The driving state of the vehicle is determined according to the text recognition result and the audio data.
20. An image processing method, characterized in that: include: Acquire a first license plate image of a vehicle and positioning data of the vehicle; Using a text recognition model, recognize a second license plate image to obtain a text recognition result on the first license plate image, wherein the second license plate image is obtained by superimposing a mask signal generated by a signal generator on the first license plate image, wherein the mask signal is used to clearly distinguish a text portion from a non-text portion of the first license plate image, and when the first license plate image is divided into different image regions, different mask signals are superimposed on different image regions; A navigation route is generated for the vehicle based on the text recognition result and the positioning data.
Citation Information
Patent Citations
A yellow double-row license plate character segmentation method based on a row segmentation line
CN109034019A