Image Recognition Method, Apparatus, Electronic Device and Readable Medium

By detecting and stitching the character position and connection relationship in the container area position code in image recognition, the problem of identifying errors in the prior art is solved, and higher recognition accuracy and container maintenance efficiency are achieved.

CN115131777BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210393386.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-06-27
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

In the prior art, when identifying images of container area position codes, it is difficult to handle changes in the arrangement of characters, resulting in recognition errors and affecting the accuracy of the recognition results.

Method used

By obtaining the image to be recognized, the image recognition model is used to detect the position and connection relationship of each character, and splicing it with the character recognition results to generate text recognition results, which are independent of the character arrangement.

Benefits of technology

It improves the accuracy of image recognition, ensures accurate identification of container area location codes, and improves the efficiency of container maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131777B_ABST
    Figure CN115131777B_ABST
Patent Text Reader

Abstract

The present application provides an image recognition method, apparatus, electronic device, and readable medium. The method includes: obtaining a to-be-recognized image including to-be-recognized text, where the to-be-recognized text includes multiple characters; performing image recognition on the to-be-recognized image to obtain a character position result of each character, a character connection result of the multiple characters, and a character recognition result of each character, where the character position result is used to indicate the position of the character in the to-be-recognized image, and the character connection result is used to indicate the adjacency relationship between each character and adjacent characters; performing character recognition on each character in the to-be-recognized image respectively according to the character position results of the respective characters to obtain the character recognition results of the respective characters; and splicing the multiple characters according to the character recognition results and the character connection result to obtain a text recognition result of the to-be-recognized text. This method can improve the accuracy of the recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an image recognition method, apparatus, electronic device, and readable medium. Background Art

[0002] During long-distance transportation, containers may have defects such as deformation, breakage, and rusting. These defects need to be reported for repair. When reporting information, it is necessary to upload a container image containing a regional location code so that the container can be located later based on the regional location code.

[0003] In the related art, for the uploaded container image, an image recognition model is usually used to detect the regional location code in the uploaded image, so as to automatically recognize the regional location code.

[0004] However, such methods are difficult to accurately recognize changes in the arrangement of the regional location code in the image, and often result in recognition errors due to changes in the arrangement of the regional location code, thus affecting the accuracy of the recognition result. Summary of the Invention

[0005] Based on the above technical problems, this application provides an image recognition method, apparatus, electronic device, and readable medium, so that during the image recognition process, the recognition of text is not affected by the arrangement of characters in the regional location code, thereby improving the accuracy of the recognition result and facilitating the improvement of the efficiency of container repair.

[0006] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.

[0007] According to one aspect of the embodiments of this application, an image recognition method is provided, including:

[0008] Obtain a to-be-recognized image containing to-be-recognized text, where the to-be-recognized text includes multiple characters;

[0009] Perform image recognition on the to-be-recognized image to obtain the character position result of each character, the character connection result of the multiple characters, and the character recognition result of each character. The character position result is used to indicate the position of the character in the to-be-recognized image, and the character connection result is used to indicate the adjacency relationship between each character and its adjacent characters;

[0010] Splice the multiple characters according to the character recognition result and the character connection result to obtain the text recognition result of the to-be-recognized text.

[0011] According to one aspect of the embodiments of this application, an image recognition apparatus is provided, including:

[0012] An image acquisition module, configured to acquire a to-be-recognized image including to-be-recognized text, where the to-be-recognized text includes multiple characters;

[0013] An image recognition module, configured to perform image recognition on the to-be-recognized image to obtain a character position result of each character, a character connection result of the multiple characters, and a character recognition result of each character, where the character position result is used to indicate the position of the character in the to-be-recognized image, and the character connection result is used to indicate the adjacency relationship between each character and adjacent characters;

[0014] A character splicing module, configured to splice the multiple characters according to the character recognition result and the character connection result to obtain a text recognition result of the to-be-recognized text.

[0015] In some embodiments of the present application, based on the above technical solution, the character position result includes the center point position of each character, and the character connection result includes a character adjacency matrix for representing the adjacency relationship between characters; the image recognition module includes:

[0016] A downsampling sub-module, configured to perform downsampling on the to-be-recognized image according to multiple scales to obtain image features at the multiple scales;

[0017] A feature fusion sub-module, configured to perform feature fusion on the image features at the multiple scales to obtain a feature fusion result at the multiple scales;

[0018] A position detection sub-module, configured to detect the positions of each character according to the feature fusion result at the multiple scales to obtain the center point position of each character;

[0019] An adjacency analysis sub-module, configured to analyze the adjacency relationship between each character according to the feature fusion result at the multiple scales and the center point position of each character to obtain a character adjacency matrix between each character;

[0020] A character recognition module, configured to perform character recognition on each character in the to-be-recognized image respectively according to the character position result of each character to obtain a character recognition result of each character.

[0021] In some embodiments of the present application, based on the above technical solution, the position detection sub-module includes:

[0022] A convolution unit, configured to perform convolution processing on the multi-scale feature information to obtain a convolution result;

[0023] A center feature prediction unit, configured to perform center feature prediction according to the convolution result to obtain a center point feature map, where each feature value in the center point feature map represents the probability that the corresponding pixel point is the center point of the character;

[0024] A center point determination unit, configured to determine the center point positions of each character according to the center point feature map.

[0025] In some embodiments of the present application, based on the above technical solutions, the position detection sub-module further includes:

[0026] A vertex distance prediction unit, configured to perform vertex distance prediction according to the convolution result to obtain a vertex feature map of each character, where each feature value in the vertex feature map represents the distance between the corresponding pixel point in the image to be recognized and each vertex of the character's border;

[0027] A border determination unit, configured to determine the border positions of each character according to the vertex feature map.

[0028] In some embodiments of the present application, based on the above technical solutions, the position detection sub-module further includes:

[0029] An offset prediction unit, configured to perform offset prediction according to the convolution result to obtain an offset feature map, where each feature value in the offset feature map represents the offset amount of the corresponding pixel point in the image to be recognized, and the offset amount is used to adjust the border position;

[0030] A border adjustment unit, configured to adjust the border positions of each character according to the offset amount.

[0031] In some embodiments of the present application, based on the above technical solutions, the adjacency analysis sub-module includes:

[0032] An angle calculation unit, configured to calculate the rotation angle and character scale of each character in the image to be recognized according to the feature fusion result at multiple scales;

[0033] A position feature determination unit, configured to determine the character position features of each character according to the center point positions of each character, as well as the rotation angle and character scale of each character;

[0034] A similarity calculation unit, configured to calculate the similarity between each character and other characters according to the center point positions of each character;

[0035] A connection graph construction unit, configured to construct a connection graph of each character according to the character position features of each character and the similarity between each character and other characters;

[0036] An adjacency relationship determination unit, configured to determine the adjacency relationship between each character according to the connection graph of each character;

[0037] A matrix construction unit, configured to construct the character adjacency matrix according to the adjacency relationship between each character.

[0038] In some embodiments of the present application, based on the above technical solutions, the adjacency relationship determination unit includes:

[0039] A graph convolution sub-unit, configured to perform graph convolution operations with the nodes corresponding to the characters as the central points for the connection graph of each character, and obtain a graph convolution result;

[0040] An adjacency determination sub-unit, configured to determine the adjacency relationship between each character according to the connection result between the nodes in the graph convolution result.

[0041] In some embodiments of the present application, based on the above technical solutions, the character recognition module includes:

[0042] An image screenshot sub-module, configured to intercept the character images corresponding to each character from the image to be recognized according to the character position results of each character;

[0043] A classification prediction sub-module, configured to input the character images corresponding to each character into a character classification model for prediction respectively, and obtain a character classification result corresponding to each character, where the character classification result includes at least one result character corresponding to the character.

[0044] In some embodiments of the present application, based on the above technical solutions, the character recognition module further includes:

[0045] A classification module training sub-module, configured to train a character classification model to be trained through character training data, and obtain a training classification result;

[0046] A loss calculation sub-module, configured to perform joint calculation of a triplet loss function and a cross-entropy loss function according to the training classification result, and obtain a training loss result;

[0047] A parameter adjustment sub-module, configured to adjust the model parameters of the character classification model to be trained according to the training loss result, and obtain the character classification model.

[0048] In some embodiments of the present application, based on the above technical solutions, the character splicing module includes:

[0049] A character sequence determination sub-module, configured to determine the character sequence of each character according to the adjacency relationship in the character connection result;

[0050] An arrangement splicing sub-module, configured to arrange and splice the character recognition results according to the character sequence, and obtain a text recognition result of the text to be recognized.

[0051] In some embodiments of the present application, based on the above technical solutions, each sequence position in the character sequence corresponds to a candidate character set; the arrangement splicing sub-module includes:

[0052] A candidate set determination unit, configured to determine, for each character, a candidate character set corresponding to the character according to the sequence position of the character in the character sequence;

[0053] A result character determination unit, configured to determine a result character for each character according to the intersection of the character classification result of each character and the determined candidate character set;

[0054] A text splicing unit, configured to arrange and splice the result characters according to the character sequence to obtain a text recognition result of the text to be recognized.

[0055] In some embodiments of the present application, based on the above technical solution, the image recognition device further includes:

[0056] A field intercepting module, configured to intercept a picture of the text to be recognized from the image to be recognized according to the character position result and the character connectivity result;

[0057] A field recognition module, configured to perform field recognition on the picture of the text to be recognized to obtain a field recognition result of the text to be recognized, where the field recognition result includes recognition results of the multiple characters;

[0058] A result correction module, configured to correct the text recognition result according to the field recognition result.

[0059] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the image recognition method in the above technical solution by executing the executable instructions.

[0060] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image recognition method in the above technical solution is implemented.

[0061] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image recognition method provided in the above various optional implementation manners.

[0062] In an embodiment of the present application, after obtaining the image to be recognized of the container, image recognition is performed on the image to be recognized through an image recognition model to obtain the character position results of each character and the character connection results of multiple characters. Subsequently, according to the character position results of each character, character recognition is performed on each character in the image to be recognized to obtain the character recognition results of each character. Then, according to the character recognition results and the character connection results, splicing is performed to obtain the image recognition result of the area position code. Through the above method, when performing image recognition, the position of a single character and the adjacency relationship of each character are output simultaneously. After recognizing a single character according to the position of the single character in the image, the recognition result is spliced according to the character connection result, so that in the image recognition process, the recognition of the text is not affected by the arrangement of the characters in the text to be recognized, thereby improving the accuracy of the recognition result and being beneficial to improving the efficiency of container maintenance.

[0063] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] In the drawings:

[0066] Figure 1 Schematically shows an exemplary system architecture diagram of the technical solution of the present application in an application scenario;

[0067] Figure 2 Is the overall process schematic diagram of the image recognition method in the embodiment of the present application;

[0068] Figure 3 Is the schematic flowchart of the content recommendation method in the embodiment of the present application;

[0069] Figure 4 Is the schematic diagram of the image to be recognized in the embodiment of the present application;

[0070] Figure 5 Is the schematic diagram of the structure of the image recognition model in the embodiment of the present application;

[0071] Figure 6 Is the schematic flowchart of the content recommendation method in the embodiment of the present application;

[0072] Figure 7Shows a schematic diagram of the calculation of the graph convolutional network in the embodiments of the present application;

[0073] Figure 8 Schematic diagram of the image to be recognized in the embodiments of the present application;

[0074] Figure 9 Schematic flowchart of the content recommendation method in the embodiments of the present application;

[0075] Figure 10 Schematically shows the block diagram of the composition of the image recognition device in the embodiments of the present application;

[0076] Figure 11 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0077] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0078] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0079] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0080] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0081] It should be understood that the solution of the present application can be applied in the field of image recognition, and specifically applied to the recognition scenario of the regional location code of containers. Specifically, various defects such as deformation, breakage, and rust may occur to the containers during the long-distance transportation process. These defects need to be reported for repair and compensation. When reporting information, specific defect locations and defect pictures are required to form a number of defect entries, and the auditor will match the entries with the pictures to review the accuracy of each defect entry. Among them, the defect pictures will include the regional location code, so as to mark the relevant information of the container defects. However, in the captured container images, the writing styles of the regional location codes are different, the arrangement styles are diverse, the brightness is different, the shooting fields of view are different, and the text scales are different. These problems have a great impact on the extraction of the regional location codes of containers. The solution of the present application can accurately recognize these pictures, so as to accurately recognize the handwritten regional location codes in various situations from the captured pictures.

[0082] In order to obtain a relatively high-accuracy image recognition result in the above various scenarios, the present application proposes an image recognition method, which is applied to Figure 1 the image recognition system shown in, please refer to Figure 1 , Figure 1 which schematically shows an exemplary system architecture diagram of the technical solution of the present application in an application scenario. The image recognition system includes a server and terminal devices. The aforementioned image recognition method can be executed by an image recognition device or an image recognition service on the server, or can be executed by a terminal device with relatively strong computing power.

[0083] Among them, as Figure 1 shown, the aforementioned terminal devices include but are not limited to tablet computers, laptop computers, palmtop computers, mobile phones, voice interaction devices and personal computers (PCs), vehicle-mounted terminals, aircraft, etc., which are not limited here. Among them, the voice interaction devices include but are not limited to smart speakers and smart home appliances. In some implementation manners, the client can be presented as a web client or an application client, and is deployed on the aforementioned terminal devices. Figure 1 The server in

[0084] can be a single server or a server cluster or a cloud computing center composed of multiple servers, etc., which are not limited specifically here. Figure 1 Although Figure 1 only five terminal devices and one server are shown in

[0085] This application introduces an image recognition method to solve the problem that the code arrangement affects the recognition accuracy in the recognition of the regional location code of containers. Specifically, please refer to Figure 2 , Figure 2 , which is a schematic diagram of the overall process of the image recognition method in the embodiment of this application. As Figure 2 shown, the overall process of this image recognition method includes four steps. First, in step 210, it is necessary to photograph the position of the regional code of the container to obtain the image to be processed as input data. Subsequently, based on deep learning network technology, the detection results of single characters and the results of character connected components are output. As Figure 2 shown, the method is divided into two branches, namely the single-character information detection branch and the character connected-component analysis branch. In step 220, for the single-character information detection branch, the value of each pixel in the output classification feature map represents the probability that the pixel is the center point of the target box, and the value of each pixel in the localization feature map represents the horizontal and vertical distances of the pixel from the four vertices of the character border. In step 230, for the character connected-component branch, several graphs are obtained based on the graph convolutional network, each graph corresponding to a character, and each node in the graph corresponding to its neighboring characters. The score between two nodes represents the probability that the two characters belong to the same field. Subsequently, in step 240, for the result output by the single-character detection branch, based on its position information, single characters are expanded and cropped from the original image, and character recognition is performed on them to determine the character category. Usually, the character category includes numbers from 0 to 9 and some capital letters from A to Z. Finally, in step 250, the results of character recognition and the character connected-component analysis branch are comprehensively and structurally processed, so as to obtain the text recognition result at the field level as the result of image recognition.

[0086] More specifically, the image recognition device may be embodied as a client deployed on a terminal device, such as all the clients shown in the above examples of the application scenarios of this application. Then, the server may send the image recognition device to the terminal device via a wireless network. The image recognition device may also be embodied as a terminal device dedicated to image recognition. Then, after generating the image recognition device, the server may also configure the image recognition device on the terminal device via a wired network or a removable storage medium, etc. The image recognition device may also be deployed on the server. After the terminal device obtains the image to be recognized, it sends the image to be recognized to the server. After the server performs the image recognition operation, it is then sent to the terminal device, etc. This application is described by taking the image recognition device deployed on the server as an example, but this should not be construed as a limitation of this application. Further, the above wireless network uses standard communication technologies and / or protocols. The wireless network is usually the Internet, but it can also be any network, including but not limited to any combination of Bluetooth, Local Area Network (LAN), Metropolitan Area Network (MAN), Wide Area Network (WAN), mobile, private network or virtual private network). In some embodiments, customized or dedicated data communication technologies may be used to replace or supplement the above data communication technologies.

[0087] The solution of this application can be implemented relying on artificial intelligence. For example, the above image recognition process can be implemented by means of machine learning.

[0088] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.

[0089] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0090] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0091] Combined with the above introduction, the technical solutions provided by this application will be described in detail below in conjunction with specific embodiments. Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the content recommendation method in the embodiments of this application. An embodiment of the content recommendation method in the embodiments of this application includes:

[0092] Step S310, obtain a to-be-recognized image including to-be-recognized text, where the to-be-recognized text includes multiple characters.

[0093] The to-be-recognized image can be obtained by taking a photo of the area on the container where the area location code is written, and the area location code is the to-be-recognized text. A to-be-recognized graphic can include multiple area location codes. The area location code is a field composed of multiple characters. Specifically, please refer to Figure 4 , Figure 4 which is a schematic diagram of the to-be-recognized image in the embodiments of this application. As Figure 4 shown, this image includes an area location code composed of 4 characters. In an actual scenario, the area location code on the container is usually handwritten by humans and then photographed by humans to obtain a photo. Therefore, various shooting situations will occur in the to-be-recognized image, such as various angles, sizes, fonts, and arrangements.

[0094] Step S320, perform image recognition on the to-be-recognized image to obtain the character position result of each character, the character connection result of the multiple characters, and the character recognition result of each character. The character position result is used to indicate the position of the character in the to-be-recognized image, and the character connection result is used to indicate the adjacency relationship between each character and its adjacent characters.

[0095] Input the image to be recognized into an image recognition model for prediction. For each character, output the character position result, and for all characters, output the character connectivity result. The character position result is used to indicate the position of the character in the image to be recognized. Specifically, the character position can be marked with a border, and the specific form of representing the position can adopt any suitable way, such as a rectangle represented by diagonal vertices or a circle represented by an origin plus a radius, etc. The image recognition model confirms the positions of the boundaries of each character in the image to be recognized through convolution and calculation of the image to be recognized, so as to obtain the character position result. The character connectivity result is used to indicate the adjacency relationship between each character and adjacent characters. The image recognition model can analyze and calculate based on the character position result to obtain the distances between each character. According to the distances between characters, judge which characters belong to the same regional position code, construct the adjacency relationship between each character and adjacent characters, and obtain the character connection result.

[0096] Step S330, splice the multiple characters according to the character recognition result and the character connectivity result to obtain the text recognition result of the text to be recognized.

[0097] According to the character connectivity result, the number of fields included in the text to be recognized in the image to be recognized can be determined. For example, an image to be recognized may include two regional position codes A and B, each regional position code consists of 4 characters, and the character connectivity result will include the connection relationship corresponding to the two regional position codes, that is, each group of 4 characters. According to the adjacency relationship in the character connectivity result, the character recognition results can be spliced together to form the corresponding fields, so as to obtain the text recognition result of the text to be recognized. Specifically, in an actual scenario, the regional position code including 4 characters may be in the form of a row, a column, or two rows, etc. For the sequence of regional position codes, there are specific rules for the characters that will appear at each position. For example, it only includes numbers or only includes uppercase English, etc. According to such specific rules, the recognition result can be further judged. For a row or a column of characters, the sequence order can be judged in turn according to the order of the recognized character sequence. For example, the regional position code starts with a letter and ends with a number, and among the four characters recognized in a row, the rightmost character is a letter and the leftmost character is a number, then the splicing order of the regional position code is spliced in the order from right to left. For two rows of characters, each character is matched according to the specific rules of each sequence position to obtain the order of the characters. If the character recognition result includes multiple possible characters, then specific characters can be further determined according to the rules of the text to be recognized itself during splicing. For example, if the text to be recognized consists of pure uppercase English letters, the results of lowercase letters and numbers can be excluded to determine the final characters.

[0098] In an embodiment of the present application, after obtaining the image to be recognized of the container, image recognition is performed on the image to be recognized through an image recognition model to obtain the character position results of each character and the character connection results of multiple characters. Subsequently, according to the character position results of each character, character recognition is performed on each character in the image to be recognized to obtain the character recognition results of each character, and then the character recognition results and the character connection results are spliced to obtain the image recognition result of the area position code. Through the above method, when performing image recognition, the position of a single character and the adjacency relationship of each character are output simultaneously. After recognizing a single character according to the position of the single character in the image, the recognition result is spliced according to the character connection result, so that in the image recognition process, the recognition of the text is not affected by the arrangement of the characters in the area position code, thereby improving the accuracy of the recognition result and being beneficial to improving the efficiency of container maintenance.

[0099] In an embodiment of the present application, based on the above technical solution, the character position result includes the center point position of each character, and the character connection result includes a character adjacency matrix for representing the adjacency relationship between characters; the above step S320 of performing image recognition on the image to be recognized to obtain the character position results of each character, the character connection results of the multiple characters, and the character recognition results of each character specifically includes the following steps:

[0100] Perform downsampling on the image to be recognized according to multiple scales to obtain image features at multiple scales;

[0101] Perform feature fusion on the image features at multiple scales to obtain feature fusion results at multiple scales;

[0102] Detect the positions of each character according to the feature fusion results at multiple scales to obtain the center point positions of each character;

[0103] Analyze the adjacency relationship of each character according to the feature fusion results at multiple scales and the center point positions of each character to obtain a character adjacency matrix between each character;

[0104] Perform character recognition on each character in the image to be recognized respectively according to the character position results of each character to obtain the character recognition results of each character.

[0105] Specifically, the image recognition model can be composed of multiple sub-models or sub-networks. For the convenience of introduction, please refer to Figure 5 , Figure 5 which is the structural schematic diagram of the image recognition model in the embodiment of the present application. As Figure 5As shown, the process of downsampling the image to be recognized according to multiple scales can be implemented by a neural network, such as using a neural network like HRNet. The input is processed by the neural network, which then downsamples the image with features to be recognized into feature maps of multiple different scales or resolutions, thereby obtaining image feature maps at multiple scales. Specifically, the image feature map is usually a matrix, and its resolution depends on the specific downsampling method. For example, if the image resolution is 256×256, by performing convolution with a stride of 2 during convolution, a feature image with a resolution half that of the original image is obtained. Subsequently, convolution is performed on the newly obtained image with a stride of 2, resulting in a feature image with a resolution one-fourth that of the original. The number of scales can depend on the requirements. Feature fusion is performed on the obtained image feature maps at different scales to obtain the feature fusion results at multiple scales. Specifically, as Figure 5 shown, the process of feature fusion can be implemented using the FPN model. It can be understood that the feature fusion results at multiple scales are actually feature maps at multiple resolutions. These feature maps can reflect the graphic features in the image to be recognized from different scales, thus facilitating the processing of characters at different scales. Specifically, for the feature fusion method, for example, the number of channels is not changed as the high resolution, the feature map at the low resolution scale is upsampled to the high resolution, and then the representations at each obtained scale are concatenated, and 1X1 convolution is used to mix these representations to obtain the feature fusion result. Subsequently, as Figure 5As shown, the image recognition model may include multiple branches to calculate the required data. Detect the positions of individual characters based on the feature fusion results at multiple scales to obtain the center point positions of the individual characters. Specifically, input the feature fusion results into the image recognition model to predict the probability that each pixel point in the feature map is the center point position, and take the point with the highest probability as the prediction result, thereby obtaining the center point positions of the individual characters. In a specific embodiment, the character bounding box will also be determined. The character bounding box can be further determined based on the center point position, or can be determined separately based on the feature fusion results at multiple scales, and then adjusted and matched according to the center point position. Analyze the adjacency relationships of the individual characters based on the feature fusion results at multiple scales and the center point positions of the individual characters to obtain the character adjacency matrix between the individual characters. The character adjacency matrix is a matrix used to represent whether there is a connection relationship between the individual characters. For example, if there are 4 characters, the character adjacency matrix is a 4X4 matrix. If there is an adjacency relationship between two characters, the value at the corresponding position is 1. Specifically, the graphic features of each character in the image to be recognized can be calculated based on the feature fusion results. Then, according to the position of the character in the image to be recognized and the center point position of the character, the distances between the individual characters can be calculated. The distance between the characters can use the Euclidean distance between the center point positions as their distance. And according to the graphic features, the deflection angles, character sizes, and character forms of the individual characters can be compared. For example, judge whether the deflection angles and character sizes are close, and judge whether the character forms (such as Chinese and English characters, English uppercase and lowercase, numbers, or Roman characters, etc.) are similar or belong to the same system. Based on the distance and graphic features, the relationships between the individual characters can be analyzed. Characters with close distances and similar image features have a higher probability of belonging to the same position code, while characters with far distances or dissimilar image features generally do not belong to the same position code. Therefore, in the character adjacency matrix, establish adjacency relationships for characters with close distances and similar graphic features, and do not establish adjacency relationships for characters with relatively far distances or large differences in graphic features, thereby constructing the character adjacency matrix between the individual characters.

[0106] Based on the character position results of the individual characters, the positions of the characters in the figure can be determined. According to the positions of the characters, the images of the individual characters can be extracted from the image to be recognized, and then the individual characters can be recognized separately to obtain the character recognition results of the individual characters. Character recognition can be performed, for example, using neural networks, learning models, or character recognition methods. The character recognition result can be the possible characters corresponding to the individual characters, specifically, it can be one character or multiple characters. For example, the character recognition result of a character can be that the character is one of the number 0, the uppercase letter O, or the lowercase letter o.

[0107] In the embodiments of the present application, by performing multi-scale downsampling and feature fusion processing on the image to be recognized, and then determining the character center points and character adjacency relationships according to the fused features, it is possible to process characters of different sizes, angles, and shapes, thereby improving the robustness of the solution.

[0108] In the embodiments of the present application, based on the above technical solution, in the above step, detecting the positions of each character according to the feature fusion results at multiple scales to obtain the center point positions of each character specifically includes the following steps:

[0109] Performing convolution processing on the multi-scale feature information to obtain a convolution result;

[0110] Performing central feature prediction according to the convolution result to obtain a center point feature map, and each feature value in the center point feature map represents the probability that the corresponding pixel point is the center point of a character;

[0111] Determining the center point positions of each character according to the center point feature map.

[0112] Specifically, please refer to Figure 5 , the process of detecting the positions of each character to obtain the center point positions can be specifically carried out through the target box center point detection branch in Figure 5 . The target box is the border of the character, used to enclose the range of the character, and the center point is the position where the center of the character is located. It can be understood that for some characters, the center point is not necessarily a point on the character. For example, for the number 0 or the letter U, etc., the center point is not on the character itself. The target box center point detection branch can be implemented through a network model. Specifically, the network model uses a 3×3 convolution kernel to perform the convolution process on the input multi-scale feature information to obtain a convolution result, and then performs central feature prediction according to the convolution result to obtain a center point feature map. The center point feature map output by the center point prediction network is a feature map with a channel number of 1. Each pixel point in the center point feature map corresponds to a pixel point in the image to be recognized, and each pixel point in the center point feature map represents the probability that the corresponding pixel point in the image to be recognized is the center point of a character. According to the center point feature map, pixel points with probabilities higher than a certain threshold or the top N pixel points ranked by corresponding probability values can be found as the center points, thereby determining the center point positions of each character. The number of center points can be determined according to the number of characters.

[0113] The loss function of the network model used by the target box center point detection branch is as follows:

[0114]

[0115] Among them, N is the total number of pixel points, xyc represents that the pixel point is located at the (x, y) position of the c-th feature map, where c = 1, Y xyc represents the true value of the pixel point, and the value is a floating point number in the range of 0 to 1, represents the probability that the network predicts that the pixel point is the center point of the target box, and both α and β are hyperparameters, with the default values of α = 2 and β = 4.

[0116] In the embodiments of the present application, by predicting the center point of the character, the center points of each character can be recognized in the case where the characters are close to each other or there are connected strokes in the character writing, which is beneficial to the division of the characters and avoids character recognition errors.

[0117] In the embodiments of the present application, based on the above technical solution, after performing convolution processing on the multi-scale feature information to obtain a convolution result, the method further includes the following steps:

[0118] Predict the vertex distance according to the convolution result to obtain the vertex feature map of each character, and each feature value in the vertex feature map represents the distance between the corresponding pixel point in the image to be recognized and each vertex of the character's border;

[0119] Determine the border position of each character according to the vertex feature map.

[0120] Specifically, the process of vertex distance prediction can be performed through Figure 5 the vertex distance regression branch in. This branch can be implemented by a vertex prediction model. The vertex prediction model performs a convolution process on the input multi-scale feature information through a 3×3 convolution kernel to obtain a convolution result. Subsequently, the vertex distance is predicted according to the convolution result to obtain the vertex feature map of each character. The vertex feature map output by the vertex prediction model is a feature map with 8 channels. Each pixel point in this vertex feature map corresponds to a pixel point in the image to be recognized and represents the distance from the pixel point to each vertex of the character's border. For example, taking a quadrilateral border as an example, a pixel point in the vertex feature map represents the horizontal and vertical axis distances from this point to the upper left, upper right, lower right, and lower left 4 vertices of the border. According to the vertex feature map, the border position of each character can be determined. Specifically, according to the change rule of the distance from the vertex, the positions of the 4 vertices can be determined, and thus the specific range of the quadrilateral border can be further determined. The loss function of the vertex prediction model is shown in the following formula:

[0121] L1 Loss = |f(x) - y|

[0122] Among them, f(x) represents the distance predicted by the network, and y represents the corresponding labeled distance, and they are both floating point numbers.

[0123] In the embodiments of the present application, the position of the border is determined by the vertex feature map obtained by predicting the distances of the vertices of each character border, so that the character border can be represented by the polygon representation method, improving the accuracy of edge detection.

[0124] In the embodiments of the present application, based on the above technical solution, after performing convolution processing on the multi-scale feature information to obtain a convolution result, the method further includes the following steps:

[0125] Perform offset prediction according to the convolution result to obtain an offset feature map, where each feature value in the offset feature map represents the offset amount of the corresponding pixel point in the image to be recognized, and the offset amount is used to adjust the border position;

[0126] Adjust the border positions of the respective characters according to the offset amount.

[0127] Specifically, the process of offset prediction can be performed through Figure 5 the offset prediction regression branch in. This branch can be implemented by an offset prediction model. The offset prediction model performs a convolution process on the input multi-scale feature information through a 3×3 convolution kernel to obtain a convolution result. Subsequently, offset prediction is performed according to the convolution result to obtain the offset feature maps of the respective characters. The offset feature map output by the offset prediction model is a feature map with 2 channels. Each pixel point on this offset feature map corresponds to a pixel point in the image to be recognized and represents the offset amount by which the character border needs to be offset. Specifically, it is represented by the horizontal and vertical axis distances that need to be offset. After determining the offset amount, the character border determined according to the above vertex feature map can be offset and corrected according to this offset feature map to determine a new feature map. The loss function of the offset prediction model is shown in the following formula:

[0128] L1 Loss = |f(x) - y|

[0129] where f(x) represents the offset amount predicted by the network, and y represents the corresponding true value, both of which are floating-point numbers.

[0130] In the embodiments of the present application, by predicting the offset amount of the character border and correcting the character border according to the prediction result, the accuracy loss caused by operations such as multiple downsamplings is compensated, and the accuracy of character border recognition is improved.

[0131] In the embodiments of the present application, based on the above technical solution, for the convenience of introduction, please refer to Figure 6 , Figure 6 which is a schematic flowchart of the content recommendation method in the embodiments of the present application, as shown in Figure 6As shown above, in the above steps, the adjacency relationship between each character is analyzed based on the feature fusion results at the multiple scales and the center point positions of each character, and a character adjacency matrix between each character is obtained. Specifically, it includes the following steps S610 to S660:

[0132] Step S610: Calculate the rotation angle and character scale of each character in the image to be recognized according to the feature fusion results at multiple scales.

[0133] Step S620: Determine the character position features of each character according to the center point position of each character, as well as the rotation angle and character scale of each character.

[0134] Step S630: Calculate the similarity between each character and other characters according to the center point position of each character.

[0135] Step S640: Construct a connection graph for each character according to the character position features of each character and the similarity between each character and other characters.

[0136] Step S650: Determine the adjacency relationship between each character according to the connection graph of each character.

[0137] Step S660: Construct the character adjacency matrix according to the adjacency relationship between each character.

[0138] Specifically, first, according to the feature fusion results at multiple scales, the rotation angle and character scale of each character in the image to be recognized are calculated. Specifically, the character bounding box can be determined according to the feature fusion results at multiple scales, and then the rotation angle and character scale of the character can be determined according to the bounding box. According to the center point position of each character, as well as the rotation angle and character scale of each character, the character position features of each character are determined. The center point position of each character is encoded, and the obtained encoding result is used as the position feature, and the rotation angle and character scale of each character are used as geometric features. The position feature and geometric feature are jointly used as the character position feature of each character. Subsequently, according to the center point position of each character, the similarity between each character and other characters is calculated. Specifically, various similarity calculation methods can be used to calculate the similarity. For example, the Euclidean distance between each character can be calculated according to the center point position of each character, and the calculated Euclidean distance is used as the similarity between the two characters.

[0139] According to the character position features of each character and the similarity between each character and other characters, a connection graph for each character can be constructed. For example, please refer to Figure 7 , Figure 7 which shows a schematic diagram of the calculation of the graph convolutional network in the embodiment of the present application. Figure 7The illustration that has not been processed by the graph convolutional network shows an example of a connection graph constructed with a solid dot as the center point. Specifically, in the connection graph of each character, it is constructed with each character itself as the center point. Using the character position features of the character as nodes and the similarity between characters as edges, according to certain preset conditions, such as according to similarity thresholds and other conditions, the connection graph of each character is constructed. If the similarity between two characters is greater than or equal to the similarity threshold, the nodes corresponding to the two characters are connected in the connection graph. If the similarity between two characters is less than the similarity threshold, the corresponding nodes are not connected. Subsequently, based on the connection graph of each character, the adjacency relationship between each character is determined. Specifically, in the connection graph, starting from the center point corresponding to the character, it is calculated whether there is a connection relationship with other nodes. The specific calculation process can be carried out according to preset rules, such as calculating the probability of connection based on similarity and character position features or classifying according to thresholds. It can also be calculated by a neural network model for probability or classification. If there is a connection relationship, it means that the two characters belong to the same field, that is, belong to the same regional position encoding. Otherwise, it means that the two characters belong to different fields, that is, belong to different regional position encodings. Finally, based on the adjacency relationship between each character, the character adjacency matrix is constructed. The character adjacency matrix includes the connected domains formed by all the node characters with adjacency relationships. Specifically, the character adjacency matrix includes all characters as rows and columns. For example, in the case of 8 characters, the character adjacency matrix is an 8X8 matrix. When there is an adjacency relationship between two characters, the element values corresponding to the two characters are set to 1, and the element values corresponding to two characters without an adjacency relationship are 0.

[0140] In the embodiments of the present application, by using the character position features and the similarity of characters to determine the connection relationship between characters, the obtained results can fully reflect the relevance of the positions of each character in the image to be recognized, which is beneficial to separately processing the uncorrelated fields in the text to be recognized, thereby reducing the internal interference of the text to be recognized.

[0141] In the embodiments of the present application, based on the above technical solution, in the above step, based on the connection graph of each character, determining the adjacency relationship between each character specifically includes the following steps:

[0142] For the connection graph of each character, perform a graph convolution operation with the node corresponding to the character as the center point to obtain a graph convolution result;

[0143] Based on the connection results between the nodes in the graph convolution result, determine the adjacency relationship between each character.

[0144] Specifically, a graph convolutional network can be used for connected component analysis to determine the adjacency relationship. Specifically, the graph convolutional network consists of 4 graph convolutional layers and 1 classification layer, and makes predictions for each node in the connection graph except the central point. The output graph convolutional result is an 8-channel feature map, and the value on each feature map represents the connection relationship between the character node and the central node. Specifically, please refer to Figure 7 , the connection graph contains a central point and several nodes. The connection graph is input into the graph convolutional network for calculation, and the graph convolutional network finally outputs the information of the connection edges between each node and the central point to indicate whether there is a connection relationship between the two.

[0145] The loss function corresponding to the graph convolutional network is a binary cross-entropy loss function. The loss is calculated for each edge separately. The ground truth of each edge is 1, indicating that the node has a connection relationship with the central node, otherwise there is no connection relationship. The expression of the loss function is as follows, where N is the total number of pixel points, xyc represents that the pixel point is located at the (x, y) position of the c-th feature map, here c = 1, Y xyc represents the category of the pixel point, taking values of 0 or 1, represents the probability that the network predicts the pixel point as a positive sample.

[0146]

[0147] In the embodiment of the present application, the adjacency relationship of each character time is specifically calculated by means of a graph convolutional network, avoiding the interference of the threshold value in the character clustering method, thereby reducing the sensitivity to noise and improving the robustness of the solution.

[0148] In the embodiment of the present application, based on the above technical solution, in the above steps, according to the character position results of each character, character recognition is performed on each character in the to-be-recognized image respectively to obtain the character recognition results of each character, which specifically includes the following steps:

[0149] According to the character position results of each character, intercept the character images corresponding to each character from the to-be-recognized image;

[0150] Input the character images corresponding to each character into the character classification model for prediction respectively to obtain the character classification results corresponding to each character, and the character classification results include at least one result character corresponding to the character.

[0151] According to the character position results of each character, the image recognition device can extract the character images corresponding to each character from the image to be recognized. Specifically, the character position results usually include the center point positions and bounding boxes of each character. The screenshot position is determined according to the center point positions, and then the bounding boxes are enlarged according to a predetermined ratio based on the center point positions. Then, screenshots are taken according to the bounding boxes at the center point positions, so that the character images of each character can be obtained. The image recognition device respectively classifies and recognizes the obtained character images, so as to determine the recognition results of each character. Specifically, the image recognition device inputs the character images corresponding to each character into the character classification model for prediction, and obtains the character classification results corresponding to each character. The character classification results include at least one result character corresponding to the character. The character classification model can be specifically implemented by the high-resolution network HRNet18. For the regional position code of the container, the classification results usually include the numbers 0 to 9, the capital English letters A to Z, and the lowercase English letters a to z. The character classification model will classify letters and numbers with very high similarity into the same class. For example, the number 0 and the capital letter O, the lowercase letter o have very high similarity, so the character classification results will include three result characters.

[0152] In the embodiment of the present application, character images are obtained from the image to be recognized by means of screenshot, and the character classification model is used to recognize the character images, so as to obtain the classification results of each character. By separately detecting individual characters, the influence of other characters on character recognition can be avoided, and the accuracy and robustness of recognition can be improved.

[0153] In the embodiment of the present application, based on the above technical solution, before inputting the character images corresponding to each character into the character classification model for prediction to obtain the character classification results corresponding to each character, the method further includes the following steps:

[0154] Training the character classification model to be trained through character training data to obtain training classification results;

[0155] Performing joint calculation of the triplet loss function and the cross-entropy loss function according to the training classification results to obtain training loss results;

[0156] Adjusting the model parameters of the character classification model to be trained according to the training loss results to obtain the character classification model.

[0157] Specifically, the character training data is single-character pictures extracted and labeled according to historical pictures taken actually. The prediction results of the pictures in the character training data are predicted by the character classification model to be trained as the training classification results. Subsequently, joint calculation of the triplet loss function and the cross-entropy loss function is performed according to the training classification results to obtain training loss results. Specifically, the triplet loss function is defined as follows:

[0158] L triple = max(d(a, p) - d(a, n) + margin, 0)

[0159] Where a represents the reference positive example, p represents the positive example belonging to the same category as a, n represents the negative example belonging to a different category from a, d() represents the distance between two samples, (a, p, n) forms a triple, margin represents the distance threshold, and the larger the value of margin, the farther the distance between the positive example and the negative example is greater than the distance between two positive examples, otherwise vice versa.

[0160] The cross-entropy loss function is defined as follows:

[0161]

[0162] Where N represents the total number of samples, x represents the sample, y is the actual value, and m is the output result.

[0163] The triple loss function and the cross-entropy loss function are jointly calculated according to the following formula:

[0164] L = cL triple + dL cross

[0165] Where c and d are the coefficients of the triple loss function and the cross-entropy loss function respectively, and c + d = 1.

[0166] Finally, according to the calculated loss result, the model parameters of the character classification model to be trained are adjusted to obtain the character classification model. Specifically, the training process can be carried out in groups or iteratively. After the result output by the training model meets the training end condition, the training can be ended, thus obtaining the character classification model.

[0167] In the embodiments of the present application, by jointly using the triple loss and the cross-entropy loss, the supervision information of the character category can be utilized, and the distance between classes can be maximally separated, thereby improving the classification effect.

[0168] In the embodiments of the present application, based on the above technical solution, in step S340, according to the character recognition result and the character connection result, the multiple characters are spliced to obtain the text recognition result of the text to be recognized, which specifically includes the following steps:

[0169] According to the adjacency relationship in the character connection result, determine the character sequence of each character;

[0170] Arrange and splice the character recognition results according to the character sequence to obtain the text recognition result of the text to be recognized.

[0171] Specifically, according to the adjacency relationships in the character connection results, the characters can be arranged into a character sequence. For a character with only one adjacent character, it can be determined that it is a character at both ends of the character, while a character with two adjacent characters is an intermediate character in the sequence. According to the adjacency relationships, each character can be arranged into a character sequence. It can be understood that there can be two opposite sorts of the character sequence, or there can be multiple sorting possibilities, and all these sorts are determined as the character sequence. Subsequently, the character recognition results are arranged and spliced according to the character sequence to obtain the text recognition result of the text to be recognized. Specifically, for the determined various character sequences, further judgment can be made according to the positional relationships of the characters in the graph. For example, the positions of the characters in the character recognition results in the image to be recognized are selected for the determined character sequences in the order from top to bottom and from left to right. For example, please refer to Figure 8 , Figure 8 which is a schematic diagram of the image to be recognized in the embodiments of the present application. There are four characters A, B, C, and D in the character recognition results, and in the character connection results, two character sequences can be determined, namely A - B - C - D and D - C - B - A. In Figure 8 , the positions of the four characters in the image to be recognized are that character A is the highest, followed by character B and character C, and character D is the lowest. Then, according to the positions of these four characters, the character sequence can finally be determined as A - B - C - D. The rule for filtering the character sequence according to the character positions can be determined according to the reading order or the writing order.

[0172] In the embodiments of the present application, by arranging and splicing the character recognition results according to the character sequence, the text recognition result of the text to be recognized is obtained, providing a specific solution for text splicing and improving the operability of the solution.

[0173] In the embodiments of the present application, based on the above technical solution, each sequence position in the character sequence corresponds to a candidate character set; for the convenience of introduction, please refer to Figure 9 , Figure 9 which is a schematic flowchart of the content recommendation method in the embodiments of the present application. As shown in Figure 9 , the above steps of arranging and splicing the character recognition results according to the character sequence to obtain the text recognition result of the text to be recognized specifically include the following steps S810 to S830:

[0174] Step S810: For each character, determine the candidate character set corresponding to the character according to the sequence position of the character in the character sequence;

[0175] Step S820: Determine the result character of each character according to the intersection of the character classification result of each character and the determined candidate character set;

[0176] Step S830: Arrange and splice the result characters according to the character sequence to obtain the text recognition result of the text to be recognized.

[0177] In a specific implementation, there are usually composition rules and standards for the text to be recognized. According to these standards, the candidate character set for each character can be determined. For example, taking the regional location code of a container as an example, the regional location code consists of four characters, and there are corresponding requirements for each character. For example, the candidate characters for the first sequence position in the sequence are L, R, O, and D, the candidate characters for the second sequence position are H, T, X, and B, the candidate characters for the third sequence position are numbers within the range of 0-9, and the candidate characters for the fourth sequence position are numbers within the range of 0-9 or N. There will be only one character in each sequence position, and the candidate characters specify the characters that may appear in the corresponding sequence position. Characters other than the candidate characters will not appear in the corresponding sequence position. For example, the character in the first sequence position will only be one of the above four characters, and the character in the third sequence position must be a number. According to the intersection of the character classification result of each character and the determined candidate character set, the result character of each character can be further determined. For example, for the third sequence position, the character classification result indicates that the character can be the number 0, the capital letter O, or the lowercase letter o, and the candidate character set is the numbers from 0 to 9. Therefore, it can be determined that the position is the number 0. And if for the first sequence position, since the candidate characters only include the capital letter O, it can be determined that the first character is the capital letter O. After determining the result characters of each character, the result characters are arranged and spliced according to the character sequence to obtain the text recognition result of the text to be recognized.

[0178] In the embodiments of the present application, by using the candidate character set corresponding to the sequence position to determine the result character of the character and then splicing it into the text recognition result, similar characters can be distinguished, which is beneficial to improving the accuracy of the scheme recognition.

[0179] In the embodiments of the present application, based on the above technical solution, after performing image recognition on the image to be recognized to obtain the character position result of each character, the character connection result of the multiple characters, and the character recognition result of each character, the method further includes the following steps:

[0180] Intercept the picture of the text to be recognized from the image to be recognized according to the character position result and the character connection result;

[0181] Perform field recognition on the picture of the text to be recognized to obtain the field recognition result of the text to be recognized, and the field recognition result includes the recognition result of the multiple characters;

[0182] Modify the text recognition result according to the field recognition result.

[0183] In this embodiment, the text to be recognized in the image to be recognized is also recognized as a whole. Taking the area position code as an example, according to the character connection result, the number of area position codes existing in the text to be recognized can be determined, and the character position result can determine the location of each area position code in the text to be recognized. Therefore, according to the location, pictures of each area position code can be cropped from the image to be recognized. Subsequently, field recognition is performed on the picture of the text to be recognized to obtain the field recognition result of the text to be recognized. The process of field recognition can be executed by a neural network, and the neural network recognizes the whole of the area position code, so as to directly give the field recognition result of each field. Finally, modify the text recognition result according to the field recognition result. Specifically, if there is a conflict between the field recognition result and the text recognition result, the recognition result in the text recognition result can be inferred according to the context relationship in the field recognition result, so as to modify the relevant recognition result. For example, due to writing problems, the first character is determined as L in the text recognition result, but the first character is recognized as I in the field recognition result. However, in the solution of the field recognition result, the text to be recognized is a word starting with I, then the first character can be modified to I to obtain the correct text recognition result.

[0184] In the embodiment of the present application, the recognition result of a single character is modified through the whole-word recognition result, so that the context information to be recognized can be used to correct the result, thereby improving the accuracy of the solution.

[0185] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0186] The device embodiment of the present application is introduced below, which can be used to execute the image recognition method in the above embodiments of the present application. Figure 10 Schematically shows the block diagram of the composition of the image recognition device in the embodiment of the present application. As Figure 10 shown, the image recognition device 900 mainly may include:

[0187] An image acquisition module 910, configured to acquire an image to be recognized including text to be recognized, where the text to be recognized includes a plurality of characters;

[0188] An image recognition module 920 is configured to perform image recognition on the image to be recognized, and obtain a character position result of each character, a character connection result of the multiple characters, and a character recognition result of each character. The character position result is used to indicate the position of the character in the image to be recognized, and the character connection result is used to indicate the adjacency relationship between each character and its adjacent characters;

[0189] A character splicing module 930 is configured to splice the multiple characters according to the character recognition result and the character connection result, and obtain a text recognition result of the text to be recognized.

[0190] In some embodiments of the present application, based on the above technical solution, the character position result includes the center point position of each character, and the character connection result includes a character adjacency matrix for representing the adjacency relationship between characters; the image recognition module 920 includes:

[0191] A downsampling sub-module is configured to perform downsampling on the image to be recognized according to multiple scales, and obtain image features at the multiple scales;

[0192] A feature fusion sub-module is configured to perform feature fusion on the image features at the multiple scales, and obtain a feature fusion result at the multiple scales;

[0193] A position detection sub-module is configured to detect the positions of each character according to the feature fusion result at the multiple scales, and obtain the center point position of each character;

[0194] An adjacency analysis sub-module is configured to analyze the adjacency relationship between each character according to the feature fusion result at the multiple scales and the center point position of each character, and obtain a character adjacency matrix between each character;

[0195] A character recognition module is configured to perform character recognition on each character in the image to be recognized respectively according to the character position result of each character, and obtain a character recognition result of each character.

[0196] In some embodiments of the present application, based on the above technical solution, the position detection sub-module includes:

[0197] A convolution unit is configured to perform convolution processing on the multi-scale feature information, and obtain a convolution result;

[0198] A center feature prediction unit is configured to perform center feature prediction according to the convolution result, and obtain a center point feature map, where each feature value in the center point feature map represents the probability that the corresponding pixel point is the center point of the character;

[0199] A center point determination unit is configured to determine the center point position of each character according to the center point feature map.

[0200] In some embodiments of the present application, based on the above technical solutions, the position detection sub-module further includes:

[0201] A vertex distance prediction unit, configured to perform vertex distance prediction according to the convolution result to obtain a vertex feature map of each character, where each feature value in the vertex feature map represents the distance between the corresponding pixel point in the image to be recognized and each vertex of the character's border;

[0202] A border determination unit, configured to determine the border position of each character according to the vertex feature map.

[0203] In some embodiments of the present application, based on the above technical solutions, the position detection sub-module further includes:

[0204] An offset prediction unit, configured to perform offset prediction according to the convolution result to obtain an offset feature map, where each feature value in the offset feature map represents the offset amount of the corresponding pixel point in the image to be recognized, and the offset amount is used to adjust the border position;

[0205] A border adjustment unit, configured to adjust the border position of each character according to the offset amount.

[0206] In some embodiments of the present application, based on the above technical solutions, the adjacency analysis sub-module includes:

[0207] An angle calculation unit, configured to calculate the rotation angle and character scale of each character in the image to be recognized according to the feature fusion result at multiple scales;

[0208] A position feature determination unit, configured to determine the character position feature of each character according to the center point position of each character, as well as the rotation angle and character scale of each character;

[0209] A similarity calculation unit, configured to calculate the similarity between each character and other characters according to the center point position of each character;

[0210] A connection graph construction unit, configured to construct a connection graph of each character according to the character position feature of each character and the similarity between each character and other characters;

[0211] An adjacency relationship determination unit, configured to determine the adjacency relationship between each character according to the connection graph of each character;

[0212] A matrix construction unit, configured to construct the character adjacency matrix according to the adjacency relationship between each character.

[0213] In some embodiments of the present application, based on the above technical solutions, the adjacency relationship determination unit includes:

[0214] A graph convolution sub - unit, which is used to perform graph convolution operations with the nodes corresponding to the characters as the central points for the connection graph of each character, and obtain graph convolution results;

[0215] An adjacency determination sub - unit, which is used to determine the adjacency relationships between characters according to the connection results between nodes in the graph convolution results.

[0216] In some embodiments of the present application, based on the above technical solutions, the image recognition module 920 includes:

[0217] An image screenshot sub - module, which is used to intercept the character images corresponding to each character from the to - be - recognized image according to the character position results of each character;

[0218] A classification prediction sub - module, which is used to input the character images corresponding to each character into a character classification model for prediction respectively, and obtain the character classification results corresponding to each character. The character classification results include at least one result character corresponding to the character.

[0219] In some embodiments of the present application, based on the above technical solutions, the image recognition module 920 further includes:

[0220] A classification module training sub - module, which is used to train the to - be - trained character classification model through character training data, and obtain training classification results;

[0221] A loss calculation sub - module, which is used to jointly calculate the triplet loss function and the cross - entropy loss function according to the training classification results, and obtain training loss results;

[0222] A parameter adjustment sub - module, which is used to adjust the model parameters of the to - be - trained character classification model according to the training loss results, and obtain the character classification model.

[0223] In some embodiments of the present application, based on the above technical solutions, the character splicing module 930 includes:

[0224] A character sequence determination sub - module, which is used to determine the character sequences of each character according to the adjacency relationships in the character connection results;

[0225] An arrangement splicing sub - module, which is used to arrange and splice the character recognition results according to the character sequences, and obtain the text recognition result of the to - be - recognized text.

[0226] In some embodiments of the present application, based on the above technical solutions, each sequence position in the character sequence corresponds to a candidate character set; the arrangement splicing sub - module includes:

[0227] A candidate set determination unit, configured to determine, for each character, a candidate character set corresponding to the character according to the sequence position of the character in the character sequence;

[0228] A result character determination unit, configured to determine the result character of each character according to the intersection of the character classification result of each character and the determined candidate character set;

[0229] A text splicing unit, configured to arrange and splice the result characters according to the character sequence to obtain a text recognition result of the text to be recognized.

[0230] In some embodiments of the present application, based on the above technical solutions, the image recognition device 900 further includes:

[0231] A field intercepting module, configured to intercept a picture of the text to be recognized from the image to be recognized according to the character position result and the character connectivity result;

[0232] A field recognition module, configured to perform field recognition on the picture of the text to be recognized to obtain a field recognition result of the text to be recognized, where the field recognition result includes recognition results of the multiple characters;

[0233] A result correction module, configured to correct the text recognition result according to the field recognition result.

[0234] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module performs operations have been described in detail in the method embodiment, and will not be elaborated here.

[0235] Figure 11 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application.

[0236] It should be noted that Figure 11 The computer system 1000 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scopes of the embodiments of the present application.

[0237] Such as Figure 11As shown, computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 1002 or programs loaded from a storage section 1008 into a Random Access Memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0238] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.

[0239] In particular, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a Central Processing Unit (CPU) 1001, various functions defined in the system of the present application are executed.

[0240] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0241] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0242] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0243] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0244] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0245] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An image recognition method, characterized in that, Including: Obtain a to-be-recognized image including to-be-recognized text, where the to-be-recognized text includes multiple characters; Perform downsampling on the to-be-recognized image according to multiple scales to obtain image features at the multiple scales; Perform feature fusion on the image features at the multiple scales to obtain feature fusion results at the multiple scales; Detect the positions of each character according to the feature fusion results at the multiple scales, obtain the center point positions of each character, and obtain the character position results of each character; According to the feature fusion results at the multiple scales, calculate the rotation angle and character scale of each character in the to-be-recognized image; Determine the character position features of each character according to the center point positions of each character and the rotation angle and character scale of each character; Calculate the similarity between each character and other characters according to the center point positions of each character; Construct a connection graph for each character according to the character position features of each character and the similarity between each character and other characters; Determine the adjacency relationship between each character according to the connection graph of each character; Construct the character adjacency matrix according to the adjacency relationship between each character to obtain the character connectivity result of the multiple characters; Perform character recognition on each character in the to-be-recognized image respectively according to the character position results of each character to obtain the character recognition results of each character; Stitch the multiple characters according to the character recognition results and the character connectivity result to obtain the text recognition result of the to-be-recognized text.

2. The method according to claim 1, wherein The detecting the positions of each character according to the feature fusion results at the multiple scales to obtain the center point positions of each character includes: Perform convolution processing on the multi-scale feature information to obtain a convolution result; Perform central feature prediction according to the convolution result to obtain a center point feature map, where each feature value in the center point feature map represents the probability that the corresponding pixel point is the center point of a character; Determine the center point positions of each character according to the center point feature map.

3. The method according to claim 2, wherein After performing convolution processing on the multi-scale feature information to obtain a convolution result, the method further includes: Perform vertex distance prediction according to the convolution result to obtain a vertex feature map of each character, where each feature value in the vertex feature map represents the distance between the corresponding pixel point in the to-be-recognized image and each vertex of the character's border; Determine the border positions of each character according to the vertex feature map.

4. The method according to claim 3, wherein After performing convolution processing on the multi-scale feature information to obtain a convolution result, the method further includes: Perform offset prediction according to the convolution result to obtain an offset feature map, where each feature value in the offset feature map represents the offset amount of the corresponding pixel point in the to-be-recognized image, and the offset amount is used to adjust the border position; Adjust the border positions of each character according to the offset amount.

5. The method according to claim 1, characterized in that, The determining the adjacency relationship between each character according to the connection graph of each character includes: For the connection graph of each character, perform graph convolution operation with the node corresponding to the character as the center point to obtain a graph convolution result; Determine the adjacency relationship between each character according to the connection result between the nodes in the graph convolution result.

6. The method according to claim 1, characterized in that, Performing character recognition on each character in the to-be-recognized image respectively according to the character position results of each character, to obtain the character recognition results of each character, including: Intercepting the character images corresponding to each character from the to-be-recognized image according to the character position results of each character; Inputting the character images corresponding to each character into a character classification model for prediction respectively, to obtain the character classification results corresponding to each character, where the character classification results include at least one result character corresponding to the character.

7. The method according to claim 6, characterized in that, Before the step of inputting the character images corresponding to each character into a character classification model for prediction respectively to obtain the character classification results corresponding to each character, the method further includes: Training the to-be-trained character classification model through character training data to obtain training classification results; Performing joint calculation of a triplet loss function and a cross-entropy loss function according to the training classification results to obtain training loss results; Adjusting the model parameters of the to-be-trained character classification model according to the training loss results to obtain the character classification model.

8. The method according to claim 1, wherein Performing splicing on the multiple characters according to the character recognition results and the character connection results to obtain the text recognition result of the to-be-recognized text, including: Determining the character sequences of each character according to the adjacency relationship in the character connection results; Arranging and splicing the character recognition results according to the character sequences to obtain the text recognition result of the to-be-recognized text.

9. The method according to claim 8, characterized in that Each sequence position in the character sequence corresponds to a candidate character set; the step of arranging and splicing the character recognition results according to the character sequences to obtain the text recognition result of the to-be-recognized text includes: For each character, determining the candidate character set corresponding to the character according to the sequence position of the character in the character sequence; Determining the result character of each character according to the intersection of the character classification result of each character and the determined candidate character set; Arranging and splicing the result characters according to the character sequences to obtain the text recognition result of the to-be-recognized text.

10. The method according to any one of claims 1 to 9, characterized in that, After performing image recognition on the to-be-recognized image to obtain the character position results of each character, the character connection results of the multiple characters, and the character recognition results of each character, the method further includes: Intercepting the picture of the to-be-recognized text from the to-be-recognized image according to the character position results and the character connection results; Performing field recognition on the picture of the to-be-recognized text to obtain the field recognition result of the to-be-recognized text, where the field recognition result includes the recognition results of the multiple characters; Correcting the text recognition result according to the field recognition result.

11. An image recognition device, characterized in that, Including: An image acquisition module, configured to acquire a to-be-recognized image including a to-be-recognized text, where the to-be-recognized text includes multiple characters; A downsampling sub-module, configured to perform downsampling on the to-be-recognized image according to multiple scales to obtain image features at the multiple scales; A feature fusion sub-module, configured to perform feature fusion on the image features at the multiple scales to obtain the feature fusion results at the multiple scales; A position detection sub-module, configured to detect the positions of each character according to the feature fusion results at the multiple scales, obtain the center point positions of each character, and obtain the character position results of each character; An angle calculation unit, configured to calculate the rotation angle and character scale of each character in the image to be recognized according to the feature fusion results at the multiple scales; A position feature determination unit, configured to determine the character position features of each character according to the center point positions of each character, as well as the rotation angle and character scale of each character; A similarity calculation unit, configured to calculate the similarity between each character and other characters according to the center point positions of each character; A connection graph construction unit, configured to construct a connection graph for each character according to the character position features of each character and the similarity between each character and other characters; An adjacency relationship determination unit, configured to determine the adjacency relationship between each character according to the connection graph of each character; A matrix construction unit, configured to construct the character adjacency matrix according to the adjacency relationship between each character, and obtain the character connectivity results of the multiple characters; A character recognition module, configured to perform character recognition on each character in the image to be recognized respectively according to the character position results of each character, and obtain the character recognition results of each character; A character splicing module, configured to splice the multiple characters according to the character recognition results and the character connectivity results, and obtain the text recognition result of the image to be recognized.

12. An electronic device, characterized in that, Comprising: A processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the image recognition method according to any one of claims 1 to 10 by executing the executable instructions.

13. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image recognition method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and the processor of the computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device executes the image recognition method according to any one of claims 1 to 10.