Text recognition model training method and device, recognition method and device and electronic equipment

By encoding and decoding the image to be recognized, determining the location of the target pixel point, and decoding the text in combination with image features, the problem of poor text recognition accuracy in the prior art is solved, and high-accuracy text recognition is achieved.

CN120279565APending Publication Date: 2025-07-08SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410035359.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the text recognition method of an image has poor text recognition accuracy due to the direct recognition of the image area.

Method used

By encoded the image to be recognized, image features are acquired, and position decoded to determine the position of the target pixel point, text decoding is performed in combination with the target position features and image features to identify the text content in the image.

Benefits of technology

Improve the accuracy of text recognition and realize accurate text recognition from the smallest recognition unit pixel points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279565A_ABST
    Figure CN120279565A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device of a text recognition model, a recognition method and device of the text recognition model, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-recognized image including text content, and performing image coding on the to-be-recognized image to obtain image features of the to-be-recognized image; based on the image features, performing position decoding on the to-be-recognized image to obtain at least one target pixel point position in the to-be-recognized image; performing position coding on the positions of the target pixel points to obtain target position features, and performing text decoding on the to-be-recognized image in combination with the target position features and the image features to obtain text content in the to-be-recognized image; and determining a text recognition result of the to-be-recognized image in combination with the text content and the target pixel point position. According to the invention, the accuracy of text recognition can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a training method, an identification method, a device, and an electronic device for a text recognition model. Background Art

[0002] Artificial Intelligence (AI) is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The artificial intelligence software technologies mainly include several directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0003] In related technologies, for text recognition of images, usually, the image area with text in the image to be recognized is directly recognized, and then the text content in the image to be recognized is recognized from the image area. Since the range of the image area is broad, the accuracy of text recognition is affected, resulting in poor accuracy of text recognition. Summary of the Invention

[0004] Embodiments of this application provide a training method, an identification method, a device, an electronic device, a computer-readable storage medium, and a computer program product for a text recognition model, which can effectively improve the accuracy of text recognition.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a text recognition method, including:

[0007] Obtain an image to be recognized including text content, and perform image encoding on the image to be recognized to obtain the image features of the image to be recognized;

[0008] Based on the image features, perform position decoding on the image to be recognized to obtain at least one target pixel position in the image to be recognized, where the target pixel position is used to indicate the position of the target pixel with text in the image to be recognized;

[0009] Perform position encoding on each target pixel position to obtain target position features, and combine the target position features and the image features to perform text decoding on the image to be recognized to obtain the text content in the image to be recognized;

[0010] Determine the text recognition result of the image to be recognized by combining the text content and the position of the target pixel point.

[0011] An embodiment of the present application provides a method for training a text recognition model, including:

[0012] Obtain an image sample to be recognized and at least one sample label carried by the image sample to be recognized. The sample label is used to indicate the label text content corresponding to the position of the target pixel point in the image sample to be recognized;

[0013] Among them, the position of the target pixel point is used to indicate the position of the target pixel point with text in the image sample to be recognized, and the number of sample labels is less than the total number of words in the text of the image sample to be recognized;

[0014] Call the initial text recognition model to perform text recognition on the image sample to be recognized, and obtain the text recognition result of the image sample to be recognized;

[0015] Train the initial text recognition model by combining the sample label and the text recognition result to obtain the text recognition model, which is used to recognize the text content in the image to be recognized.

[0016] An embodiment of the present application provides a text recognition device, including:

[0017] An image encoding module, configured to obtain an image to be recognized including text content, and perform image encoding on the image to be recognized to obtain the image feature of the image to be recognized;

[0018] A position decoding module, configured to perform position decoding on the image to be recognized based on the image feature to obtain at least one target pixel point position in the image to be recognized. The position of the target pixel point is used to indicate the position of the target pixel point with text in the image to be recognized;

[0019] A position encoding module, configured to perform position encoding on each target pixel point position to obtain a target position feature, and combine the target position feature and the image feature to perform text decoding on the image to be recognized to obtain the text content in the image to be recognized;

[0020] A determination module, configured to determine the text recognition result of the image to be recognized by combining the text content and the position of the target pixel point.

[0021] In the above solution, the position decoding is implemented through a position decoding network. The position decoding network includes a position recognition layer and a position correction layer. The image features include the pixel point features corresponding to each pixel point in the image to be recognized. The above position decoding module is further configured to call the position recognition layer, and based on the image features, perform position recognition on the image to be recognized to obtain at least one candidate pixel point position in the image to be recognized; call the position correction layer, and based on the pixel point features corresponding to each candidate pixel point position, perform position correction on each candidate pixel point position to obtain at least one target pixel point position in the image to be recognized.

[0022] In the above solution, the above position decoding module is further configured to, when the number of candidate pixel point positions is one, determine the candidate pixel point position as the target pixel point position; when the number of candidate pixel point positions is multiple, call the position correction layer, and based on the pixel point features corresponding to each candidate pixel point position, perform position correction on each candidate pixel point position to obtain the correction information corresponding to each candidate pixel point position; for each candidate pixel point position, when the correction information corresponding to the candidate pixel point position indicates that the candidate pixel point position is the position where text exists in the image to be recognized, determine the candidate pixel point position as the target pixel point position.

[0023] In the above solution, the above text decoding is implemented through a text decoding network. The text decoding network includes a fusion layer and a text decoding layer. The above position encoding module is further configured to call the fusion layer, fuse the target position features and the image features to obtain fusion features; call the text decoding layer, and based on the fusion features, perform text decoding on the image to be recognized to obtain the text content in the image to be recognized.

[0024] In the above solution, the above fusion layer includes a position fusion layer and a feature fusion layer. The above position encoding module is further configured to, for each target pixel point position, determine the pixel point position in the image to be recognized whose distance from the target pixel point position is less than a distance threshold as the reference pixel point position corresponding to the target pixel point position; perform position encoding on the reference pixel point position to obtain reference position features; and call the position fusion layer to perform position fusion on the target position features and the reference position features to obtain position fusion features; call the feature fusion layer to fuse the position fusion features and the image features to obtain the fusion features.

[0025] In the above solution, the above-mentioned determination module is further configured to aggregate the positions of the target pixel points to obtain at least one target position group, where the distances between the positions of the target pixel points in the target position group are less than a preset distance threshold; and determine the text recognition result of the image to be recognized by combining the text content and the target position group.

[0026] In the above solution, the text content includes the sub-text contents respectively corresponding to the positions of the target pixel points. The above-mentioned determination module is further configured to perform position aggregation on the positions of the target pixel points to obtain at least one candidate position group, where the distances between the positions of the target pixel points in the candidate position group are less than a preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the target position group; when the number of candidate position groups is multiple, select at least one target position group from the multiple candidate position groups based on the sub-text contents.

[0027] In the above solution, the above-mentioned determination module is further configured to perform the following processing for each of the candidate position groups: perform text content semantic analysis on the sub-text contents respectively corresponding to the positions of the target pixel points in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group.

[0028] In the above solution, the text content includes the sub-text contents respectively corresponding to the positions of the target pixel points, and the text recognition result includes the sub-text recognition results respectively corresponding to the target position groups. The above-mentioned determination module is further configured to perform the following processing for each of the target position groups: when the number of the target pixel points in the target position group is one, determine the sub-text content corresponding to the target pixel point position as the sub-text recognition result corresponding to the target position group; when the number of the target pixel points in the target position group is multiple, perform text fusion on the sub-text contents respectively corresponding to the target pixel point positions in the target position group to obtain the sub-text recognition result corresponding to the target position group.

[0029] An embodiment of the present application provides a training device for a text recognition model, including: an acquisition module, configured to acquire an image sample to be recognized and at least one sample label carried by the image sample to be recognized, where the sample label is used to indicate the label text content corresponding to the position of the target pixel point in the image sample to be recognized; wherein, the position of the target pixel point is used to indicate the position where the target pixel points with text in the image sample to be recognized are located, and the number of sample labels is less than the total number of characters in the text in the image sample to be recognized; a recognition module, configured to call an initial text recognition model to perform text recognition on the image sample to be recognized to obtain a text recognition result of the image sample to be recognized; a training module, configured to train the initial text recognition model by combining the sample label and the text recognition result to obtain the text recognition model, where the text recognition model is used to recognize the text content in the image to be recognized.

[0030] In the above solution, the initial text recognition model includes an image encoding network, a position decoding network, and a text decoding network. The above recognition module is further configured to call the image encoding network to perform image encoding on the image sample to be recognized to obtain a sample image feature of the image sample to be recognized; call the position decoding network to perform position decoding on the image sample to be recognized based on the sample image feature to obtain at least one predicted pixel point position in the image to be recognized; perform position encoding on each of the predicted pixel point positions to obtain a predicted position feature, and call the text decoding network to perform text decoding on the image sample to be recognized to obtain a predicted text content in the image sample to be recognized; aggregate each of the predicted pixel point positions to obtain at least one predicted position group, and determine the text recognition result of the image sample to be recognized by combining the text content and the predicted position group.

[0031] In the above solution, the text recognition result includes sub-text recognition results corresponding to each predicted position group. The above training module is further configured to, for each of the sub-text recognition results, obtain the similarity between the text content indicated by the sub-text recognition result and each of the label text contents, and determine the loss value corresponding to the sub-text recognition result based on the maximum similarity; train the initial text recognition model based on the loss value to obtain the text recognition model.

[0032] In the above solution, the above training module is further configured to, when the number of sub-text recognition results is one, determine the loss value corresponding to the sub-text recognition result as the target loss value; when the number of sub-text recognition results is multiple, sum the loss values corresponding to each of the sub-text recognition results to obtain the target loss value; train the initial text recognition model based on the target loss value to obtain the text recognition model.

[0033] An embodiment of the present application provides an electronic device, including:

[0034] A memory for storing computer-executable instructions or computer programs;

[0035] A processor, when executing the computer-executable instructions or computer programs stored in the memory, implements the text recognition method and the training method of the text recognition model provided by the embodiments of the present application.

[0036] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to cause a processor to implement the text recognition method and the training method of the text recognition model provided by the embodiments of the present application when executed.

[0037] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions. The computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the text recognition method and the training method of the text recognition model as described above in the embodiments of the present application.

[0038] The embodiments of the present application have the following beneficial effects:

[0039] By performing image encoding on the image to be recognized, the image features of the image to be recognized are obtained. Based on the image features, position decoding is performed on the image to be recognized to obtain the position of the target pixel points in the image to be recognized, and position encoding is performed on the position of the target pixel points to obtain the target position features. Combining the target position features and the image features, text encoding is performed on the image to be recognized to obtain the text content in the image to be recognized. Combining the text content and the position of the target pixel points, the text recognition result of the image to be recognized is determined. In this way, by determining the target position features indicating the position of the target pixel points where text exists in the image to be recognized, and combining the target position features and the image features, the text content in the image to be recognized is decoded. Since the pixel points are the smallest recognition units of the image to be recognized, the text content is determined by the target position features indicating the position of the target pixel points where text exists in the image to be recognized, thereby realizing text recognition from the smallest recognition granularity, and effectively improving the accuracy of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic structural diagram of the text recognition system provided by the embodiments of the present application;

[0041] Figure 2 is a schematic structural diagram of the electronic device 500 for text recognition provided by the embodiments of the present application;

[0042] Figure 3 It is a schematic structural diagram of an electronic device 600 provided by an embodiment of the present application for training a text recognition model;

[0043] Figure 4 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 1 ;

[0044] Figure 5 It is a schematic principle diagram of a text recognition method provided by an embodiment of the present application;

[0045] Figure 6 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 2 ;

[0046] Figure 7 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 3 ;

[0047] Figure 8 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 4 ;

[0048] Figure 9 It is a schematic flow chart of a training method of a text recognition model provided by an embodiment of the present application Figure 1 ;

[0049] Figure 10 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 5 ;

[0050] Figure 11 It is a schematic flow chart of a text recognition method provided by an embodiment of the present application Figure 6 ;

[0051] Figure 12 It is a schematic principle diagram of the annotation principle of sample labels provided by an embodiment of the present application;

[0052] Figure 13 It is a schematic principle diagram of a training method of a text recognition model provided by an embodiment of the present application;

[0053] Figure 14 It is a schematic diagram of the effect of point annotation provided by an embodiment of the present application;

[0054] Figure 15 It is a schematic diagram of the text recognition effect of a text recognition method provided by an embodiment of the present application. Detailed implementation manners

[0055] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0056] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0057] In the following description, the terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0059] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0060] 1) Artificial Intelligence (AI): It is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0061] 2) Machine Learning (ML): It is an interdisciplinary subject involving multiple fields, including probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0062] 3) Convolutional Neural Networks (CNN): It is a type of feed-forward neural network (FNN) with convolutional calculations and a deep structure, and it is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input images according to their hierarchical structure.

[0063] 4) In response to: It is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or have a set delay; without special instructions, there is no restriction on the execution order of multiple executed operations.

[0064] 5) Convolutional layer: Each convolutional layer in a convolutional neural network consists of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of convolutional operations is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. More layers of the network can iteratively extract more complex features from low-level features.

[0065] 6) Pooling layer: After feature extraction in the convolutional layer, the output feature map will be passed to the pooling layer for feature selection and information filtering. The pooling layer contains a preset pooling function, and its function is to replace the result of a single point in the feature map with the feature map statistic of its adjacent area. The pooling layer selects the pooling area in the same way as the convolutional kernel scans the feature map, which is controlled by the pooling size, stride, and padding.

[0066] 7) Fully-Connected Layer: The fully-connected layer in a convolutional neural network is equivalent to the hidden layer in a traditional feedforward neural network. The fully-connected layer is located at the end of the hidden layer of the convolutional neural network and only transmits signals to other fully-connected layers. The feature map loses its spatial topology in the fully-connected layer, is unfolded into a vector, and passes through an activation function.

[0067] In the implementation process of the embodiments of the present application, the applicant found the following problems in the related art:

[0068] In the related art, for text recognition of an image, usually, the image area with text in the image to be recognized is directly recognized, and then the text content in the image to be recognized is recognized from the image area. Since the range of the image area is broad, the accuracy of text recognition is affected, resulting in poor accuracy of text recognition.

[0069] The embodiments of the present application provide a text recognition method, a training method for a text recognition model, a device, an electronic device, a computer-readable storage medium, and a computer program product, which can effectively improve the accuracy of text recognition. The following describes an exemplary application of the text recognition system provided by the embodiments of the present application.

[0070] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the text recognition system 100 provided by the embodiments of the present application. The terminal (exemplarily shows the terminal 400) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0071] The terminal 400 is used for the user to use the client 410 to display the text recognition result on the graphical interface 410-1 (exemplarily shows the graphical interface 410-1). The terminal 400 and the server 200 are connected to each other through a wired or wireless network.

[0072] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The electronic device provided in the embodiments of the present application may be implemented as a terminal or as a server. The terminal and the server may be directly or indirectly connected by wired or wireless communication methods, and no limitation is made in the embodiments of the present application.

[0073] In some embodiments, the server 200 obtains an image to be recognized including text content, performs text recognition on the image to be recognized to obtain a text recognition result, and sends the text recognition result to the terminal 400.

[0074] In some other embodiments, the terminal 400 obtains an image to be recognized including text content, performs text recognition on the image to be recognized to obtain a text recognition result, and sends the text recognition result to the server 200.

[0075] In some other embodiments, the embodiments of the present application may be implemented by means of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve the calculation, storage, processing, and sharing of data.

[0076] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, and application technology applied based on the cloud computing business model. It can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.

[0077] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device 500 for text recognition provided by the embodiments of the present application, where Figure 2 the shown electronic device 500 may be Figure 1 the server 200 or the terminal 400 in Figure 2The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. Each component in the electronic device 500 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to achieve connection communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0078] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0079] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. Optionally, the memory 450 includes one or more storage devices that are physically located far from the processor 430.

[0080] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0081] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0082] An operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0083] A network communication module 452, for reaching other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0084] In some embodiments, the text recognition device provided by the embodiments of the present application can be implemented in software. Figure 2 FIG. shows a text recognition device 455 stored in a memory 450, which can be software in the form of a program and a plug-in, etc., including the following software modules: an image encoding module 4551, a position decoding module 4552, a position encoding module 4553, and a determination module 4554. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0085] See Figure 3 , Figure 3 FIG. is a schematic structural diagram of an electronic device 600 for training a text recognition model provided by an embodiment of the present application. Among them, Figure 3 the shown electronic device 600 can be Figure 1 the server 200 or the terminal 400 in Figure 3 The shown electronic device 600 includes: at least one processor 530, a memory 550, and at least one network interface 520. Each component in the electronic device 600 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 3 all kinds of buses are labeled as the bus system 540.

[0086] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0087] The memory 550 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disk drives, etc. The memory 550 optionally includes one or more storage devices physically located far from the processor 530.

[0088] The memory 550 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0089] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.

[0090] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0091] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0092] In some embodiments, the training device of the text recognition model provided in the embodiments of the present application can be implemented in software. Figure 3 Shown is the training device 555 of the text recognition model stored in the memory 550, which can be software in the form of programs and plugins, etc., including the following software modules: the acquisition module 5551, the recognition module 5552, and the training module 5553. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0093] In other embodiments, the text recognition device provided in the embodiments of the present application can be implemented in hardware. As an example, the text recognition device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the text recognition method provided in the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more Application Specific Integrated Circuits (ASICs), DSPs, Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), or other electronic components.

[0094] In some embodiments, a terminal or a server may implement the text recognition method and the training method of the text recognition model provided in the embodiments of the present application by running a computer program or computer-executable instructions. For example, the computer program may be a native program in an operating system (for example, a dedicated text recognition program) or a software module. For example, it may be a text recognition module embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it may be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above computer program may be any form of application program, module or plug-in.

[0095] The text recognition method provided in the embodiments of the present application will be described in conjunction with the exemplary applications and implementations of the server or terminal provided in the embodiments of the present application.

[0096] See Figure 4 , Figure 4 is a flowchart of the text recognition method provided in the embodiments of the present application Figure 1 will be described in conjunction with Figure 4 Steps 101 to 106 shown below. The text recognition method provided in the embodiments of the present application may be implemented independently by a server or a terminal, or jointly implemented by a server and a terminal. The following will take the independent implementation by the server as an example for description.

[0097] In step 101, an image to be recognized including text content is obtained.

[0098] In some embodiments, the image to be recognized includes text content. Through the text recognition method provided in the embodiments of the present application, the text content in the image to be recognized can be recognized.

[0099] As an example, see Figure 5 , Figure 5 is a schematic diagram of the principle of the text recognition method provided in the embodiments of the present application, Figure 5 The image to be recognized 51 shown includes the text content 52 "FLYINGPIG".

[0100] In step 102, the image to be recognized is encoded to obtain an image feature of the image to be recognized.

[0101] In some embodiments, the above image encoding refers to a processing process of converting the image to be recognized into a vector to obtain an image feature in vector form.

[0102] In some embodiments, the above image encoding can be implemented by an image encoding network, and the image encoding network can be any artificial intelligence model with image encoding capabilities. For example, it can be a Transformer network or the like. Then, step 102 can be implemented as follows: call the image encoding network to perform image encoding on the image to be recognized, and obtain the image features of the image to be recognized.

[0103] As an example, refer to Figure 5 , call the image encoding network 53 to perform image encoding on the image to be recognized 51, and obtain the image features 54 of the image to be recognized 51.

[0104] In step 103, based on the image features, perform position decoding on the image to be recognized to obtain the positions of at least one target pixel point in the image to be recognized.

[0105] In some embodiments, the target pixel point position is used to indicate the position of the target pixel point with text in the image to be recognized.

[0106] In some embodiments, the above position decoding is used to identify the position of the target pixel point in the image to be recognized. The pixel points of the image to be recognized include target pixel points and other pixel points. The target pixel points are the pixel points with text, and the other pixel points are the pixel points without text.

[0107] In some embodiments, the above position decoding can identify the positions of at least some of the target pixel points in the image to be recognized.

[0108] In some embodiments, refer to Figure 6 , Figure 6 is a schematic flow chart of the text recognition method provided by an embodiment of the present application Figure 2 , the position decoding is implemented by a position decoding network. The position decoding network includes a position recognition layer and a position correction layer. The image features include the pixel point features corresponding to each pixel point in the image to be recognized. Figure 4 The step 103 shown can be implemented by Figure 6 the steps 1031 to 1032 shown.

[0109] In step 1031, call the position recognition layer to perform position recognition on the image to be recognized based on the image features, and obtain at least one candidate pixel point position in the image to be recognized.

[0110] In some embodiments, the above candidate pixel point position is the pixel point position where text may exist in the image to be recognized. The above position recognition layer can be any position recognition network based on artificial intelligence. The position recognition layer is used to recognize the pixel point position where text may exist in the image to be recognized.

[0111] In some embodiments, the above-mentioned image features include pixel point features respectively corresponding to each pixel point in the image to be recognized. The above-mentioned position recognition layer is called to perform position recognition on the image to be recognized based on the image features, and at least one candidate pixel point position in the image to be recognized is obtained. This can be achieved in the following manner: For each pixel point in the image to be recognized, the position recognition layer is called, and pixel point recognition is performed on the pixel point based on the pixel point feature of the pixel point to obtain a pixel point recognition result. When the pixel point recognition result indicates that the pixel point is a pixel point where text may exist, the position of the pixel point in the image to be recognized is determined as the candidate pixel point position.

[0112] As an example, referring to Figure 5 , the position recognition layer 55 is called to perform position recognition on the image to be recognized 51 based on the image feature 54, and at least one candidate pixel point position 56 in the image to be recognized 51 is obtained.

[0113] In step 1032, the position correction layer is called to perform position correction on each candidate pixel point position based on the pixel point feature respectively corresponding to each candidate pixel point position, and at least one target pixel point position in the image to be recognized is obtained.

[0114] In some embodiments, the above-mentioned position correction layer may be any position correction network based on artificial intelligence, and the position correction layer is used to recognize the positions of pixel points where text may exist in the image to be recognized.

[0115] In some embodiments, the above-mentioned image features include pixel point features respectively corresponding to each pixel point in the image to be recognized, and the pixel point feature corresponding to the candidate pixel point position is the pixel point feature corresponding to the candidate pixel point.

[0116] As an example, referring to Figure 5 , the position correction layer 57 is called to perform position correction on each candidate pixel point position 56 based on the pixel point feature respectively corresponding to each candidate pixel point position 56, and at least one target pixel point position in the image to be recognized is obtained.

[0117] In some embodiments, the above-mentioned step 1032 can be implemented in the following manner: When the number of candidate pixel point positions is one, the candidate pixel point position is determined as the target pixel point position; when the number of candidate pixel point positions is multiple, the position correction layer is called to perform position correction on each candidate pixel point position based on the pixel point feature respectively corresponding to each candidate pixel point position, and correction information corresponding to each candidate pixel point position is obtained; for each candidate pixel point position, when the correction information corresponding to the candidate pixel point position indicates that the candidate pixel point position is a position where text exists in the image to be recognized, the candidate pixel point position is determined as the target pixel point position.

[0118] In some embodiments, the above correction information is used to indicate whether the candidate pixel point position is a position where text exists in the image to be recognized. When the number of candidate pixel point positions is multiple, for each candidate pixel point position, a position correction layer is called, and based on the pixel point features corresponding to the candidate pixel point position, the position of the candidate pixel point position is corrected to obtain the correction information corresponding to the candidate pixel point position. The correction information is used to indicate whether the candidate pixel point position is a position where text exists in the image to be recognized.

[0119] In this way, by calling the position correction layer and based on the pixel point features corresponding to each candidate pixel point position respectively, the positions of each candidate pixel point position are corrected to obtain at least one target pixel point position in the image to be recognized, thereby realizing the position correction of the candidate pixel point position and making the obtained target pixel point position more accurate.

[0120] In step 104, position encoding is performed on each target pixel point position to obtain target position features.

[0121] In some embodiments, the above position encoding is used to convert the target pixel point position in coordinate form into a target position feature in vector form. The above target position feature includes sub-position features corresponding to each target pixel point position respectively.

[0122] In some embodiments, the above position encoding can be implemented through a position encoding network, and the position encoding network can be any artificial intelligence model with a position encoding function. Then step 104 can be implemented in the following manner: for each target pixel point position, a position encoding network is called to perform position encoding on the target pixel point position to obtain the sub-position feature corresponding to the target pixel point position. The sub-position features are fused to obtain the target position feature.

[0123] In step 105, combining the target position feature and the image feature, text decoding is performed on the image to be recognized to obtain the text content in the image to be recognized.

[0124] In some embodiments, text decoding is implemented through a text decoding network. The text decoding network includes a fusion layer and a text decoding layer. Refer to Figure 7 , Figure 7 which is the flowchart of the text recognition method provided by the embodiments of the present application Figure 3 , Figure 4 shown, step 105 can be implemented through Figure 7 steps 1051 to 1052 shown.

[0125] As an example, refer to Figure 5 , text decoding is implemented through a text decoding network 58. The text decoding network 58 includes a fusion layer 60 and a text decoding layer 59.

[0126] In step 1051, the fusion layer is called to fuse the target position feature and the image feature to obtain a fused feature.

[0127] In some embodiments, the above-mentioned fusion layer includes a position fusion layer and a feature fusion layer. The above step 1051 can be implemented in the following manner: for each target pixel position, the pixel positions in the to-be-recognized image whose distance from the target pixel position is less than the distance threshold are determined as the reference pixel positions corresponding to the target pixel position; position encoding is performed on the reference pixel positions to obtain reference position features; and the position fusion layer is called to perform position fusion on the target position feature and the reference position feature to obtain a position fusion feature; the feature fusion layer is called to fuse the position fusion feature and the image feature to obtain a fused feature.

[0128] In some embodiments, for each target pixel position, since there is a high probability of text in the image area around the target pixel in the to-be-recognized image, the pixel positions in the to-be-recognized image whose distance from the target pixel position is less than the distance threshold can be determined as the reference pixel positions corresponding to the target pixel position, and the position fusion of the target position feature and the reference position feature is performed. The obtained position fusion feature can more comprehensively cover all the text-containing areas in the to-be-recognized image, thereby making the characterization accuracy of the obtained position fusion feature higher.

[0129] In some embodiments, the above-mentioned position fusion layer can be any artificial intelligence model with a feature fusion function. The above-mentioned feature fusion layer can also be any artificial intelligence model with a feature fusion function.

[0130] In some embodiments, the above-mentioned position fusion of the target position feature and the reference position feature to obtain a position fusion feature can be implemented in the following manner: the target position feature and the reference position feature are feature concatenated to obtain a position fusion feature.

[0131] In step 1052, the text decoding layer is called to perform text decoding on the to-be-recognized image based on the fused feature to obtain the text content in the to-be-recognized image.

[0132] In some embodiments, the above-mentioned text decoding is used to recognize the text content in the to-be-recognized image. The above-mentioned text decoding layer can be any artificial intelligence-based text decoding network, and the text decoding layer is used to recognize the text content in the to-be-recognized image.

[0133] Thus, for each target pixel position, since there is a high probability of text in the image area around the target pixel in the image to be recognized, the pixel positions in the image to be recognized whose distance from the target pixel position is less than the distance threshold can be determined as the reference pixel positions corresponding to the target pixel position, and the target position feature and the reference position feature are fused in position. The resulting position fusion feature can more comprehensively cover all the text-containing regions in the image to be recognized, so that the representation accuracy of the resulting fusion feature is higher. By using the fusion feature with higher representation accuracy to perform text decoding on the image to be recognized, the resulting text content is more accurate.

[0134] In step 106, combining the text content and the target pixel positions to determine the text recognition result of the image to be recognized.

[0135] In some embodiments, the above text content includes the sub-text content corresponding to each target pixel position respectively, and the text recognition result includes the sub-text recognition result corresponding to each target position group respectively.

[0136] In some embodiments, referring to Figure 8 , Figure 8 is the flowchart of the text recognition method provided by the embodiments of the present application Figure 4 , Figure 4 As shown, step 106 can be implemented by Figure 8 steps 1061 to 1062 shown in

[0137] In step 1061, aggregate the target pixel positions to obtain at least one target position group.

[0138] In some embodiments, the above target position group includes at least one target pixel position, and the distance between the target pixel positions in the target position group is less than the preset distance threshold.

[0139] In some embodiments, the text content includes the sub-text content corresponding to each target pixel position respectively. The above step 1061 can be implemented in the following manner: perform position aggregation on the target pixel positions to obtain at least one candidate position group, where the distance between the target pixel positions in the candidate position group is less than the preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the target position group; when the number of candidate position groups is multiple, select at least one target position group from the multiple candidate position groups based on the sub-text content.

[0140] As an example, perform position aggregation on the target pixel point positions A, B, and C to obtain a candidate position group {A, B, C}. If the number of the obtained candidate position groups is one, then directly determine the candidate position group {A, B, C} as the target position group.

[0141] As an example, perform position aggregation on the target pixel point positions A, B, and C to obtain a candidate position group {A, B} and a candidate position group {C}. At this time, the number of candidate position groups is multiple. Based on the content of each sub-text, select at least one target position group from the candidate position groups {A, B} and {C}.

[0142] In some embodiments, the above-mentioned selecting at least one target position group from multiple candidate position groups based on the content of each sub-text can be implemented in the following manner: perform the following processing for each candidate position group respectively: perform text content semantic analysis on the sub-text content corresponding to each target pixel point position in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group.

[0143] In some embodiments, the above-mentioned semantic analysis result is used to indicate whether the sub-text contents in the candidate position group can form a complete semantics.

[0144] In some embodiments, the above-mentioned performing text content semantic analysis on the sub-text content corresponding to each target pixel point position in the candidate position group to obtain a semantic analysis result can be implemented in the following manner: perform text combination on the sub-text content corresponding to each target pixel point position in the candidate position group to obtain at least one combined text corresponding to the candidate position group, and perform text content semantic analysis on each combined text to obtain a combined semantic analysis result corresponding to each combined text; when there is a combined semantic analysis result indicating that the combined text can form a complete semantics, determine the semantic analysis result of the candidate position group as the first semantic analysis result; when all the combined semantic analysis results indicate that the combined text cannot form a complete semantics, determine the semantic analysis result of the candidate position group as the second semantic analysis result.

[0145] In some embodiments, the above-mentioned first semantic analysis result is used to indicate that the sub-text contents in the candidate position group can form a complete semantics, and the above-mentioned second semantic analysis result is used to indicate that the sub-text contents in the candidate position group cannot form a complete semantics.

[0146] As an example, when the sub - text content corresponding to the target pixel position A in the candidate position group {A, B} is "scallions", and the sub - text content corresponding to the target pixel position B is "stir - fried tofu with scallions", when performing semantic analysis on the sub - text content corresponding to the candidate position group {A, B}, if the obtained semantic analysis result indicates that the sub - text contents in the candidate position group {A, B} can form a complete semantics, then the candidate position group {A, B} is determined as the target position group.

[0147] Continuing with the above example, when the sub - text content corresponding to the target pixel position A in the candidate position group {A, B} is "scallions", and the sub - text content corresponding to the target pixel position B is "stir - fried tofu with scallions", perform text combination on the sub - text contents corresponding to each target pixel position in the candidate position group to obtain the combined text {A, B} (scallions stir - fried tofu with scallions) and the combined text {B, A} (stir - fried tofu with scallions scallions) corresponding to the candidate position group. Perform semantic analysis on each combined text to obtain the combined semantic analysis result corresponding to each combined text. When there is a combined semantic analysis result indicating that the combined text {A, B} can form a complete semantics, determine the semantic analysis result of the candidate position group as the first semantic analysis result. The first semantic analysis result is used to indicate that the sub - text contents in the candidate position group can form a complete semantics.

[0148] In this way, perform position aggregation on each target pixel position to obtain at least one candidate position group, where the distance between each target pixel position in the candidate position group is less than a preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the target position group; when the number of candidate position groups is multiple, perform the following processing for each candidate position group respectively: perform semantic analysis on the sub - text contents corresponding to each target pixel position in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub - text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group, so as to perform position aggregation on the target pixel positions through whether they can form a complete semantics and obtain a target position group that can form a complete semantics.

[0149] In step 1062, combine the text content and the target position group to determine the text recognition result of the image to be recognized.

[0150] In some embodiments, the text content includes the sub - text contents corresponding to each target pixel position respectively, and the text recognition result includes the sub - text recognition results corresponding to each target position group respectively.

[0151] In some embodiments, step 1062 above may be implemented in the following manner: the following processing is performed for each target position group respectively: when the number of target pixel point positions in a target position group is one, the sub-text content corresponding to the target pixel point position is determined as the sub-text recognition result corresponding to the target position group; when the number of target pixel point positions in a target position group is multiple, the sub-text contents corresponding to the target pixel point positions in the target position group are textually fused to obtain the sub-text recognition result corresponding to the target position group.

[0152] In some embodiments, when the number of target pixel point positions in a target position group is multiple, the sub-text contents corresponding to the target pixel point positions in the target position group are textually fused in the sorting order of the target pixel point positions in the target position group to obtain the sub-text recognition result corresponding to the target position group.

[0153] Continuing with the example, when the sub-text content corresponding to the target pixel point position A in the target position group {A, B} is "scallions" and the sub-text content corresponding to the target pixel point position B is "stir-fried tofu", the sub-text contents corresponding to the target pixel point positions in the target position group are textually fused in the sorting order of the target pixel point positions in the target position group to obtain the sub-text recognition result corresponding to the target position group, "scallions stir-fried tofu".

[0154] In this way, by performing image encoding on the image to be recognized, the image features of the image to be recognized are obtained. Based on the image features, position decoding is performed on the image to be recognized to obtain the target pixel point positions in the image to be recognized, and position encoding is performed on the target pixel point positions to obtain the target position features. Then, combining the target position features and the image features, text encoding is performed on the head image to be recognized to obtain the text content in the image to be recognized. Combining the text content and the target pixel point positions, the text recognition result of the image to be recognized is determined. In this way, by determining the target position features that can indicate the positions of the target pixel points where text exists in the image to be recognized, and combining the target position features and the image features, the text content in the image to be recognized is decoded. Since the pixel point is the smallest recognition unit of the image to be recognized, the text content is determined by the target position features that can indicate the positions of the target pixel points where text exists in the image to be recognized, thereby achieving text recognition from the smallest recognition granularity, and thus effectively improving the accuracy of text recognition.

[0155] See Figure 9 , Figure 9 is the flow schematic of the training method of the text recognition model provided by the embodiments of the present application Figure 1 will be combined with Figure 9The steps 201 to 203 shown will be described. The training method of the text recognition model provided by the embodiments of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in cooperation. Hereinafter, the implementation by the server alone will be taken as an example for description.

[0156] In step 201, an image sample to be recognized and at least one sample label carried by the image sample to be recognized are obtained.

[0157] In some embodiments, the sample label is used to indicate the label text content corresponding to the position of the target pixel points in the image sample to be recognized.

[0158] In some embodiments, the position of the target pixel points is used to indicate the position of the target pixel points with text in the image sample to be recognized, and the number of sample labels is less than the total number of characters in the image sample to be recognized.

[0159] In some embodiments, the number of sample labels can be less than the total number of characters in the text of the image sample to be recognized, so as to effectively reduce the annotation cost of the sample labels and thus effectively improve the training efficiency of the text recognition model.

[0160] In some embodiments, the arrangement rule between different sample labels in the image sample to be recognized is the same as the arrangement rule of the text in the image sample to be recognized. For example, if the arrangement rule of the text in the image sample to be recognized is evenly arranged on the diagonal of the image to be recognized, then the arrangement rule between different sample labels in the image sample to be recognized can also be evenly arranged on the diagonal of the image to be recognized. By setting the arrangement rule between different sample labels in the image sample to be recognized to be the same as the arrangement rule of the text in the image sample to be recognized, the text recognition model obtained by training with the image sample to be recognized can accurately recognize the text direction in the image to be recognized, effectively improving the recognition performance of the text recognition model.

[0161] In step 202, an initial text recognition model is called to perform text recognition on the image sample to be recognized, and a text recognition result of the image sample to be recognized is obtained.

[0162] In some embodiments, the initial text recognition model includes an image encoding network, a position decoding network, and a text decoding network. Refer to Figure 10 , Figure 10 which is the flow schematic of the text recognition method provided by the embodiments of the present application Figure 5 , Figure 9 The step 202 shown can be implemented through Figure 10 the steps 2021 to 2024 shown.

[0163] In some embodiments, the above-mentioned initial text recognition model has the same network structure as the text recognition model, but different network parameters. That is, the text recognition model includes an image encoding network, a position decoding network, and a text decoding network.

[0164] In step 2021, the image encoding network is called to perform image encoding on the image sample to be recognized, and the sample image features of the image sample to be recognized are obtained.

[0165] In some embodiments, the above-mentioned image encoding network is used to perform image encoding on the image sample to be recognized and convert the image sample to be recognized into sample image features in vector form.

[0166] In step 2022, the position decoding network is called to perform position decoding on the image sample to be recognized based on the sample image features, and at least one predicted pixel point position in the image sample to be recognized is obtained.

[0167] In some embodiments, the above-mentioned position decoding network includes a position recognition layer and a position correction layer, and the sample image features include the pixel point features respectively corresponding to each pixel point in the image sample to be recognized.

[0168] In some embodiments, step 2022 can be implemented in the following manner: the position recognition layer is called to perform position recognition on the image sample to be recognized based on the sample image features, and at least one candidate pixel point position in the image sample to be recognized is obtained; the position correction layer is called to perform position correction on each candidate pixel point position based on the pixel point features respectively corresponding to each candidate pixel point position, and at least one predicted pixel point position in the image sample to be recognized is obtained.

[0169] In some embodiments, the above-mentioned calling the position correction layer to perform position correction on each candidate pixel point position based on the pixel point features respectively corresponding to each candidate pixel point position to obtain at least one predicted pixel point position in the image sample to be recognized can be implemented in the following manner: when the number of candidate pixel point positions is one, the candidate pixel point position is determined as the predicted pixel point position; when the number of candidate pixel point positions is multiple, the position correction layer is called to perform position correction on each candidate pixel point position based on the pixel point features respectively corresponding to each candidate pixel point position, and the correction information corresponding to each candidate pixel point position is obtained; for each candidate pixel point position, when the correction information corresponding to the candidate pixel point position indicates that the candidate pixel point position is the position where there is text in the image to be recognized, the candidate pixel point position is determined as the predicted pixel point position.

[0170] In step 2023, position encoding is performed on each predicted pixel point position to obtain predicted position features, and the text decoding network is called to perform text decoding on the image sample to be recognized, and the predicted text content in the image sample to be recognized is obtained.

[0171] In some embodiments, the text decoding network includes a fusion layer and a text decoding layer. To call the text decoding network to perform text decoding on the image sample to be recognized and obtain the predicted text content in the image sample to be recognized, the following method can be used: call the fusion layer to fuse the predicted position features and the sample image features to obtain fused features; call the text decoding layer to perform text decoding on the image to be recognized based on the fused features to obtain the text content in the image to be recognized.

[0172] In step 2024, aggregate the positions of each predicted pixel point to obtain at least one predicted position group, and combine the text content and the predicted position group to determine the text recognition result of the image sample to be recognized.

[0173] In some embodiments, the text content includes sub-text contents corresponding to the positions of each predicted pixel point respectively. To aggregate the positions of each predicted pixel point to obtain at least one predicted position group, the following method can be used: perform position aggregation on the positions of each predicted pixel point to obtain at least one candidate position group, where the distance between the positions of each predicted pixel point in the candidate position group is less than a preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the predicted position group; when the number of candidate position groups is multiple, select at least one predicted position group from the multiple candidate position groups based on the sub-text contents.

[0174] In some embodiments, to select at least one predicted position group from the multiple candidate position groups based on the sub-text contents, the following method can be used: perform the following processing for each candidate position group respectively: perform text content semantic analysis on the sub-text contents corresponding to the positions of each predicted pixel point in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group.

[0175] In some embodiments, the text content includes sub-text contents corresponding to the positions of each predicted pixel point respectively, and the text recognition result includes sub-text recognition results corresponding to each predicted position group respectively.

[0176] In some embodiments, to combine the text content and the predicted position group to determine the text recognition result of the image sample to be recognized, the following method can be used: perform the following processing for each predicted position group respectively: when the number of target pixel point positions in the predicted position group is one, determine the sub-text content corresponding to the target pixel point position as the sub-text recognition result corresponding to the predicted position group; when the number of target pixel point positions in the predicted position group is multiple, perform text fusion on the sub-text contents corresponding to the target pixel point positions in the predicted position group to obtain the sub-text recognition result corresponding to the predicted position group.

[0177] In step 203, the initial text recognition model is trained by combining the sample labels and the text recognition results to obtain a text recognition model.

[0178] In some embodiments, the text recognition model is used to recognize the text content in the image to be recognized. The model structure of the text recognition model is the same as that of the initial text recognition model, and the model parameters of the text recognition model are different from those of the initial text recognition model.

[0179] In some embodiments, the text recognition results include sub-text recognition results corresponding to each predicted position group. Refer to Figure 11 , Figure 11 which is a schematic flowchart of the text recognition method provided by the embodiments of the present application. Figure 6 , Figure 9 As shown in Figure 11 step 203 shown in

[0180] In step 2031, for each sub-text recognition result, the similarity between the text content indicated by the sub-text recognition result and each label text content is obtained, and based on the maximum similarity, the loss value corresponding to the sub-text recognition result is determined.

[0181] In some embodiments, the similarity between the text content indicated by the sub-text recognition result and each label text content is used to indicate the degree of content similarity between the text content indicated by the sub-text recognition result and the label text content.

[0182] As an example, the expression of the loss value corresponding to the sub-text recognition result can be:

[0183] L1 = logP(1)

[0184] where P is used to indicate the maximum similarity, and L1 is used to indicate the loss value corresponding to the sub-text recognition result.

[0185] In step 2032, the initial text recognition model is trained based on the loss value to obtain a text recognition model.

[0186] In some embodiments, step 1032 can be implemented in the following manner: when the number of sub-text recognition results is one, the loss value corresponding to the sub-text recognition result is determined as the target loss value; when the number of sub-text recognition results is multiple, the loss values corresponding to each sub-text recognition result are summed to obtain the target loss value; based on the target loss value, the initial text recognition model is trained to obtain a text recognition model.

[0187] As an example, when the number of sub - text recognition results is multiple, the expression of the corresponding target loss value can be:

[0188] L = ∑L1 (2)

[0189] Where L is used to indicate the target loss value, and L1 is used to indicate the loss value corresponding to the sub - text recognition result.

[0190] In this way, by setting the number of sample labels to be less than the total number of words in the text of the image to be recognized, the annotation cost of the sample labels can be effectively reduced, thereby effectively improving the training efficiency of the text recognition model. The arrangement rule between different sample labels in the image to be recognized is the same as the arrangement rule of the text in the image to be recognized. For example, if the arrangement rule of the text in the image to be recognized is evenly arranged on the diagonal of the image to be recognized, then the arrangement rule between different sample labels in the image to be recognized can also be evenly arranged on the diagonal of the image to be recognized. By setting the arrangement rule between different sample labels in the image to be recognized to be the same as the arrangement rule of the text in the image to be recognized, the text recognition model obtained by training with the image to be recognized can accurately recognize the text direction in the image to be recognized, effectively improving the recognition performance of the text recognition model.

[0191] In this way, by performing image encoding on the image to be recognized, the image features of the image to be recognized are obtained. Based on the image features, position decoding is performed on the image to be recognized to obtain the position of the target pixel points in the image to be recognized, and position encoding is performed on the position of the target pixel points to obtain the target position features. Combining the target position features and the image features, text encoding is performed on the image to be recognized to obtain the text content in the image to be recognized. Combining the text content and the position of the target pixel points, the text recognition result of the image to be recognized is determined. In this way, by determining the target position features that can indicate the position of the target pixel points where text exists in the image to be recognized, and combining the target position features and the image features to decode the text content in the image to be recognized. Since the pixel point is the smallest recognition unit of the image to be recognized, the text content is determined by the target position features that can indicate the position of the target pixel points where text exists in the image to be recognized, thereby realizing text recognition from the smallest recognition granularity, and effectively improving the accuracy of text recognition.

[0192] Next, an exemplary application of the embodiments of the present application in an actual text recognition application scenario will be described.

[0193] The text recognition method and the training method of the text recognition model provided by the embodiments of the present application can be applied to solve the scene text localization task. By discretizing each text instance and image features into a token sequence, and using the proposed explicit point compound query modeling mechanism, a learnable coordinate bias is constructed into a compound query to optimize the localization point coordinates and reduce the impact brought by point annotation noise. In addition, a spatial compatibility attention mechanism that enhances the layout awareness ability is used for content decoding to better handle the relationship between detection and recognition and solve the problem of poor recognition performance on strongly rotated and inverted texts. Quantitative experiments on public benchmark tests show that the embodiments of the present application can achieve a fully supervised effect with a small number of text localization points, the annotation cost is much lower than that of polygons, and the state-of-the-art results are obtained on widely used benchmark tests, and better training efficiency is achieved.

[0194] The embodiments of the present application propose a weakly supervised text recognition method based on point annotation for OCR text recognition in general scenarios. The embodiments of the present application use a single point with low labor cost to annotate text instances, instead of requiring complex and fine polygon box annotations (text line annotations, word or character-level annotations) like other current mainstream methods. The embodiments of the present application treat text recognition as a sequence prediction task. Given a general scene image as input, the embodiments of the present application encode the image into a discrete feature symbol and use an autoregressive decoder to predict the text sequence and text coordinates. After passing through the decoder, the sequence becomes a feature sequence with text semantics and positions. Due to the parallel decoding architecture of Transformer, this method can decode the feature sequence into the center point of the text, the text content, and the confidence through a parallel prediction head and a feature mapping function. The proposed method is simple and efficient and can obtain advanced results on benchmarks widely used in the academic community. In addition, the embodiments of the present application also introduce a text matching criterion to provide a more accurate supervision signal, enabling the model to converge quickly.

[0195] The embodiments of the present application propose a recognition model structure with low annotation cost, simplicity and efficiency for application scenarios such as large-scale and high-concurrency recognition. Using the Transformer encoder-decoder structure, the modality difference between image pixels and text sequences is alleviated in an end-to-end manner. Without sacrificing the recognition accuracy, the commonly used serial RNN module and the architecture of multi-layer fully connected layers with high time consumption are discarded, and text recognition is realized simply and efficiently.

[0196] The embodiments of the present application aim to improve the efficiency and user experience of OCR projects and products. By optimizing the existing high-quality recognition models, faster models can be created quickly, improving the model training and inference speed while maintaining the accuracy unchanged. This method is applicable to multiple OCR application scenarios such as advertisement review, video subtitle recognition, image / video feature extraction, photo translation, intelligent marking, and similarity systems. In a typical application scenario such as photo text recognition, the embodiments of the present application can quickly train a recognition model using a large amount of multi-scenario data of any length, improving the iteration speed of the effect; during the process of re-learning the model when changing the application scenario, the embodiments of the present application can update and iterate the model at a low cost (by changing the dataset and annotation).

[0197] In some embodiments, referring to Figure 5 , the text recognition model provided by the embodiments of the present application adopts an encoder-decoder architecture design, which includes a feature encoding module that generates multi-scale feature maps and two cascaded decoders to achieve end-to-end text localization. The core idea of the text recognition model is mainly in the two decoder parts. First, in the coordinate decoder with the ability to generate composite queries, a composite query is constructed by introducing learnable coordinate bias parameters, which allows for inaccurate point annotations. In addition, this module updates the bias parameters between the text decoder layers, improving the localization performance of the detection task. Second, to solve the problem of the inconsistency between the text spatial information and the implicit text content information modalities, spatial compatibility attention with enhanced layout awareness is proposed to perform feature interaction between the two and guide the decoding of the text sequence. In this way, the OCR task is solved with high precision. Finally, the captured features are first decoded into coordinates by the coordinate decoder, and the proposed composite query is used to generate deep features with spatial information. Next, in the text decoding part with spatial awareness, an attention module with spatial compatibility is introduced for recognition. This means that most of the calculations are completed by the Transformer, and the Transformer is jointly optimized for these two tasks.

[0198] In some embodiments, referring to Figure 12 , Figure 12 is a schematic diagram of the annotation principle of the sample labels provided by the embodiments of the present application. Traditional text localization methods usually use rectangular boxes or polygon annotations. However, rectangular boxes will introduce background noise in the non-text areas within the box when representing arbitrarily shaped text, while polygon annotations have a high cost. To solve the problems of high annotation cost and poor representation, the embodiments of the present application are inspired by the reading mechanism of the human eye moving from one character to the next along the center line of the text area during reading. According to the text reading order, a certain number of fuzzy positioning points are used (for example, as Figure 12The fuzzy positioning points 61 and 62 shown in [figure] are used to flexibly fit the shape of the scene text. It can better represent the problems of arbitrary shape text and reverse reading, which is beneficial to the detection and recognition tasks. Especially for the scene text with non-traditional reading directions, this annotation method is extremely simple, greatly improving the representation form of the text and spatial information. At the same time, due to the relatively sparse number of positioning points, it also helps to eliminate character-level annotations. To enhance the effectiveness of the model for fuzzy point annotation, random offset noise is added to the coordinate points sampled along the text center curve in the reading order, hoping that the model can learn the relevant noise interference.

[0199] In some embodiments, referring to Figure 13 , Figure 13 is a schematic diagram of the principle of the training method of the text recognition model provided by the embodiments of the present application, initial training data preparation. First, in the embodiments of the present application, a large number of pictures are extracted from the original massive text picture data for annotation to obtain training data. The annotation data for each picture only requires point annotation (the annotation position can be at the center of the text and any character). The image features of the sample image 30 are extracted through the CNN network to obtain image features. Based on the image features, coordinate encoding and feature dimensionality reduction are performed to obtain reference features. The feature encoding is performed through the Transformer encoder to obtain a coordinate feature sequence. Through character coordinate mapping of the feature sequence output by the encoder, the final prediction sequence is obtained, and the final prediction sequence is decoded to obtain the prediction result 31.

[0200] In some embodiments, referring to Figure 14 , Figure 14 is a schematic diagram of the effect of point annotation provided by the embodiments of the present application. Point annotation is realized by annotating multiple annotation points 41 at the center of the text or any character in the picture.

[0201] Preprocessing of arbitrarily long text lines. During the training process, an image with text is fed into a feature extraction network (CNN structure) to obtain image features, and the center point of the text coordinates is encoded to be embedded with the image features of the same dimension. In the embodiments of the present application, the maximum length of the decoded text is set to 30. If the variable-length text in the image is less than 30, it is padded with blanks at the end; otherwise, it is truncated to 30 characters. Specifically, a text sequence is represented by three discrete characters, namely the abscissa of the point, the ordinate of the point, and the padded text, where each coordinate is uniformly discretized into an integer between [1, 1000]. In the embodiments of the present application, a vocabulary is used to map all discrete characters, and the size of the vocabulary is equal to the number of coordinates of the discretized points plus the number of characters. Through this quantization scheme, the embodiments of the present application can achieve high-precision transcription using a small vocabulary. For example, a 600x600 image only requires 600 to be uniformly discretized into [1, 600] to achieve zero quantization error. This is much smaller than modern language models with a vocabulary of 32K or higher.

[0202] Data augmentation is adopted during the training / inference phase, including but not limited to data augmentation methods such as Gaussian blur, Gaussian noise, random slicing, exposure, brightness adjustment, flipping, mirroring, etc. Normalization, such as subtracting the mean and dividing by the standard deviation for each pixel point's RGB or grayscale value, or subtracting 127 and dividing by 128 to normalize to the (-1, 1) interval, etc.

[0203] In some embodiments, the coordinate and the text sequence decoder, and the text sequence are divided into two parts of coordinate regression and sequence by a Transformer decoder that shares the same feature weights. In the embodiments of the present application, <start>and <eos>Markers are inserted at the head and tail of the sequence to indicate the start and end of sequence decoding. Among them, the text feature sequence will be randomly sorted. In this way, randomly sorted text instances can be effectively learned, so as to implicitly implement label assignment for different hidden features, which cleverly avoids explicit label assignment like using bipartite matching. In fact, compared with other label assignment methods, NMS and bipartite matching use a large amount of space to detect and identify text, which will consume additional computational resources for non-text area decoding, and the coordinate decoder is more efficient intuitively. Regarding the recognition of text content as a sequence classification problem of indefinite length. Due to the possible misalignment problem of the sequence, more computational resources are consumed during the training process. To eliminate such problems, the embodiments of the present application first pad or truncate the text to a fixed length of 30, where <pad>Tokens are used to fill in the gaps of shorter text instances. Additionally, assuming there are 2786 categories of characters (e.g., Chinese, English characters, symbols), three additional entries are added to the vocabulary of the token sequence dictionary, where the three additional categories are used for <pad>, SOS> and <eos>Mark.

[0204] In some embodiments, during the inference phase, the coordinate sequence is predicted sequentially first until the sequence is marked as ended <eos>It occurs. Then, the sequence integrating the coordinate center feature is used to autoregressively predict the text content in parallel. Thus, the tokens can be easily converted into point coordinates and transcriptions, generating the text sequence result. In addition, the average confidence score of the probabilities of all tokens in the corresponding segment is used to filter the original output, effectively removing redundant and false positive prediction features.

[0205] In some embodiments, training and application. Training is performed using the Adam optimizer, and the loss function is the cross-entropy loss. The recognition confidence is calculated directly based on the softmax layer probabilities during text decoding. The embodiments of the present application train the model of the present application on a mixture of the SynthText and real datasets. In the pre-training stage, the model is first trained on the synthetic dataset SynthText 150k, real datasets MLT-2017, ICDAR 2013, ICDAR 2015, as well as Total-Text and TextOCR. All experiments are distributed training on 8 NVIDIA A100 GPUs, with a batch size of 2 per GPU. Additionally, the embodiments of the present application use ResNet-50 as the backbone network.

[0206] In some embodiments, during the process of decoding coordinates and text, SPTS decodes fixed coordinate points and text by using a decoder with shared weights, but the performance of the decoder is significantly affected by the positions of the annotated points. Specifically, if the positions of the annotated points are unreasonable, it may cause the decoder to fail to correctly recognize the text, thereby affecting the accuracy of the entire text recognition task. Since there may be a difference between the optimal coordinates marked by humans and the understanding ability of the text decoder, in order to obtain the optimal coordinate information considered by the decoder, this coordinate needs to be optimized by the recognizer loss and dynamically corrected to ensure its match with the understanding ability of the decoder. In order to learn the positioning space information hidden in the fuzzy point annotations and reduce the position interference caused by point annotations, the embodiments of the present application further propose a composite query modeling mechanism with point calibration. By introducing a learnable coordinate bias to form a composite query with text positioning result P( i ) to guide text recognition. By feeding the query into the text recognizer, the recognizer predicts this bias to iteratively optimize the calibration points. The embodiments of the present application can use the following complete formula to generate the calibrated composite query:

[0207] Q (i) =P (i) +(ΔX1,ΔY1…,ΔX4,ΔY4) (3)

[0208] The sequence P (i) represents the 4 fuzzy positioning point coordinates of a word (ΔX1, ΔY1..., ΔX4, ΔY4) represents the learnable bias parameters of 4 positioning points in the x and y directions, and are used to generate the composite query Q (i) After that, it is sent to the TextTransformer Decoder Layer of layer L, and the coordinate biases are updated layer by layer. The new composite query points are formed and sent to the next layer. The reference points are updated by using the loss optimization of the recognizer to replace the reference points of the previous level as query information, so that the recognizer can guide the detection in the reverse direction and output the optimal coordinate reference points considered by the recognizer. Different levels of positioning prior features can be sensed between the decoding layers of the recognizer, and these positioning prior features are used to indicate the decoder, so as to allow inaccurate annotations to still achieve the best text decoding results and can resist random offset noise. In the specific implementation process, it is found that only using the composite point query is not enough, because the recognizer does not perceive the global feature sequence of the image. In order to allow the recognizer to sense the previous spatial information and the global feature sequence of the image, another branch is introduced in this embodiment of the application, and the coordinate decoder obtains N coordinate positioning points are fused with the position features Feat of the hidden text instance to obtain the fusion information R with the text spatial form (i) , which is used as the input of the text decoder spatial compatibility module to enable the recognizer to perceive the previous positioning information. This process can be expressed as follows:

[0209] R (i) =(MLP(Emb p (Q (i) ))+Feat where i∈{1,2,...,N} (4)

[0210] Among them, Embp represents the sine position encoding function, and this embodiment of the application also uses a 2-layer MLP head with RELU activation for further projection.

[0211] Since there may be huge differences in the position, size, rotation angle, etc. of the text in the image, it will be affected during the process of processing texts with different forms. Since the fusion information R( i ) with the text spatial form is not introduced, when the form and rotation angle differences are large, the method of sharing the decoder by SPT S may cause deviations in the results of text detection and text recognition, thus affecting the accuracy of the entire text recognition task. This is the main reason for the poor performance of SPTS in recognizing strongly rotated and inverted texts.

[0212] To enable the text recognizer to sense information with text spatial layout features and allow cross-modal interaction for detecting and recognizing information, an embodiment of this application introduces a new spatial compatibility self-attention mechanism in the text decoder to accommodate the interaction between the fused information with text spatial morphology and the recognition query information, enabling the recognizer to make good use of this information and correctly decode the text. This mechanism improves the model's processing ability for spatial data by introducing explicit spatial information into the input data and using the attention mechanism to learn the correlations in the space.

[0213] Specifically, in the text decoder, the SCA module is first adopted to capture the relationship between the composite query vector Q (i) and the text query vector. The Text Content Queries Tq input to each decoding layer are concatenated to their corresponding composite query vectors Q (i) and self-attention is calculated, where the key is the same as q, and the value does not include partial information of the embedding position. Meanwhile, an embodiment of this application introduces an autonomous attention mechanism with spatial compatibility attention bias to simulate the spatial compatibility relationship between the fused information R (i) of the text spatial features and the text query entity K concatenated with the composite query. The attention score between the query qi and the key ki is expressed as:

[0214] α qk =(CONCAT(c q ,p q )) T (CONCAT(c k ,p k ))+FFN(R (i) ,k) (5)

[0215] where c q ,c k are the q and k corresponding to the Text Content Queries, p q ,p k are the q and k of the Complex query, and concat represents vector concatenation. FFN(r qi ,k i ) is the spatial compatibility bias, where FFN represents a two-layer feed-forward network that matches the dimension of R (i) with the dimension of k. The feature vector output by each self-attention layer is the weighted sum of the embedding values in all contents according to the normalized attention scores. After SCA, the query zi is obtained and then fed into the deformable cross-attention module. As the explicit point information flows in the decoder, an embodiment of this application adopts a three-layer MLP head to predict the offset and updates the query vector Q (i) Finally, output the text class.

[0216] If the neural network knows the location and content of the object, the embodiments of the present application only need to teach it to read them out. By learning to describe the object, the model can learn to build language on the basis of pixel observations, thereby generating an effective object representation. The embodiments of the present application regard text localization as a language sequence generation task conditional on pixel inputs. For this task, the model architecture and loss function are general and relatively simple, and people can easily extend this framework to different fields or applications, or incorporate it into a perception system that supports general intelligence. For this purpose, it provides a language interface for a wide range of visual tasks.

[0217] Since the sequence tokens are predicted without a specific task head, the embodiments of the present application use cross-entropy loss to train the model. The purpose is similar to language modeling, predicting tokens given an image and the previous tokens, with a maximum likelihood loss, that is:

[0218]

[0219] where x is the given image, y and are the input and target sequences associated with x, and L is the target sequence length. y and are the same in the language modeling setting, but they can also be different (just like the augmented sequence construction in the embodiments of the present application later). wj is the pre-assigned weight of the j-th token in the sequence. To better identify the text sequence and reduce the weight of the imprecise weakly supervised coordinate points, the embodiments of the present application set wj to 1 and 2 respectively when decoding the coordinate points and the text sequence.

[0220] During inference, the embodiments of the present application draw tokens from the model likelihood, that is This can be achieved by adopting argmax sampling or using other random sampling techniques. When the EOS token is generated, the sequence decoding ends, and the predicted point coordinates and character class labels can be directly obtained through a simple classification mapping.

[0221] In some embodiments, during the training process, the embodiments of the present application trained the model of the embodiments of the present application on a mixed dataset of SynthText and real data. In the pre-training stage, the model was first trained on the synthetic dataset SynthText containing 150,000 samples for 150 epochs, and was also trained on real datasets such as MLT-2017, ICDAR 2013, ICDAR 2015, Total-Text, and TextOCR. The optimizer used was Adamw, with an initial learning rate of 5x10^(-4), and the learning rate was decreased by 1x10^(-5) after each epoch. After the training was completed, the embodiments of the present application fine-tuned the model on each dataset respectively, with a fixed learning rate of 1x10^(-5) for each optimization task, and performed 120 epochs of iteration. All experiments were conducted on 8 GPUs using distributed training, with a batch size of 2 for each GPU. In addition, the embodiments of the present application used ResNet-50 as the backbone network, and both the Transformer encoder and decoder had 6 layers, with each layer containing 8 heads. The embodiments of the present application adopted Pre-LN Transformer as the Transformer architecture of the embodiments of the present application.

[0222] The embodiments of the present application utilize the Transformer encoding and decoding structure to alleviate the modality difference between image pixels and text sequences in an end-to-end manner. Without significantly changing the recognition accuracy, the commonly used serial RNN module and the architecture of multi-layer fully connected layers, which are time-consuming, are discarded, and text recognition is achieved simply and efficiently.

[0223] In some embodiments, refer to Figure 15 , Figure 15 is a schematic diagram of the text recognition effect of the text recognition method provided by the embodiments of the present application. Through the text recognition method provided by the embodiments of the present application, the text content 51 in the image to be recognized can be accurately recognized.

[0224] Thus, a compatible mechanism of spatial attention is added to the text recognition model provided by the embodiments of the present application, which can capture the relationship between characters of words, and this helps with the recognition task. While SPTS only uses a simple shared decoder to decode coordinates and content, ignoring the internal connection therein. Compared with other methods that need to correct arbitrarily shaped text into a regular shape, the need for sampling and independently designed modules is reduced, thereby reducing the complexity of the model. Since the recognition result is better, the detection result can be guided and corrected by the recognition decoder. Secondly, in the aspect of general scene text recognition technology, the speed and model complexity problems of RNN have always been important bottlenecks of recognition algorithms. Solving this problem helps to improve the speed and performance of the recognition model, reduce the occupation of machine computing power, and broaden the deployment scenarios. In the embodiments of the present application, for most general text line recognition scenarios, the Transformer encoding and decoding structure is used to alleviate the modality difference between image pixels and text sequences in an end-to-end manner, and the commonly used serial RNN module and the architecture of multi-layer fully connected layers with high time consumption are discarded while the recognition accuracy remains basically unchanged, realizing text recognition simply and efficiently.

[0225] It can be understood that in the embodiments of the present application, data related to images to be recognized and the like are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0226] Next, the implementation of the text recognition device 455 provided by the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, as Figure 1 shown, the software modules in the text recognition device 455 stored in the memory 450 may include: an image encoding module 4551, configured to obtain an image to be recognized including text content, and perform image encoding on the image to be recognized to obtain image features of the image to be recognized; a position decoding module 4552, configured to perform position decoding on the image to be recognized based on the image features to obtain at least one target pixel point position in the image to be recognized, where the target pixel point position is used to indicate the position of the target pixel point with text in the image to be recognized; a position encoding module 4553, configured to perform position encoding on each of the target pixel point positions to obtain target position features, and combine the target position features and the image features to perform text decoding on the image to be recognized to obtain the text content in the image to be recognized; a determination module 4554, configured to determine the text recognition result of the image to be recognized by combining the text content and the target pixel point position.

[0227] In some embodiments, the position decoding is implemented by a position decoding network, which includes a position recognition layer and a position correction layer. The image features include the pixel features corresponding to each pixel point in the image to be recognized. The above-mentioned position decoding module is further configured to call the position recognition layer, based on the image features, perform position recognition on the image to be recognized, and obtain at least one candidate pixel point position in the image to be recognized; call the position correction layer, and based on the pixel features corresponding to each candidate pixel point position, perform position correction on each candidate pixel point position to obtain at least one target pixel point position in the image to be recognized.

[0228] In some embodiments, the above-mentioned position decoding module is further configured to, when the number of candidate pixel point positions is one, determine the candidate pixel point position as the target pixel point position; when the number of candidate pixel point positions is multiple, call the position correction layer, and based on the pixel features corresponding to each candidate pixel point position, perform position correction on each candidate pixel point position to obtain the correction information corresponding to each candidate pixel point position; for each candidate pixel point position, when the correction information corresponding to the candidate pixel point position indicates that the candidate pixel point position is the position where text exists in the image to be recognized, determine the candidate pixel point position as the target pixel point position.

[0229] In some embodiments, the above-mentioned text decoding is implemented by a text decoding network, which includes a fusion layer and a text decoding layer. The above-mentioned position encoding module is further configured to call the fusion layer to fuse the target position features and the image features to obtain fused features; call the text decoding layer, and based on the fused features, perform text decoding on the image to be recognized to obtain the text content in the image to be recognized.

[0230] In some embodiments, the above-mentioned fusion layer includes a position fusion layer and a feature fusion layer. The above-mentioned position encoding module is further configured to, for each target pixel point position, determine the pixel point positions in the image to be recognized whose distance from the target pixel point position is less than a distance threshold as the reference pixel point positions corresponding to the target pixel point position; perform position encoding on the reference pixel point positions to obtain reference position features; and call the position fusion layer to perform position fusion on the target position features and the reference position features to obtain position fusion features; call the feature fusion layer to fuse the position fusion features and the image features to obtain the fused features.

[0231] In some embodiments, the above-mentioned determination module is further configured to aggregate the positions of the target pixel points to obtain at least one target position group; and determine the text recognition result of the image to be recognized in combination with the text content and the target position group.

[0232] In some embodiments, the text content includes sub-text contents respectively corresponding to the positions of the target pixel points. The above-mentioned determination module is further configured to perform position aggregation on the positions of the target pixel points to obtain at least one candidate position group, where the distance between the positions of the target pixel points in the candidate position group is less than a preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the target position group; when the number of candidate position groups is multiple, select at least one target position group from the multiple candidate position groups based on the sub-text contents.

[0233] In some embodiments, the above-mentioned determination module is further configured to perform the following processing on each of the candidate position groups: perform text content semantic analysis on the sub-text contents corresponding to the positions of the target pixel points in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group.

[0234] In some embodiments, the text content includes sub-text contents respectively corresponding to the positions of the target pixel points, and the text recognition result includes sub-text recognition results respectively corresponding to the target position groups; the above-mentioned determination module is further configured to perform the following processing on each of the target position groups: when the number of the target pixel points in the target position group is one, determine the sub-text content corresponding to the target pixel point position as the sub-text recognition result corresponding to the target position group; when the number of the target pixel points in the target position group is multiple, perform text fusion on the sub-text contents corresponding to the target pixel point positions in the target position group to obtain the sub-text recognition result corresponding to the target position group.

[0235] Next, the implementation of the training device 555 of the text recognition model provided by the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, as Figure 1 As shown, the software modules in the training device 555 of the text recognition model stored in the memory 550 may include: an acquisition module 5551, configured to acquire an image sample to be recognized and at least one sample label carried by the image sample to be recognized, where the sample label is used to indicate the label text content corresponding to the position of the target pixel points in the image sample to be recognized; wherein, the position of the target pixel points is used to indicate the position where the target pixel points with text exist in the image sample to be recognized, and the number of the sample labels is less than the total number of characters in the text of the image sample to be recognized; a recognition module 5552, configured to call an initial text recognition model to perform text recognition on the image sample to be recognized to obtain a text recognition result of the image sample to be recognized; a training module 5553, configured to train the initial text recognition model in combination with the sample label and the text recognition result to obtain the text recognition model, and the text recognition model is used to recognize the text content in the image to be recognized.

[0236] In some embodiments, the initial text recognition model includes an image encoding network, a position decoding network, and a text decoding network. The above recognition module is further configured to call the image encoding network to perform image encoding on the image sample to be recognized to obtain sample image features of the image sample to be recognized; call the position decoding network to perform position decoding on the image sample to be recognized based on the sample image features to obtain at least one predicted pixel point position in the image to be recognized; perform position encoding on each of the predicted pixel point positions to obtain predicted position features, and call the text decoding network to perform text decoding on the image sample to be recognized to obtain predicted text content in the image sample to be recognized; aggregate each of the predicted pixel point positions to obtain at least one predicted position group, and combine the text content and the predicted position group to determine the text recognition result of the image sample to be recognized.

[0237] In some embodiments, the text recognition result includes sub-text recognition results respectively corresponding to each predicted position group. The above training module is further configured to, for each of the sub-text recognition results, obtain the similarity between the text content indicated by the sub-text recognition result and each of the label text contents, and determine the loss value corresponding to the sub-text recognition result based on the maximum similarity; train the initial text recognition model based on the loss value to obtain the text recognition model.

[0238] In some embodiments, the above training module is further configured to, when the number of sub - text recognition results is one, determine the loss value corresponding to the sub - text recognition result as the target loss value; when the number of sub - text recognition results is multiple, sum up the loss values corresponding to each sub - text recognition result to obtain the target loss value; and train the initial text recognition model based on the target loss value to obtain the text recognition model.

[0239] An embodiment of the present application provides a computer program product, which includes a computer program or computer - executable instructions. The computer program or computer - executable instructions are stored in a computer - readable storage medium. The processor of the electronic device reads the computer - executable instructions from the computer - readable storage medium, and the processor executes the computer - executable instructions, so that the electronic device executes the text recognition method and the training method of the text recognition model in the above embodiments of the present application.

[0240] An embodiment of the present application provides a computer - readable storage medium storing computer - executable instructions, where the computer - executable instructions are stored. When the computer - executable instructions are executed by a processor, the processor will be caused to execute the text recognition method and the training method of the text recognition model provided in the embodiments of the present application. For example, Figure 4 the text recognition method shown.

[0241] In some embodiments, the computer - readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD - ROM; or may be various electronic devices including one or any combination of the above memories.

[0242] In some embodiments, the computer - executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, sub - routine, or other unit suitable for use in a computing environment.

[0243] As an example, the computer - executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, sub - routines, or code portions).

[0244] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.

[0245] As an example, computer-executable instructions can be deployed to be executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected through a communication network.

[0246] In summary, the embodiments of the present application have the following beneficial effects:

[0247] (1) By performing image encoding on the image to be recognized, the image features of the image to be recognized are obtained. Based on the image features, position decoding is performed on the image to be recognized to obtain the positions of the target pixel points in the image to be recognized, and position encoding is performed on the positions of the target pixel points to obtain the target position features. By combining the target position features and the image features, text encoding is performed on the head image to be recognized to obtain the text content in the image to be recognized. By combining the text content and the positions of the target pixel points, the text recognition result of the image to be recognized is determined. In this way, by determining the target position features that can indicate the positions of the target pixel points where text exists in the image to be recognized, and combining the target position features and the image features, the text content in the image to be recognized is decoded. Since the pixel point is the smallest recognition unit of the image to be recognized, the text content is determined by the target position features that can indicate the positions of the target pixel points where text exists in the image to be recognized, thereby achieving text recognition from the smallest recognition granularity, and thus effectively improving the accuracy of text recognition.

[0248] (2) By calling the position correction layer, based on the pixel point features respectively corresponding to the positions of each candidate pixel point, position correction is performed on the positions of each candidate pixel point to obtain at least one position of the target pixel point in the image to be recognized, thereby achieving position correction of the positions of the candidate pixel points and making the obtained positions of the target pixel points more accurate.

[0249] (3)For each target pixel position, since there is a high probability that there is text in the image area around the target pixel in the image to be recognized, the pixel positions in the image to be recognized whose distance from the target pixel position is less than the distance threshold can be determined as the reference pixel positions corresponding to the target pixel position, and the target position feature and the reference position feature are fused in position. The obtained position fusion feature can more comprehensively cover all the text-containing areas in the image to be recognized, so that the representation accuracy of the obtained fusion feature is higher. By using the fusion feature with higher representation accuracy to perform text decoding on the image to be recognized, the obtained text content is more accurate.

[0250] (4)Perform position aggregation on each target pixel position to obtain at least one candidate position group, where the distance between the target pixel positions in the candidate position group is less than the preset distance threshold; when the number of candidate position groups is one, determine the candidate position group as the target position group; when the number of candidate position groups is multiple, perform the following processing on each candidate position group respectively: perform text content semantic analysis on the sub-text contents corresponding to the target pixel positions in the candidate position group to obtain a semantic analysis result; when the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determine the candidate position group as the target position group, so as to perform position aggregation on the target pixel positions through whether they can form a complete semantics and obtain a target position group that can form a complete semantics.

[0251] (5)The arrangement rules between different sample labels in the image sample to be recognized are the same as the arrangement rules of the text in the image sample to be recognized. For example, if the arrangement rule of the text in the image sample to be recognized is uniform arrangement on the diagonal of the image to be recognized, then the arrangement rule between different sample labels in the image sample to be recognized can also be uniform arrangement on the diagonal of the image to be recognized. By setting the arrangement rules between different sample labels in the image sample to be recognized to be the same as the arrangement rules of the text in the image sample to be recognized, the text recognition model obtained by training with the image sample to be recognized can accurately recognize the text direction in the image to be recognized, effectively improving the recognition performance of the text recognition model.

[0252] (6) By setting the number of sample labels to be less than the total number of characters in the text of the image sample to be recognized, the annotation cost of the sample labels can be effectively reduced, thereby effectively improving the training efficiency of the text recognition model. The arrangement rule between different sample labels in the image sample to be recognized is the same as the arrangement rule of the text in the image sample to be recognized. For example, if the arrangement rule of the text in the image sample to be recognized is uniform arrangement on the diagonal of the image to be recognized, then the arrangement rule between different sample labels in the image sample to be recognized can also be uniform arrangement on the diagonal of the image to be recognized. By setting the arrangement rule between different sample labels in the image sample to be recognized to be the same as the arrangement rule of the text in the image sample to be recognized, the text recognition model obtained by training with the image sample to be recognized can accurately recognize the text direction in the image to be recognized, effectively improving the recognition performance of the text recognition model.

[0253] (7) A compatible mechanism of spatial attention is added to the text recognition model provided by the embodiments of the present application, which can capture the relationship between characters of words and is helpful for the recognition task. While SPTS only uses a simple shared decoder to decode coordinates and content, ignoring the internal connection. Compared with other methods that need to correct arbitrarily shaped text into a regular shape, the need for sampling and independent design modules is reduced, thereby reducing the complexity of the model. Since the recognition result is better, the detection result can be guided and corrected by the recognition decoder. Secondly, in the aspect of general scene text recognition technology, the speed and model complexity problems of RNN have always been important bottlenecks of recognition algorithms. Solving this problem helps to improve the speed and performance of the recognition model, reduce the occupancy of machine computing power, and broaden the deployment scenario. The embodiments of the present application use the Transformer encoding and decoding structure for most general text line recognition scenarios to alleviate the modality difference between image pixels and text sequences in an end-to-end manner, and discard the commonly used serial RNN module and the architecture of multi-layer fully connected layers with high time consumption under the condition that the recognition accuracy remains basically unchanged, and realize text recognition simply and efficiently.

[0254] (8) The embodiments of the present application use the Transformer encoding and decoding structure to alleviate the modality difference between image pixels and text sequences in an end-to-end manner, and discard the commonly used serial RNN module and the architecture of multi-layer fully connected layers with high time consumption under the condition that the recognition accuracy remains basically unchanged, and realize text recognition simply and efficiently.

[0255] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.< / eos> < / eos> < / pad> < / pad> < / eos> < / start>

Claims

1. A text recognition method, characterized in that, The method includes: Obtain a to-be-recognized image including text content, and perform image encoding on the to-be-recognized image to obtain image features of the to-be-recognized image; Based on the image features, perform position decoding on the to-be-recognized image to obtain at least one target pixel point position in the to-be-recognized image, where the target pixel point position is used to indicate the position of the target pixel point with text in the to-be-recognized image; Perform position encoding on each of the target pixel point positions to obtain target position features, and combine the target position features and the image features to perform text decoding on the to-be-recognized image to obtain the text content in the to-be-recognized image; Combine the text content and the target pixel point position to determine the text recognition result of the to-be-recognized image.

2. The method according to claim 1, wherein The position decoding is implemented through a position decoding network, the position decoding network includes a position recognition layer and a position correction layer, and the image features include pixel point features respectively corresponding to each pixel point in the to-be-recognized image; The performing position decoding on the to-be-recognized image based on the image features to obtain at least one target pixel point position in the to-be-recognized image includes: Invoke the position recognition layer to perform position recognition on the to-be-recognized image based on the image features to obtain at least one candidate pixel point position in the to-be-recognized image; Invoke the position correction layer to perform position correction on each of the candidate pixel point positions based on the pixel point features respectively corresponding to each of the candidate pixel point positions to obtain at least one target pixel point position in the to-be-recognized image.

3. The method according to claim 2, characterized in that, The invoking the position correction layer to perform position correction on each of the candidate pixel point positions based on the pixel point features respectively corresponding to each of the candidate pixel point positions to obtain at least one target pixel point position in the to-be-recognized image includes: When the number of the candidate pixel point positions is one, determine the candidate pixel point position as the target pixel point position; When the number of the candidate pixel point positions is multiple, invoke the position correction layer to perform position correction on each of the candidate pixel point positions based on the pixel point features respectively corresponding to each of the candidate pixel point positions to obtain correction information corresponding to each of the candidate pixel point positions; For each of the candidate pixel point positions, when the correction information corresponding to the candidate pixel point position indicates that the candidate pixel point position is the position with text in the to-be-recognized image, determine the candidate pixel point position as the target pixel point position.

4. The method according to claim 1, characterized in that, The text decoding is implemented through a text decoding network, the text decoding network includes a fusion layer and a text decoding layer, and the combining the target position features and the image features to perform text decoding on the to-be-recognized image to obtain the text content in the to-be-recognized image includes: Invoke the fusion layer to fuse the target position features and the image features to obtain fusion features; Invoke the text decoding layer to perform text decoding on the to-be-recognized image based on the fusion features to obtain the text content in the to-be-recognized image.

5. The method according to claim 4, characterized in that, The fusion layer includes a position fusion layer and a feature fusion layer. Invoking the fusion layer to fuse the target position feature and the image feature to obtain a fusion feature includes: For each of the target pixel positions, determining, as the reference pixel position corresponding to the target pixel position, the pixel positions in the image to be recognized whose distances from the target pixel position are less than a distance threshold; Performing position encoding on the reference pixel positions to obtain reference position features; and invoking the position fusion layer to perform position fusion on the target position feature and the reference position features to obtain a position fusion feature; Invoking the feature fusion layer to fuse the position fusion feature and the image feature to obtain the fusion feature.

6. The method according to claim 1, wherein Combining the text content and the target pixel positions to determine the text recognition result of the image to be recognized includes: Aggregating the target pixel positions to obtain at least one target position group, where the distances between the target pixel positions in the target position group are less than a preset distance threshold; Combining the text content and the target position group to determine the text recognition result of the image to be recognized.

7. The method according to claim 6, wherein The text content includes the sub-text contents corresponding to the respective target pixel positions. Aggregating the target pixel positions to obtain at least one target position group includes: Performing position aggregation on the target pixel positions to obtain at least one candidate position group, where the distances between the target pixel positions in the candidate position group are less than the preset distance threshold; When the number of candidate position groups is one, determining the candidate position group as the target position group; When the number of candidate position groups is multiple, selecting at least one target position group from the multiple candidate position groups based on the respective sub-text contents.

8. The method according to claim 7, characterized in that, Selecting at least one target position group from the multiple candidate position groups based on the respective sub-text contents includes: Performing the following processing for each candidate position group respectively: Performing text content semantic analysis on the sub-text contents corresponding to the target pixel positions in the candidate position group to obtain a semantic analysis result; When the semantic analysis result indicates that the sub-text contents in the candidate position group can form a complete semantics, determining the candidate position group as the target position group.

9. The method according to claim 6, wherein The text content includes the sub-text contents corresponding to the respective target pixel positions, and the text recognition result includes the sub-text recognition results corresponding to the respective target position groups; Combining the text content and the target position group to determine the text recognition result of the image to be recognized includes: Performing the following processing for each target position group respectively: When the number of target pixel positions in the target position group is one, determining the sub-text content corresponding to the target pixel position as the sub-text recognition result corresponding to the target position group; When the number of the target pixel point positions in the target position group is multiple, text fusion is performed on the sub-text contents corresponding to the target pixel point positions in the target position group to obtain the sub-text recognition result corresponding to the target position group.

10. A training method for a text recognition model, characterized in that, The method includes: Obtaining a to-be-recognized image sample and at least one sample label carried by the to-be-recognized image sample, where the sample label is used to indicate the label text content corresponding to the target pixel point position in the to-be-recognized image sample; Wherein, the target pixel point position is used to indicate the position where the target pixel point with text exists in the to-be-recognized image sample, and the number of the sample labels is less than the total number of characters in the text in the to-be-recognized image sample; Invoking an initial text recognition model to perform text recognition on the to-be-recognized image sample to obtain the text recognition result of the to-be-recognized image sample; Combining the sample label and the text recognition result to train the initial text recognition model to obtain the text recognition model, where the text recognition model is used to recognize the text content in the to-be-recognized image.

11. The method according to claim 10, wherein The initial text recognition model includes an image encoding network, a position decoding network, and a text decoding network. Invoking the initial text recognition model to perform text recognition on the to-be-recognized image sample to obtain the text recognition result of the to-be-recognized image sample includes: Invoking the image encoding network to perform image encoding on the to-be-recognized image sample to obtain the sample image feature of the to-be-recognized image sample; Invoking the position decoding network to perform position decoding on the to-be-recognized image sample based on the sample image feature to obtain at least one predicted pixel point position in the to-be-recognized image; Performing position encoding on each of the predicted pixel point positions to obtain a predicted position feature, and invoking the text decoding network to perform text decoding on the to-be-recognized image sample to obtain the predicted text content in the to-be-recognized image sample; Aggregating each of the predicted pixel point positions to obtain at least one predicted position group, and combining the text content and the predicted position group to determine the text recognition result of the to-be-recognized image sample.

12. The method according to claim 10, wherein The text recognition result includes sub-text recognition results respectively corresponding to each predicted position group. Combining the sample label and the text recognition result to train the initial text recognition model to obtain the text recognition model includes: For each of the sub-text recognition results, obtaining the similarity between the text content indicated by the sub-text recognition result and each of the label text contents, and determining the loss value corresponding to the sub-text recognition result based on the maximum similarity; Training the initial text recognition model based on the loss value to obtain the text recognition model.

13. The method according to claim 12, characterized in that, Training the initial text recognition model based on the loss value to obtain the text recognition model includes: When the number of the sub-text recognition results is one, determining the loss value corresponding to the sub-text recognition result as the target loss value; When the number of the sub-text recognition results is multiple, sum up the loss values corresponding to the respective sub-text recognition results to obtain the target loss value; Based on the target loss value, train the initial text recognition model to obtain the text recognition model.

14. A text recognition device, characterized in that, The device includes: An image encoding module, configured to obtain a to-be-recognized image including text content, and perform image encoding on the to-be-recognized image to obtain the image features of the to-be-recognized image; A position decoding module, configured to perform position decoding on the to-be-recognized image based on the image features to obtain at least one target pixel point position in the to-be-recognized image, where the target pixel point position is used to indicate the position of the target pixel point with text in the to-be-recognized image; A position encoding module, configured to perform position encoding on each of the target pixel point positions to obtain target position features, and combine the target position features and the image features to perform text decoding on the to-be-recognized image to obtain the text content in the to-be-recognized image; A determination module, configured to determine the text recognition result of the to-be-recognized image by combining the text content and the target pixel point position.

15. A training device for a text recognition model, characterized in that, The device includes: An acquisition module, configured to acquire a to-be-recognized image sample and at least one sample label carried by the to-be-recognized image sample, where the sample label is used to indicate the label text content corresponding to the target pixel point position in the to-be-recognized image sample; wherein, the target pixel point position is used to indicate the position of the target pixel point with text in the to-be-recognized image sample, and the number of the sample labels is less than the total number of characters in the text of the to-be-recognized image sample; A recognition module, configured to call an initial text recognition model to perform text recognition on the to-be-recognized image sample to obtain the text recognition result of the to-be-recognized image sample; A training module, configured to train the initial text recognition model by combining the sample label and the text recognition result to obtain the text recognition model, where the text recognition model is used to recognize the text content in the to-be-recognized image.

16. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions or a computer program; A processor, configured to implement the method according to any one of claims 1 to 13 when executing the computer-executable instructions or the computer program stored in the memory.

17. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.

18. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, The computer program or the computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.