Character recognition method and training method and device of character recognition model

By introducing a style branch network into the text recognition model, the style features of the document image are acquired and decoupled from the text features, the problem of poor recognition effect of handwritten Chinese characters in the prior art is solved, and higher recognition accuracy and reliability are achieved.

CN120071367APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311638887.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the existing handwritten Chinese character recognition technology only recognizes based on text features, the recognition effect is poor, making it difficult to effectively extract the text content in the document.

Method used

By combining the style characteristics of the text content, the style branch network is used to obtain the style characteristics of the document image and decouple it from the text features to improve the recognition accuracy.

Benefits of technology

It improves the recognition effect of text content, improves the recognition accuracy of text recognition models for different style samples, and improves the accuracy and reliability of text recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071367A_ABST
    Figure CN120071367A_ABST
Patent Text Reader

Abstract

The invention discloses a character recognition method and a character recognition model training method and device, and relates to the technical field of machine vision. A computing device not only inputs a document image into a character recognition model to obtain character features of the document image, but also inputs a plurality of sub-images obtained by dividing the document image into the character recognition model to obtain style features of the document image. The character feature indicates the first character recognition result of the character content, and the style feature indicates the writing mode of the character content, that is, after the computing device decouples the style feature and the character feature, when the computing device performs character recognition in combination with the first character recognition result of the character content and the writing mode, the character recognition result of the character content is obtained. The result of the second character recognition result can reflect the real character content in the document image more accurately. Therefore, the accuracy of the character recognition result of the character content determined by the computing device is improved, and the recognition effect of the character content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine vision, and particularly to a method for text recognition, a method for training a text recognition model, and an apparatus therefor. Background Art

[0002] As an important part of intelligent document processing, handwritten Chinese character recognition has a wide range of application scenarios in many business scenarios. During the operation of an enterprise, a large number of document images containing handwritten Chinese characters, such as contracts and forms, have been accumulated in various departments and businesses. With the help of handwritten Chinese character recognition algorithms, the content of the documents can be effectively extracted. Usually, a neural network model is used to extract the image features of the document image, and based on these image features, the text in the document image is predicted. Since the above-mentioned image features are the text features of the document image, this results in a poor recognition effect of the text content during the process of recognizing the document image only based on the text features. Summary of the Invention

[0003] The present application provides a method for text recognition, a method for training a text recognition model, and an apparatus therefor, which recognize the text in the document image by combining the style features of the text content, improve the recognition effect of the text content, and are beneficial to improving the accuracy of text recognition.

[0004] The present application adopts the following technical solutions.

[0005] In a first aspect, the present application provides a method for text recognition. This method for text recognition can be executed by a computing device or a chip in the computing device. Taking the execution of the method for text recognition of the present application by the computing device as an example, this method for text recognition includes: the computing device obtains a document image containing text content and divides the document image to obtain a plurality of sub-images. And, the computing device inputs the document image and the plurality of sub-images into a text recognition model, and outputs the text features and style features of the document image; wherein, the text features indicate the first text recognition result of the text content, and the aforementioned text recognition model includes a style branch network, and this style branch network is used to obtain the style features of the document image, and the style features indicate the writing manner of the text content. Finally, the computing device obtains the second text recognition result of the document image according to the text features and the style features.

[0006] In this application, the computing device not only inputs the document image into the text recognition model to obtain the text features of the document image, but also inputs multiple sub-images obtained by dividing the document image into the text recognition model to obtain the style features of the document image. Since the text features indicate the first text recognition result of the text content and the style features indicate the writing style of the text content, when the computing device combines the first text recognition result and the writing style of the text content for text recognition, the result of the second text recognition result can more accurately reflect the true text content in the document image. Thus, the accuracy of the text recognition result of the text content determined by the computing device is improved, and the recognition effect of the text content is improved. That is to say, after the computing device decouples the style features from the text features, it can effectively improve the recognition accuracy of the text recognition model for samples of different styles (document images containing text content).

[0007] Combined with the text recognition method provided in the first aspect, in an alternative implementation, the foregoing text recognition model further includes a text branch network, and the text branch network includes one or more of the following combinations: a convolutional layer, a fully connected layer, a recurrent neural network layer, and a deep neural network layer based on a self-attention mechanism. The style branch network includes one or more of the following combinations: a convolutional layer, a fully connected layer, and a deep neural network layer based on a self-attention mechanism. The foregoing computing device inputs the document image and multiple sub-images into the text recognition model, and outputs the text features and style features of the document image, including: the computing device inputs the document image into the text branch network to obtain the text features of the document image; and the computing device inputs the multiple sub-images into the style branch network to obtain the style features of the document image. By using branch networks that extract different features to obtain the text features and style features of the document image, the decoupling between the style features and text features of the text content is achieved, thereby improving the recognition accuracy of the text recognition model for the text content.

[0008] Combined with the text recognition method provided in the first aspect, in an alternative implementation, the computing device inputs multiple sub-images into the style branch network to obtain the style features of the document image, including: for each sub-image among the multiple sub-images, the computing device inputs the sub-image into the style branch network to obtain the style features of the sub-image. And the computing device fuses the multiple style features of the multiple sub-images to obtain the style features of the document image. In this application, the content of an image is divided into multiple sub-images, so that the text recognition model will recognize these multiple sub-images as images with different style features, achieving the decoupling between the style features of different sub-images. The style features obtained after the computing device fuses the style features of these multiple sub-images can more accurately reflect the writing style of the document image, thereby improving the recognition accuracy of the text recognition model for the text content.

[0009] Combined with the text recognition method provided in the first aspect, in an optional implementation, the text features of the document image include multiple text features. The aforementioned computing device obtains a second text recognition result of the document image according to the text features and the style features, including: for each text feature among the multiple text features, the computing device fuses the text feature and the style feature to obtain a fused feature. And, the computing device analyzes the multiple fused features of the multiple text features to obtain the second text recognition result of the document image. In this application, the computing device enhances features of different styles within different text features, which is beneficial to improving the recognition effect of the text recognition model on the document image, thereby enhancing the recognition effect of the text recognition model on samples of different styles.

[0010] Combined with the text recognition method provided in the first aspect, in an optional implementation, the text recognition method provided in this application further includes: the computing device obtains a first confidence level of the first text recognition result and a second confidence level of the second text recognition result; and, the computing device determines the larger confidence level among the first confidence level and the second confidence level, and outputs the text recognition result corresponding to the larger confidence level. Selecting the text recognition result with a larger confidence level among different text recognition results for output further improves the reliability and accuracy of the output text recognition result.

[0011] Combined with the text recognition method provided in the first aspect, in an optional implementation, the text recognition method provided in this application further includes: the computing device performs image enhancement processing on the document image to obtain an enhanced image, and inputs the enhanced image into the style branch network to obtain the style feature of the enhanced image. And, the computing device trains the style branch network according to the style feature of the document image and the style feature of the enhanced image to obtain a trained style branch network. Using the style feature of the enhanced image corresponding to the document image and the style feature of the document image to implement the training of the style branch network, on the basis of decoupling the text features and style features in the document image, it realizes feature enhancement based on self-supervised style feature extraction, which is beneficial to improving the recognition accuracy of the text recognition model for text content with different style features.

[0012] Second aspect, the present application provides a method for training a text recognition model. The method for training the text recognition model can be executed by a computing device or a chip in the computing device. Taking the execution of the method for training the text recognition model of the present application by the computing device as an example, the method for training the text recognition model includes: The computing device acquires a document image containing text content, and performs image enhancement processing on the document image to obtain an enhanced image. The computing device inputs the document image and the enhanced image into the text recognition model, and outputs the text feature and style feature of the document image, and the style feature of the enhanced image; wherein, the text feature indicates the text recognition result of the text content, and the text recognition model includes a style branch network, and the style branch network is used to obtain the style feature of the image, and the style feature indicates the writing style of the text content in the image. And, the computing device trains the style branch network according to the style feature of the document image and the style feature of the enhanced image to obtain the trained text recognition model.

[0013] In the present application, the style feature of the enhanced image corresponding to the document image and the style feature of the document image are used to train the style branch network. On the basis of decoupling the text feature and style feature in the document image, feature enhancement based on self-supervised style feature extraction is realized, which is beneficial to improving the recognition accuracy of the text recognition model for text content with different style features.

[0014] Combined with the method for training the text recognition model provided in the second aspect, in an optional implementation manner, the method for training the text recognition model provided in the present application further includes: The computing device acquires an image to be recognized, and divides the image to be recognized into multiple sub-images. The computing device inputs the image to be recognized and the multiple sub-images into the trained text recognition model, and outputs the text feature and style feature of the image to be recognized. The text feature of the image to be recognized indicates the first text recognition result of the text content in the image to be recognized, and the style feature of the image to be recognized indicates: the writing style of the text content in the image to be recognized. And, the computing device obtains a second text recognition result according to the text feature and style feature of the image to be recognized.

[0015] Since the text feature indicates the first text recognition result of the text content in the image to be recognized, and the style feature indicates the writing style of the text content in the image to be recognized, therefore, when the computing device performs text recognition by combining the first text recognition result and the writing style of the text content, the result of the second text recognition result can more accurately reflect the real text content in the image to be recognized. Thus, the accuracy of the text recognition result of the text content determined by the computing device is improved, and the recognition effect of the text content is improved. That is to say, after the computing device decouples the style feature and the text feature, the recognition accuracy of the text recognition model for samples of different styles (any document image containing text content) can be effectively improved.

[0016] Combined with the text recognition method provided in the first aspect and the training method of the text recognition model provided in the second aspect, in an optional implementation manner, the above-mentioned text content includes handwritten Chinese characters.

[0017] Combined with the text recognition method provided in the first aspect and the training method of the text recognition model provided in the second aspect, in an optional implementation manner, there are no overlapping regions among the above-mentioned multiple sub-images.

[0018] Combined with the text recognition method provided in the first aspect and the training method of the text recognition model provided in the second aspect, in an optional implementation manner, the above-mentioned multiple sub-images are of the same or different sizes.

[0019] In a third aspect, the present application provides a text recognition device, and the text recognition device includes a module that executes the method according to the first aspect or any one of the implementation manners in the first aspect.

[0020] Exemplarily, the text recognition device includes: an image acquisition module, a model processing module, and a text recognition module. The image acquisition module is used to acquire a document image containing text content and divide the document image to obtain multiple sub-images. The model processing module is used to input the document image and the multiple sub-images into a text recognition model and output the text feature and style feature of the document image, where the text feature indicates the first text recognition result of the text content, the text recognition model includes a style branch network, and the style branch network is used to acquire the style feature of the document image, and the style feature indicates the writing manner of the text content. The text recognition module is used to obtain the second text recognition result of the document image according to the text feature and the style feature.

[0021] Optionally, the text recognition device further includes: a confidence level acquisition module and an output module. The confidence level acquisition module is used to: acquire the first confidence level of the first text recognition result and the second confidence level of the second text recognition result. The output module is used to: determine the larger confidence level between the first confidence level and the second confidence level and output the text recognition result corresponding to the larger confidence level.

[0022] Optionally, the text recognition device further includes: an image enhancement module and a model training module. The image enhancement module is used to: perform image enhancement processing on the document image to obtain an enhanced image. The model processing module is further used to: input the enhanced image into the style branch network to obtain the style feature of the enhanced image. The model training module is used to: train the style branch network according to the style feature of the document image and the style feature of the enhanced image to obtain a trained style branch network.

[0023] Optionally, the text content includes handwritten Chinese characters.

[0024] Optionally, there are no overlapping regions among the multiple sub-images.

[0025] Optionally, the sizes of the multiple sub-images are the same or different.

[0026] In a fourth aspect, the present application provides a training device for a text recognition model, and the training device for the text recognition model includes a module that executes the method according to the second aspect or any implementation manner of the second aspect.

[0027] Exemplarily, the training device includes: an image acquisition module, a model processing module, and a model training module. The image acquisition module is configured to: acquire a document image containing text content, and perform image enhancement processing on the document image to obtain an enhanced image. The model processing module is configured to: input the document image and the enhanced image into the text recognition model, and output the text feature and style feature of the document image, and the style feature of the enhanced image. Among them, the text feature indicates the text recognition result of the text content. The text recognition model includes a style branch network, and the style branch network is used to obtain the style feature of the image, and the style feature indicates the writing manner of the text content in the image. The model training module is configured to: train the style branch network according to the style feature of the document image and the style feature of the enhanced image to obtain a trained text recognition model.

[0028] Optionally, the image acquisition module is further configured to: acquire an image to be recognized, and divide the image to be recognized into multiple sub-images. The model processing module is further configured to: input the image to be recognized and the multiple sub-images into the trained text recognition model, and output the text feature and style feature of the image to be recognized; the text feature of the image to be recognized indicates the first text recognition result of the text content in the image to be recognized, and the style feature of the image to be recognized indicates: the writing manner of the text content in the image to be recognized. The training device for the text recognition model provided by the present application further includes: a text recognition module configured to: obtain the second text recognition result of the image to be recognized according to the text feature and style feature of the image to be recognized.

[0029] In a fifth aspect, the present application provides a chip. The chip includes: a control circuit and an interface circuit. The interface circuit is configured to acquire a document image, and cooperate with the control circuit to execute the method according to any implementation manner of the first aspect and the second aspect.

[0030] In a sixth aspect, the present application provides an electronic device. The electronic device includes: a processor and a memory. The memory is configured to store a set of computer instructions, and when the processor executes the set of computer instructions, the method according to any implementation manner of the first aspect and the second aspect is executed. For example, the electronic device may include, but is not limited to: a computing device, a server, a terminal, or other types of electronic devices with data processing capabilities.

[0031] In a seventh aspect, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on a computing device, the computing device executes the method according to any one of the implementation manners in the first aspect and the second aspect above.

[0032] In an eighth aspect, the present application provides a computer program product. When the computer program product runs on a computing device, the computing device executes the method according to any one of the implementation manners in the first aspect and the second aspect above.

[0033] The beneficial effects of the third aspect to the eighth aspect above can be referred to the description of any one of the implementation manners in the first aspect or the second aspect, and will not be elaborated here. Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the architecture of a text recognition system provided by the present application;

[0035] Figure 2 It is a schematic flow chart of a text recognition method provided by the present application Figure 1 ;

[0036] Figure 3 It is a schematic flow chart of a text recognition method provided by the present application Figure 2 ;

[0037] Figure 4 It is a schematic flow chart of a text recognition method provided by the present application Figure 3 ;

[0038] Figure 5 It is a schematic flow chart of a text recognition method provided by the present application Figure 4 ;

[0039] Figure 6 It is a schematic flow chart of a method for training a text recognition model provided by the present application;

[0040] Figure 7 It is a schematic diagram of an image in a private dataset provided by the present application;

[0041] Figure 8A It is a schematic diagram of a visualization result of style features provided by the present application;

[0042] Figure 8B It is a schematic flow chart of a text recognition method provided by the present application Figure 5 ;

[0043] Figure 9 It is a schematic diagram of the structure of a text recognition device provided by the present application;

[0044] Figure 10 Schematic structural diagram of a training device provided for this application;

[0045] Figure 11 Schematic structural diagram of an electronic device provided for this application. Specific embodiments

[0046] This application provides a text recognition method. The computing device not only inputs a document image into a text recognition model to obtain the text features of the document image, but also inputs multiple sub-images obtained by dividing the document image into the text recognition model to obtain the style features of the document image. Since the text features indicate the first text recognition result of the text content and the style features indicate the writing manner of the text content, when the computing device combines the first text recognition result and the writing manner of the text content for text recognition, the result of the second text recognition result can more accurately reflect the real text content in the document image. Thus, the accuracy of the text recognition result of the text content determined by the computing device is improved, and the recognition effect of the text content is improved. That is to say, after the computing device decouples the style features and the text features, the recognition accuracy of the text recognition model for samples of different styles (document images containing text content) can be effectively improved.

[0047] For the sake of clear and concise description of the following embodiments, a brief introduction to the related technologies is given first.

[0048] (1) Neural network

[0049] A neural network can be composed of neurons. A neuron can refer to an operation unit that takes x s and the intercept 1 as inputs. The output of this operation unit satisfies the following formula (1).

[0050]

[0051] where h W,b is the output of the operation unit, x is the input of the operation unit, s = 1, 2,..., n, n is a natural number greater than 1, and W s is x sThe weight is w, b is the bias of the computing unit, and f is the activation function of the neuron, which is used to introduce non-linearity into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons. The weight represents the strength of the connection between different neurons. The weight determines the influence of the input on the output. A weight close to 0 means that changing the input does not change the output. A negative weight means that increasing the input reduces the output.

[0052] In some feasible embodiments, the input signal of the neural network can be various forms of signals such as video signals, image signals, matrix data, graph structure data (graph data), etc. The input signal of the neural network also includes various other computer-processable engineering signals, which will not be listed one by one here. If a neural network is used for deep learning in base station text recognition, etc., the accuracy of base station text recognition by the neural network can be improved.

[0053] (2) Graph Neural Network

[0054] A graph neural network refers to a general term for algorithms that use neural networks to learn graph data, extract and discover the features and patterns in graph data, and meet the requirements of graph learning tasks such as clustering, classification, prediction, segmentation, and generation. A graph neural network is a framework for directly learning graph data using deep learning. By formulating certain strategies on the nodes and edges in the graph, the graph data is transformed into a standardized and standard representation and input into various different neural networks for training, achieving excellent results in tasks such as text recognition, node classification, edge information propagation, and graph clustering.

[0055] The graph neural network mentioned in this application is not limited to a specific type. For example, graph convolutional network (GCN), graph auto-encoder (gAE), graph generative network (GGN), graph recurrent network (GRN), graph attention network (GAT), etc.

[0056] The following will describe the implementation manners of the embodiments of the present application in detail with reference to the accompanying drawings.

[0057] Figure 1The figure is a schematic architecture diagram of a text recognition system provided by this application. As Figure 1 shown, the text recognition system 100 includes a computing device 110, a training device 120, a database 130, a terminal device 140, a data storage system 150, and a data acquisition device 160.

[0058] The computing device 110 can be a terminal, such as a computer, a mobile phone terminal, a tablet computer, a laptop computer, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, an extended reality (ER) device, a camera, or an in-vehicle computer, etc. The computing device 110 can also be an edge device (for example, a box carrying a processing-capable chip), etc. In this application, the computing device 110 can be a computing device connected to a base station, or a computing device deployed in a base station, such as a server or a cloud device, etc.

[0059] The training device 120 can be a terminal, or other computing devices that support integer computing, such as a server or a cloud device, etc.

[0060] As a possible embodiment, the computing device 110 and the training device 120 are deployed on different physical devices (such as: servers in a server or a cluster), or the computing device 110 and the training device 120 are different physical devices. Exemplarily, the computing device 110 and the training device 120 are processors deployed on different physical devices. For example, the computing device 110 can be a graphics processing unit (GPU), a central processing unit, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The training device 120 can be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of this application solution.

[0061] In another possible embodiment, the computing device 110 and the training device 120 are deployed on the same physical device, or the computing device 110 and the training device 120 are the same physical device.

[0062] The data acquisition device 160 is used to acquire training data and store the training data in the database 120. The data acquisition device 160 and the computing device 110 and the training device 120 may be the same or different devices. In the embodiments of the present application, the data acquisition device 160 may be a camera, a camera phone, a mobile phone, a tablet computer, or a computing device or network device with an image acquisition function.

[0063] The training device 120 is used to train the neural network using the training data until the loss function in the neural network converges, and when the value of the loss function is less than a specific threshold, the neural network training is completed, so that the neural network reaches a certain accuracy. For example, the training device 120 uses the style features of the document image and the style features of the enhanced image corresponding to the document image as the input training data, and the character recognition result of the document image as the output to train the neural network. Another example is that the training device 120 uses the style features of the document image and the style features of the enhanced image corresponding to the document image as the input training data, and the difference between the two style features as the output to train the neural network. Or, if all the training data in the database 130 is used for training, the neural network training is completed, and the trained neural network has functions such as character recognition, extraction of character features of text content, and extraction of style features. Furthermore, the training device 120 configures the trained neural network 101 to the computing device 110. The computing device 110 is used to implement the character recognition function for the document image containing text content according to the trained neural network 101.

[0064] In this embodiment, the above neural network model may be referred to as a character recognition model. When the character recognition model is used to recognize handwritten Chinese characters, the character recognition model may also be referred to as a Chinese character recognition model.

[0065] Optionally, the neural network 101 is used to recognize the text content in the document image and output the character recognition result of the document image. Among them, the neural network 101 may be a network type suitable for character recognition. For example, the neural network 101 is a convolutional neural network (CNN), a recurrent neural network, or a graph neural network, etc.

[0066] In some embodiments, the computing device 110 and the training device 120 are the same computing device, and the computing device can configure the trained neural network 101 to itself and use the trained neural network 101 to implement the above-mentioned character recognition function.

[0067] In some other embodiments, the training device 120 may configure the trained neural network 101 to multiple computing devices 110. Each computing device 110 utilizes the trained neural network 101 to implement the above-mentioned text recognition function.

[0068] In combination with the text recognition system 100, the text recognition method provided in this embodiment can be applied to any scenario that requires text recognition of text content. The text content may include, but is not limited to: Chinese characters (simplified or traditional Chinese), English, or text of other language types, etc. When the text recognition model is applied to text content of different language types, the training and prediction processes of the text recognition model will be adaptively adjusted according to the differences in input data and output data, which will not be elaborated here.

[0069] In some embodiments, the text recognition method can be applied to various text recognition scenarios in a city, such as identification or document verification in commercial areas, schools, parks, stadiums, etc.

[0070] In some other embodiments, the text recognition system 100 can use the text recognition method to assist in recognizing handwritten content, such as document recognition of financial bills or recognition of logistics forms, etc.

[0071] It should be noted that in actual applications, the training data maintained in the database 130 may not necessarily all come from the data acquisition device 160, and it is also possible to be received from other devices. In addition, the training device 120 does not necessarily train the neural network completely based on the training data maintained in the database 130, and it is also possible to obtain training data from the cloud or other places to train the neural network. The above description should not be used as a limitation to the embodiments of the present application.

[0072] Further, according to the functions performed by the computing device 110, the computing device 110 can be further subdivided into an architecture as shown in Figure 1 As shown in Figure 1 The computing device 110 is configured with a computing module 111, an I / O interface 112, and a preprocessing module 113.

[0073] Taking the computing device 110 as a computing device connected to access stations such as base stations and wireless access points as an example, the computing module 111 can be a GPU, CPU, other general-purpose processors, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. on a vehicle-mounted computer.

[0074] For example, the computing module 111 is used to run the neural network 101 to implement the function of the neural network 101 for text recognition of text content, so as to obtain the text recognition result of the document image containing the aforementioned text content.

[0075] The I / O interface 112 is used for data interaction with external devices. A user can input data to the I / O interface 112 through the terminal device 140, such as an instruction for instructing the computing device 110 to start performing an optical character recognition method on a document image containing text content. Additionally, the input data can also come from the database 130.

[0076] The preprocessing module 113 is used to perform preprocessing based on the input data received by the I / O interface 112. In the embodiments of the present application, the preprocessing module 113 can be used to generate training data, such as a training set, a validation set, and a test set, based on the input data received from the I / O interface 112. Optionally, the preprocessing module 113 can also perform preprocessing operations such as denoising on the input data, such as matrix data and graph data of a base station, to eliminate irrelevant information and restore useful real information.

[0077] When the computing device 110 performs preprocessing on the input data, or when the computing module 111 of the computing device 110 performs related processing such as computing, the computing device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or can also store the data and instructions obtained from the corresponding processing into the data storage system 150.

[0078] Finally, the I / O interface 112 returns the processing result to the terminal device 140, so as to provide it to the user for the user to view the processing result. In the embodiments of the present application, the processing result can be the optical character recognition result of the document image, such as the real text content included in the document image.

[0079] The terminal device 140 can be used as a data acquisition end to acquire the input data input to the I / O interface 112 as shown in the figure and the processing result output from the I / O interface 112 as new sample data, and store them in the database 130. Of course, the sample data can also be acquired without passing through the terminal device 140, but the I / O interface 112 can store the input data input to the I / O interface 112 as shown in the figure and the processing result output from the I / O interface 112 as new sample data in the database 130.

[0080] Figure 1 This is only a schematic diagram of a system architecture provided by the embodiments of the present application. Figure 1 The positional relationship between the devices, components, modules, etc. shown does not constitute any limitation. According to the user's requirements for optical character recognition, the optical character recognition system and the computing device may include more or fewer hardware components, which are not limited in the present application. For example, in Figure 1 In this case, the data storage system 150 is an external memory relative to the computing device 110. In other cases, the data storage system 150 can also be placed in the computing device 110.

[0081] The scenarios to which the present application can be applied include, but are not limited to: document processing and digitization scenarios such as contracts, invoices, handwritten manuscripts, etc. The present application can effectively convert the handwritten Chinese character images in the document into corresponding electronic text information, realizing electronic archiving and document processing. The present application can also be applied to scenarios that require the introduction of handwritten Chinese character writing style information, such as handwriting comparison and image generation.

[0082] Next, based on the Figure 1 shown text recognition system and computing device, the text recognition method provided by the present application and the model training method of the text recognition model will be described in detail.

[0083] As Figure 2 shown, Figure 2 is a flowchart of a text recognition method provided by the present application. Figure 1 The computing device that executes the text recognition method provided by the embodiments of the present application may refer to: computing device 110, the processor or processor core (core) included in computing device 110, such as a CPU or ASIC, other general-purpose processors, DSPs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. As Figure 2 shown, here, taking the computing device executing the text recognition method provided in this embodiment as an example for description, the text recognition method provided in this embodiment includes the following steps S210 to S240.

[0084] S210. The computing device acquires a document image containing text content.

[0085] In some alternative implementation manners, the above text content includes one or a combination of the following: Chinese characters (simplified or traditional Chinese), English, or text of other language types. It should be noted that the text content in the document image can include both handwritten text and text in a standard format (such as Song typeface, Kai typeface, etc.).

[0086] This embodiment will be described by taking the text content including handwritten Chinese characters as an example, and will not be repeated hereinafter.

[0087] For example, in Figure 2 , the text content in the document image is: "Plus, strictly prohibited from drawing images", but this text content cannot be directly recognized by the computing device.

[0088] S220. The computing device divides the document image obtained in S210 to obtain a plurality of sub-images.

[0089] Optionally, the number of these sub-images is an integer of two or more.

[0090] Exemplarily, there is no overlapping area between these multiple sub-images. For example, Figure 2 one of the sub-images includes the text content "strictly prohibited from drawing", and the other sub-image includes the text content "image information". In this way, the text contents included in different sub-images are different, which can be used to decouple the text in the same document image. It should be understood that in some alternative ways, the overlapping area is also called the coincidence area, redundant area, or identical area, etc.

[0091] In some possible examples, there may also be partially overlapping areas among these multiple sub-images.

[0092] In an alternative scenario, these multiple sub-images are of the same size. When the multiple sub-images are of the same size, the computing device can extract the style features of the sub-images in the same or similar way, which is beneficial to improving the efficiency of text recognition.

[0093] In another alternative scenario, these multiple sub-images are of different sizes. For example, each sub-image includes at least one complete piece of text. When the multiple sub-images are of different sizes, the adaptability between each sub-image and the text content in the document image is better, thereby improving the accuracy of the style features of the sub-images extracted by the computing device, which is beneficial to improving the accuracy of text recognition.

[0094] The above-mentioned size and division method of the sub-images are only alternative ways provided in this embodiment, and should not be construed as a limitation to this application.

[0095] In some cases, the computing device can also perform image enhancement processing on the multiple sub-images obtained by dividing the document image. For example, randomly select one or more images from the multiple sub-images for image enhancement processing to obtain the sub-images after image enhancement. The image enhancement processing mentioned in this example refers to processing certain degraded image features, such as edges, contours, contrast, etc. through a certain image processing method to improve the visual effect of the image, increase the clarity of the image, or highlight some "useful" information in the image, compress other "useless" information, and convert the image into a form more suitable for human or computer analysis and processing.

[0096] S230. The computing device inputs the document image and the multiple sub-images into the text recognition model, and outputs the text features and style features of the document image.

[0097] Among them, the text features of the document image indicate the first text recognition result of the text content. In some alternative cases, the first text recognition result may refer to the initial recognition result of the text content.

[0098] The style features of the above document image indicate the writing style of the text content. The writing style may include, but is not limited to, the following attributes: the thickness of the strokes, the inclination angle of the text in the document image, the similarity between strokes of the same type, the similarity between different types of strokes, the connection method between different strokes, the distance between different strokes, the proportion of a specified stroke in the entire text, and other attributes, etc.

[0099] The text recognition model provided in this embodiment may include: a backbone network, a text branch network, and a style branch network.

[0100] The backbone network is used to extract the general features of the image, and then the general features of the document image are passed into the text branch network, and the multiple groups of general features of multiple sub-images are passed into the style branch network, with one sub-image corresponding to one group of general features.

[0101] The above text branch model is used to extract the text features of the image and realize the preliminary recognition of the text content. Exemplarily, the text branch network includes one or several of the following combinations: a convolutional layer, a fully connected layer, a recurrent neural networks (RNN) layer, and a deep neural network layer based on the self-attention mechanism (i.e., a transformer).

[0102] The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolutional processing on the input signal. Each layer in the fully connected layer is a tiled structure composed of many neurons (such as 1×4096 or other parameters). The RNN layer will remember the previous information and apply it to the calculation of the current output, that is, the nodes between the hidden layers of this layer are no longer unconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the previous hidden layer. The transformer is a model that uses the attention mechanism to improve the model training speed.

[0103] For example, the word branch network includes a convolutional layer, a fully connected layer, an RNN layer, and a transformer. For a more detailed description of each network layer, reference can be made to the content of the general technology, which will not be elaborated here.

[0104] The above style branch network is used to obtain the style features of the document image, so as to determine the writing style of the text content in the document image. Exemplarily, the style branch network includes one or several of the following combinations: a convolutional layer, a fully connected layer, and a deep neural network layer based on the self-attention mechanism (transformer).

[0105] For example, the style branch network includes convolutional layers, fully connected layers, and transformers. For a more detailed description of each network layer, reference can be made to the content of common techniques, which will not be elaborated here.

[0106] As an alternative implementation, the computing device inputs the document image and multiple sub-images into the text recognition model, and outputs the text features and style features of the document image, which may include the steps shown as Figure 3 follows. Figure 3 This is a schematic flowchart of a text recognition method provided by this application. Figure 2 The above S230 may include the following S231 and S232.

[0107] S231. The computing device inputs the document image into the text branch network to obtain the text features of the document image.

[0108] Exemplarily, the computing device inputs the general features extracted from the document image through the backbone network into the text branch network, and learns the text features corresponding to the document image through the connectionist temporal classification (CTC) loss function (abbreviated as CTC loss).

[0109] For more content about the text branch network and text features, reference can be made to the description of the foregoing embodiments, which will not be elaborated here.

[0110] S232. The computing device inputs multiple sub-images into the style branch network to obtain the style features of the document image.

[0111] As a feasible example, the process of the above S232 may include the following steps: For each of the multiple sub-images, the computing device inputs the sub-image into the style branch network to obtain the style features of the sub-image, and fuses the multiple style features of the multiple sub-images to obtain the style features of the document image.

[0112] For example, assuming the number of multiple sub-images is 2, the computing device inputs the general features of the two sub-images into the style branch network, and obtains the style features of each sub-image through contrastive learning. The computing device takes the sub-images from the same document image as positive samples, and takes the sub-images from other images as negative samples, so as to perform self-supervised enhancement on the style features of the sub-images. Finally, the computing device fuses the style features corresponding to the two sub-images to obtain the style features of the document image.

[0113] In this embodiment, the computing device divides the content of an image into multiple sub-images, such that the text recognition model will recognize these multiple sub-images as images with different style features, achieving the decoupling between the style features of different sub-images. After the computing device fuses the style features of these multiple sub-images, the obtained style features can more accurately reflect the writing style of the document image, thereby improving the recognition accuracy of the text recognition model for the text content.

[0114] In summary, the computing device uses branch networks that extract different features to obtain the text features and style features of the document image, thereby achieving the decoupling between the style features and text features of the text content, and thus improving the recognition accuracy of the text recognition model for the text content.

[0115] Please continue to refer to Figure 2 , the text recognition method provided in this embodiment further includes the following S240.

[0116] S240: The computing device obtains a second text recognition result of the document image according to the text features and style features.

[0117] In this way, the computing device not only inputs the document image into the text recognition model to obtain the text features of the document image, but also inputs the multiple sub-images obtained by dividing the document image into the text recognition model to obtain the style features of the document image. Since the text features indicate the first text recognition result of the text content and the style features indicate the writing style of the text content, when the computing device performs text recognition by combining the first text recognition result and the writing style of the text content, the result of the second text recognition result can more accurately reflect the real text content in the document image. Thus, the accuracy of the text recognition result of the text content determined by the computing device is improved, and the recognition effect of the text content is improved. That is to say, after the computing device decouples the style features and text features, the recognition accuracy of the text recognition model for samples of different styles (document images containing text content) can be effectively improved.

[0118] As an optional implementation manner, the text features of the document image include multiple text features. For the above S240, this embodiment provides a feasible example, as Figure 4 shown, Figure 4 is a schematic flowchart of a text recognition method provided by the present application Figure 3 . The above S240 includes the following S241 and S242.

[0119] S241: For each text feature among the multiple text features, the computing device fuses the text feature and the style feature to obtain a fused feature.

[0120] The fusion methods may include, but are not limited to: adding or subtracting the vector corresponding to the text feature and the vector corresponding to the style feature, or fusing these two vectors through convolution or other types of operations to obtain a fused feature. For example, the fused feature indicates the recognition result of the text content and the writing style of the text content.

[0121] S242. The computing device analyzes multiple fused features of multiple text features to obtain a second text recognition result of the document image.

[0122] Exemplarily, the computing device randomly fuses the text feature and the style feature, that is, each text feature randomly selects one from the style features for fusion, and finally the fused feature uses a CTC decoder to predict the recognition result.

[0123] In this way, the computing device enhances features of different styles within different text features, which is beneficial to improving the recognition effect of the text recognition model on the document image, thereby enhancing the recognition effect of the text recognition model on samples of different styles.

[0124] It should be noted that in the text recognition model provided by the embodiments of the present application, there is not only the text feature of the document image (indicating the first text recognition result), but also the second text recognition result. To improve the accuracy of the text recognition model, on the basis of the above embodiments, the present application provides a possible implementation manner, as Figure 5 shown Figure 5 is a flowchart of a text recognition method provided by the present application. Figure 4 After S240 above, the text recognition method provided by this embodiment further includes the following steps S250 and S260.

[0125] S250. The computing device obtains a first confidence level of the first text recognition result and a second confidence level of the second text recognition result.

[0126] Exemplarily, the first confidence level is the confidence level determined by the text branch network using CTC loss1, and the second confidence level is the confidence level determined by the style branch network using CTC loss2. The above confidence levels are used to indicate the reliability of the text recognition results. For example, the first confidence level indicates the reliability of the first text recognition result, and the second confidence level indicates the reliability of the second text recognition result.

[0127] S260. The computing device determines the larger confidence level among the first confidence level and the second confidence level, and outputs the text recognition result corresponding to the larger confidence level.

[0128] In this embodiment, the text recognition model not only outputs an initial first text recognition result (the text features of the document image), but also outputs a second text recognition result obtained after fusing the style features. That is to say, the text recognition model outputs text recognition results at two places. The calculation combines the two text recognition results to select the best confidence (the larger confidence) as the final prediction result of the overall model, so as to further improve the reliability and accuracy of the output text recognition result.

[0129] In addition, please continue to refer to Figure 5 , the computing device can also set a gradient flow truncation measure before fusing the text features and style features. This gradient flow truncation can be used to prevent the gradient of the contrast loss (CTC loss2) of the second text recognition result from backpropagating to the style branch network, which is beneficial to maintaining the reliability of the style branch network, thereby improving the recognition accuracy of the text recognition model for text content.

[0130] In an alternative implementation, the computing device can also train and optimize the text recognition model. As Figure 6 shown, Figure 6 is a schematic flowchart of a training method for a text recognition model provided by this application. This training method can be executed by the above-mentioned computing device or training device, or a chip in the computing device or a chip in the training device. This embodiment takes the computing device executing this training method as an example for illustration. As Figure 6 shown in (A) and (B) in

[0131] S610. The computing device obtains a document image containing text content.

[0132] For the detailed process of S610, please refer to the content of S210 and will not be elaborated here.

[0133] S620. The computing device performs image enhancement processing on the document image to obtain an enhanced image.

[0134] Here, the image enhancement processing refers to processing some degraded image features, such as edges, contours, contrast, etc. through a certain image processing method to improve the visual effect of the image, increase the clarity of the image, or highlight some "useful" information in the image, compress other "useless" information, and convert the image into a form more suitable for human or computer analysis and processing.

[0135] Figure 6 One of the differences between (A) and (B) in

[0136] In an alternative scenario, as Figure 6 shown in (A) in

[0137] In another alternative scenario, as Figure 6 in (B) of , the enhanced image includes multiple enhanced sub-images obtained by dividing the document image into two or more sub-images and performing image enhancement processing on the divided sub-images.

[0138] S630. The computing device inputs the document image and the enhanced image into the text recognition model, and outputs the text features and style features of the document image and the style features of the enhanced image.

[0139] Regarding the content of the text recognition model, reference can be made to the description of the above embodiments, which will not be elaborated here.

[0140] As a feasible specific example, the computing device inputs the enhanced image into the style branch network in the text recognition model to obtain the style features of the enhanced image. The style features of the enhanced image indicate the writing style of the text content in the enhanced image.

[0141] S640. The computing device trains the style branch network according to the style features of the document image and the style features of the enhanced image to obtain the trained text recognition model.

[0142] Here, taking the style branch network in the text recognition model as an example, the computing device can train the style branch network according to the style features of the document image and the style features of the enhanced image to obtain the trained style branch network, thereby determining the above-mentioned trained text recognition model.

[0143] Figure 6 The second difference point between (A) and (B) in is that different positive and negative samples are used in the training.

[0144] For example, when there is only one enhanced image, as Figure 6 in (A) of , the computing device can use the style features of the document image and the style features of the enhanced image as positive samples, and the style features of other images as negative samples to train the style branch network to obtain the trained style branch network. Since the above-mentioned enhanced image is obtained by performing image enhancement processing on the document image, the above Figure 6 shown training method can also be called the model training process of self-supervised style feature extraction based on image enhancement.

[0145] Again, for example, when there are multiple enhanced images, as Figure 6In (B) among them, the computing device can use the style features of the enhanced sub-images as positive samples and the style features of other images as negative samples to train the style branch network to obtain the trained style branch network. Among them, the style features of multiple (multiple) enhanced sub-images may be different, which can be used to improve the recognition accuracy of the text recognition model for different style samples and also improve the generalization of the text recognition process of the text recognition model for different document images.

[0146] In this way, the computing device uses the style features of the enhanced image corresponding to the document image and the style features of the document image to train the style branch network. On the basis of decoupling the text features and style features in the document image, feature enhancement based on self-supervised style feature extraction is achieved, which is beneficial to improving the recognition accuracy of the text content with different style features by the text recognition model.

[0147] Refer to Figure 6 For the relevant description, the training method provided in this embodiment can be used alone or in combination with the text recognition method shown in the foregoing Figures 2 to 5 This application does not limit this.

[0148] To further illustrate the beneficial effects of the text recognition method and the training method of the text recognition model provided by this application, the beneficial effects of the embodiments of this application are compared and described below in combination with two open datasets, CASIA-HWDB and TAL_OCR_CHN, and two private datasets.

[0149] Regarding the content of the open datasets CASIA-HWDB and TAL_OCR_CHN, reference can be made to the description of the general technology, which will not be elaborated here. Figure 7 It is a schematic diagram of an image in a private dataset provided by this application. As Figure 7 shown, it shows the document image in Private Dataset 1. The embodiments of this application can also use financial bills or other types of document images as Private Dataset 2.

[0150] The effectiveness of the text recognition model provided in this application was verified on common SOTA models. In the experimental setup, two sets of comparative experiments were introduced in this embodiment to verify the effectiveness of the feature fusion method provided in this application, namely: Example ①, fusing the text features of a document image with the style features of that document image; Example ②, fusing the text features of a document image with the style features with noise; Example ③, fusing the text features of a document image with the style features of other images. Example ① means that the text features of each sample are only fused with the style features extracted from itself. Example ② means that after adding random noise to the style features of each sample, they are then fused with the text features of itself. Example ③ means that the style features of each sample are fused with the style features of other images. After applying the above three examples to the SOTA model, the experimental results of the accuracy rate (AR) and generalization are shown in Table 1 below.

[0151] Table 1

[0152]

[0153] Referring to Table 1, the text recognition model proposed in this application can outperform the SOTA model in both accuracy and generalization.

[0154] In Example ①: the experiment of fusing text features with its own style features, it can be seen that this method can effectively improve the accuracy and generalization of the model compared with the SOTA model.

[0155] In Example ②: the experiment of fusing text features with the style features with noise, it can be seen that although adding perturbations to the style features will affect the accuracy of the model, it is very helpful for improving generalization.

[0156] The solution proposed in this application combines the advantages of both experiments. In terms of the AR index, it is 0.96% higher than the SOTA model on the ICDAR 2013 dataset, and 36.48%, 11.97%, and 18.14% higher than the SOTA model on the TAL_OCR_CHN, Private Dataset 1, and Private Dataset 2 respectively, becoming the new SOTA effect on these four datasets.

[0157] This application also set up a comparative experiment to verify the effectiveness of the training method of the text recognition model provided in this application, that is, the effectiveness of the text recognition model based on self-supervised style feature extraction. The computing device modified the definition of positive and negative samples for contrastive learning to samples from the same writer as positive samples and samples from different writers as negative samples. The experimental results are shown in Table 2 below.

[0158] Table 2

[0159]

[0160] Among them, the text recognition model corresponding to the self-supervised style feature extraction of image enhancement in Table 2 can refer to Figure 6 the text recognition model in (A) of Figure 6 , and the text recognition model corresponding to the self-supervised writing style feature extraction that decouples the text content can refer to the text recognition model in (B) of

[0161]

[0162] , which will not be elaborated here.

[0163] Regarding the visualization results of the style features extracted by the above-mentioned handwritten text recognition mode based on self-supervised style feature extraction, this embodiment provides an optional implementation manner, such as Figure 8A shown in Figure 8A a schematic diagram of the style feature visualization result provided by this application. The color represents the writer (e.g., black - writer 1, white - writer 2), and the points of the same color represent the samples from the same writer. It can be seen from the figure that most of the samples of the same writer are clustered together, indicating that this model can effectively extract the writing style features of the writer. These style features can be effectively used in other downstream tasks, such as handwriting comparison, handwritten image generation, etc.

[0164] It should be noted that the text recognition model and the style branch network provided in this application can also be applied to tasks such as handwriting comparison and image generation in the field of handwritten text, or to technical fields such as speech synthesis and painting style transfer in other fields that require the introduction of style information, which will not be elaborated here.

[0165] For example, when the solution provided in this application is applied to a product, it can be applied to optical character recognition (OCR) recognition services, such as providing services and selling in the forms of web pages, application programming interfaces (APIs), and mobile applications.

[0166] Regarding the text recognition process provided in this application, this embodiment provides a feasible specific example, as Figure 8B shown Figure 8B is a schematic flow chart of a text recognition method provided in this application Figure 5 . The text recognition method provided in this embodiment includes the following steps ① to ③.

[0167] Step ①: The text recognition model extracts general features from the input image.

[0168] Combined Figure 2 for explanation, the input image may include: a document image and two sub-images obtained by dividing the document image.

[0169] Step ②-1: The text recognition model extracts text features from the general feature 1 corresponding to the document image in the input image to obtain the text features of the input image.

[0170] Step ②-2: The text recognition model extracts style features from the general feature 2 corresponding to the two sub-images in the input image to obtain the style features of the input image.

[0171] Step ③: The text recognition model fuses the text features of the input image and the style features of the input image, so as to enhance diverse style features in the text features, and uses a CTC decoder to parse the fused features to obtain the final prediction result (text recognition result).

[0172] In summary, this embodiment designs a network structure for feature decoupling and re-fusion of handwritten Chinese character images, which is the first model in the OCR field to adopt this architecture. The structural schematic diagram is as Figure 8BAs shown, first, the input image is passed through a backbone network to extract the general features of the image. Subsequently, the general features are input into a text feature extraction module and a style feature extraction module to obtain text features and style features respectively. Then, these two parts of features are fused with the aim of enhancing diverse style features in the text features. Finally, a CTC decoder is used to obtain the final prediction result.

[0173] It can be understood that, in order to implement the functions in the above embodiments, the computing device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application scenario and design constraints of the technical solution.

[0174] As described above in combination with Figures 1 to 8B , the text recognition method and training method provided according to this embodiment are described in detail. Next, in combination with Figure 9 and Figure 10 , the text recognition device and training device provided according to this embodiment will be described.

[0175] Figure 9 FIG. is a schematic structural diagram of a text recognition device provided by the present application. The text recognition device 900 can be used to implement the functions of the computing device in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, the text recognition device 900 can be a computing device as shown in Figure 1 , or the computing device provided in subsequent embodiments. It should be understood that the text recognition device 900 can also be a module (such as a chip) applied to any of the foregoing computing devices.

[0176] As shown in Figure 9 , the text recognition device 900 includes: an image acquisition module 910, a model processing module 920, and a text recognition module 930.

[0177] Exemplarily, the image acquisition module 910 is used to acquire a document image containing text content and divide the document image into multiple sub-images. The model processing module 920 is used to input the document image and the multiple sub-images into a text recognition model and output the text features and style features of the document image. Among them, the text features indicate the first text recognition result of the text content. The text recognition model includes a style branch network, and the style branch network is used to acquire the style features of the document image, and the style features indicate the writing style of the text content. The text recognition module 930 is used to obtain the second text recognition result of the document image according to the text features and style features.

[0178] Optionally, the text recognition model further includes a text branch network, which includes one or a combination of the following: a convolutional layer, a fully connected layer, a recurrent neural network layer, and a deep neural network layer based on a self-attention mechanism. The style branch network includes one or a combination of the following: a convolutional layer, a fully connected layer, and a deep neural network layer based on a self-attention mechanism. The model processing module 920 is specifically configured to: input the document image into the text branch network to obtain the text features of the document image. Input multiple sub-images into the style branch network to obtain the style features of the document image.

[0179] Optionally, the model processing module 920 is specifically configured to: for each of the multiple sub-images, input the sub-image into the style branch network to obtain the style features of the sub-image. Fuse the multiple style features of the multiple sub-images to obtain the style features of the document image.

[0180] Optionally, the text features of the document image include multiple text features. The text recognition module 930 is specifically configured to: for each of the multiple text features, fuse the text feature and the style feature to obtain a fused feature. Analyze the multiple fused features of the multiple text features to obtain a second text recognition result of the document image.

[0181] Optionally, the text recognition device 900 further includes: a confidence level acquisition module and an output module. The confidence level acquisition module is configured to: acquire a first confidence level of the first text recognition result and a second confidence level of the second text recognition result. The output module is configured to: determine the larger confidence level between the first confidence level and the second confidence level, and output the text recognition result corresponding to the larger confidence level.

[0182] Optionally, the text recognition device 900 further includes: an image enhancement module and a model training module. The image enhancement module is configured to: perform image enhancement processing on the document image to obtain an enhanced image. The model processing module 920 is further configured to: input the enhanced image into the style branch network to obtain the style features of the enhanced image. The model training module is configured to: train the style branch network according to the style features of the document image and the style features of the enhanced image to obtain a trained style branch network.

[0183] Optionally, the text content includes handwritten Chinese characters.

[0184] Optionally, there is no overlapping area among the multiple sub-images.

[0185] Optionally, the multiple sub-images may have the same or different sizes.

[0186] The image acquisition module 910, the model processing module 920, the text recognition module 930, and other possible modules can cooperate to implement each step in the above method embodiments. For a more detailed description of the above image acquisition module 910, model processing module 920, and text recognition module 930, reference can be directly made to the relevant description of the computing device in the method embodiments shown in the foregoing drawings, and details are not repeated here.

[0187] Figure 10 FIG. is a schematic structural diagram of a training device provided by the present application. The training device 1000 can be used to implement the functions of the computing device in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. In this embodiment, the training device 1000 can be a Figure 1 computing device or a training device as shown, or a computing device or a training device provided in subsequent embodiments. It should be understood that the training device 1000 can also be a module (such as a chip) applied to any of the foregoing devices.

[0188] Exemplarily, the training device 1000 includes: an image acquisition module 1010, a model processing module 1020, and a model training module 1030. The image acquisition module 1010 is configured to: acquire a document image containing text content, and perform image enhancement processing on the document image to obtain an enhanced image. The model processing module 1020 is configured to: input the document image and the enhanced image into a text recognition model, and output the text features and style features of the document image, and the style features of the enhanced image. Among them, the text features indicate the text recognition result of the text content. The text recognition model includes a style branch network, and the style branch network is used to acquire the style features of the image, and the style features indicate the writing manner of the text content in the image. The model training module 1030 is configured to: train the style branch network according to the style features of the document image and the style features of the enhanced image to obtain a trained text recognition model.

[0189] Optionally, the image acquisition module 1010 is further configured to: acquire an image to be recognized, and divide the image to be recognized into a plurality of sub-images. The model processing module 1020 is further configured to: input the image to be recognized and the plurality of sub-images into the trained text recognition model, and output the text features and style features of the image to be recognized. The text features of the image to be recognized indicate the first text recognition result of the text content in the image to be recognized, and the style features of the image to be recognized indicate: the writing manner of the text content in the image to be recognized. The training device 1000 provided in this embodiment further includes: a text recognition module, configured to obtain a second text recognition result of the image to be recognized according to the text features and style features of the image to be recognized.

[0190] The image acquisition module 1010, the model processing module 1020, the model training module 1030, and other possible modules can cooperate to implement each step in the above method embodiments. For a more detailed description of the above image acquisition module 1010, model processing module 1020, and model training module 1030, reference can be directly made to the relevant descriptions of the computing device and the training device in the method embodiments shown in the foregoing drawings, and details are not repeated here.

[0191] When the text recognition device implements the text recognition method shown in any of the foregoing drawings through software, or the training device implements the training method shown in any of the foregoing drawings through software, the text recognition device and the training device and their respective units can also be software modules. The above text recognition method or training method is implemented by the processor calling the software module. The processor can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The above PLD can be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0192] It can be understood that Figure 9 and Figure 10 The devices shown are only examples provided in this embodiment. According to the differences in the text recognition process and the training process, the text recognition device or the training device may include more or fewer units, which are not limited in this application.

[0193] When the text recognition device or the training device is implemented by hardware, the hardware can be implemented by a processor, a chip, or a chip system. The chip system includes one or more chips, and each chip includes an interface circuit and a control circuit. The interface circuit is used to receive data from other devices outside the chip and transmit it to the control circuit, or send the data from the control circuit to other devices outside the chip. The control circuit and the interface circuit are used to implement the method in any possible implementation manner in the above embodiments through logic circuits or by executing code instructions. The beneficial effects can be referred to the description of any aspect in the above embodiments, and details are not repeated here.

[0194] It can be understood that the processor in the embodiments of the present application may be a CPU, or may also be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0195] In addition, Figure 9 and Figure 10 the device shown can also be implemented by an electronic device. For example, Figure 11 as shown, Figure 11 is a schematic structural diagram of the electronic device provided by the present application. The electronic device 700 includes: a memory 710 and at least one processor 720. The processor 720 can implement the text recognition method or training method provided in the above embodiments, and the memory 710 is used to store software instructions corresponding to the above text recognition method or training method.

[0196] As an alternative implementation, in terms of hardware implementation, the electronic device 700 may refer to a chip or a chip system encapsulating one or more processors 720. For example, when the electronic device 700 is used to implement the method steps in the above embodiments, the processor 720 included in the electronic device 700 executes the steps of the computing device and its possible sub-steps in the above method. In an alternative scenario, the electronic device 700 may further include a communication interface 730, and the communication interface 730 can be used to send and receive data. For example, the communication interface 730 is used to receive image data or send video streams, etc.; the communication interface 730 can be implemented through the interface circuit included in the electronic device 700. Therefore, in some examples, the communication interface 730 can also be referred to as the transceiver of the electronic device. In this embodiment, the communication interface 730 supports wired connection using the unified multimedia interconnection interface.

[0197] In an embodiment of the present application, the communication interface 730, the processor 720, and the memory 710 can be connected through a bus 740, and the bus 740 can be divided into an address bus, a data bus, a control bus, etc. The bus 740 can be a Peripheral Component Interconnect Express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a Cache-Coherent Interconnect for Accelerators (CCIX), or other types of buses, etc.

[0198] It should be noted that the electronic device 700 can also execute Figure 9 or Figure 10 the functions of the device shown, which will not be elaborated here.

[0199] The electronic device 700 provided in this embodiment can be the above training device or computing device, etc., or other devices with an optical character recognition function. The present application does not limit this.

[0200] The method steps in the embodiments of the present application can also be implemented by the processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a network device or a terminal device. Of course, the processor and the storage medium can also exist as discrete components in the training device or the computing device.

[0201] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disc (DVD); or it may be a semiconductor medium, such as a solid state drive (SSD).

[0202] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A text recognition method, characterized in that, the method includes: Obtain a document image containing text content, and divide the document image to obtain multiple sub-images; Input the document image and the multiple sub-images into a text recognition model, and output the text features and style features of the document image; Wherein, the text features indicate the first text recognition result of the text content; the text recognition model includes a style branch network, and the style branch network is used to obtain the style features of the document image, and the style features indicate the writing style of the text content; According to the text features and the style features, obtain the second text recognition result of the document image.

2. The method according to claim 1, characterized in that, The text recognition model further includes a text branch network, and the text branch network includes one or a combination of the following: a convolutional layer, a fully connected layer, a recurrent neural network layer, and a deep neural network layer based on a self-attention mechanism; the style branch network includes one or a combination of the following: a convolutional layer, a fully connected layer, and a deep neural network layer based on a self-attention mechanism; Inputting the document image and the multiple sub-images into a text recognition model, and outputting the text features and style features of the document image, includes: Input the document image into the text branch network to obtain the text features of the document image; Input the multiple sub-images into the style branch network to obtain the style features of the document image.

3. The method according to claim 2, characterized in that, Inputting the multiple sub-images into the style branch network to obtain the style features of the document image, includes: For each sub-image among the multiple sub-images, input the sub-image into the style branch network to obtain the style feature of the sub-image; Fuse the multiple style features of the multiple sub-images to obtain the style features of the document image.

4. The method according to any one of claims 1-3, characterized in that, The text features of the document image include multiple text features; According to the text features and the style features, obtaining the second text recognition result of the document image, includes: For each text feature among the multiple text features, fuse the text feature and the style feature to obtain a fused feature; Parse the multiple fused features of the multiple text features to obtain the second text recognition result of the document image.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtain the first confidence of the first text recognition result and the second confidence of the second text recognition result; Determine the larger confidence among the first confidence and the second confidence, and output the text recognition result corresponding to the larger confidence.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Perform image enhancement processing on the document image to obtain an enhanced image; Input the enhanced image into the style branch network to obtain the style features of the enhanced image; Training the style branch network according to the style features of the document image and the style features of the enhanced image to obtain a trained style branch network.

7. The method according to any one of claims 1-6, wherein, the text content includes handwritten Chinese characters.

8. The method according to any one of claims 1-7, wherein, there are no overlapping regions among the multiple sub-images.

9. The method according to any one of claims 1-8, wherein, the multiple sub-images are of the same or different sizes.

10. A method for training a text recognition model, wherein, the method includes: Obtaining a document image containing text content, and performing image enhancement processing on the document image to obtain an enhanced image; Inputting the document image and the enhanced image into a text recognition model, and outputting the text features and style features of the document image, and the style features of the enhanced image; wherein, the text features indicate the text recognition result of the text content; the text recognition model includes a style branch network, and the style branch network is used to obtain the style features of the image, and the style features indicate the writing method of the text content in the image; Training the style branch network according to the style features of the document image and the style features of the enhanced image to obtain a trained text recognition model.

11. The method according to claim 10, wherein, the method further includes: Obtaining an image to be recognized, and dividing the image to be recognized into multiple sub-images; Inputting the image to be recognized and the multiple sub-images into the trained text recognition model, and outputting the text features and style features of the image to be recognized; the text features of the image to be recognized indicate the first text recognition result of the text content in the image to be recognized, and the style features of the image to be recognized indicate: the writing method of the text content in the image to be recognized; Obtaining a second text recognition result of the image to be recognized according to the text features and style features of the image to be recognized.

12. A text recognition device, wherein, the device includes: An image acquisition module, configured to acquire a document image containing text content, and divide the document image to obtain multiple sub-images; A model processing module, configured to input the document image and the multiple sub-images into a text recognition model, and output the text features and style features of the document image; wherein, the text features indicate the first text recognition result of the text content; the text recognition model includes a style branch network, and the style branch network is used to obtain the style features of the document image, and the style features indicate the writing method of the text content; A text recognition module, configured to obtain a second text recognition result of the document image according to the text features and the style features.

13. The device according to claim 12, wherein, The text recognition model further includes a text branch network, and the text branch network includes one or more combinations of the following: a convolutional layer, a fully connected layer, a recurrent neural network layer, and a deep neural network layer based on a self-attention mechanism; the style branch network includes one or more combinations of the following: a convolutional layer, a fully connected layer, and a deep neural network layer based on a self-attention mechanism; The model processing module is specifically configured to: input the document image into the text branch network to obtain the text features of the document image; input the multiple sub-images into the style branch network to obtain the style features of the document image.

14. The apparatus according to claim 13, wherein, The model processing module is specifically configured to: for each of the multiple sub-images, input the sub-image into the style branch network to obtain the style feature of the sub-image; fuse the multiple style features of the multiple sub-images to obtain the style feature of the document image.

15. The apparatus according to any one of claims 12-14, wherein, The text features of the document image include multiple text features; The text recognition module is specifically configured to: for each of the multiple text features, fuse the text feature and the style feature to obtain a fused feature; analyze the multiple fused features of the multiple text features to obtain a second text recognition result of the document image.

16. The apparatus according to any one of claims 12-15, wherein, The apparatus further includes: A confidence level acquisition module, configured to: acquire a first confidence level of the first text recognition result and a second confidence level of the second text recognition result; An output module, configured to: determine the larger confidence level among the first confidence level and the second confidence level, and output the text recognition result corresponding to the larger confidence level.

17. The apparatus according to any one of claims 12-16, wherein, The apparatus further includes: An image enhancement module, configured to: perform image enhancement processing on the document image to obtain an enhanced image; The model processing module is further configured to: input the enhanced image into the style branch network to obtain the style feature of the enhanced image; A model training module, configured to: train the style branch network according to the style feature of the document image and the style feature of the enhanced image to obtain a trained style branch network.

18. The apparatus according to any one of claims 12-17, wherein, The text content includes handwritten Chinese characters.

19. The apparatus according to any one of claims 12-18, wherein, There are no overlapping regions among the multiple sub-images.

20. The apparatus according to any one of claims 12-19, wherein, The multiple sub-images are of the same or different sizes.

21. A training apparatus for a text recognition model, wherein, The training apparatus includes: An image acquisition module, configured to: acquire a document image containing text content, and perform image enhancement processing on the document image to obtain an enhanced image; A model processing module, configured to: input the document image and the enhanced image into an optical character recognition (OCR) model, and output the text features and style features of the document image, and the style features of the enhanced image; wherein, the text features indicate the OCR results of the text content; the OCR model includes a style branch network, and the style branch network is configured to obtain the style features of an image, and the style features indicate the writing style of the text content in the image. A model training module, configured to: train the style branch network according to the style features of the document image and the style features of the enhanced image, and obtain a trained OCR model.

22. The training device according to claim 21, wherein, the image acquisition module is further configured to: acquire an image to be recognized, and divide the image to be recognized into a plurality of sub-images; the model processing module is further configured to: input the image to be recognized and the plurality of sub-images into the trained OCR model, and output the text features and style features of the image to be recognized; the text features of the image to be recognized indicate the first OCR result of the text content in the image to be recognized, and the style features of the image to be recognized indicate: the writing style of the text content in the image to be recognized; the training device further includes: an OCR module, configured to obtain a second OCR result of the image to be recognized according to the text features and style features of the image to be recognized.

23. A chip, wherein, it includes: a control circuit and an interface circuit; the interface circuit is configured to acquire a document image, and cooperate with the control circuit to execute the method according to any one of claims 1-11.

24. An electronic device, wherein, it includes: a processor and a memory; the memory is configured to store a set of computer instructions, and when the processor executes the set of computer instructions, the method according to any one of claims 1-11 is executed.

25. A computer-readable storage medium, wherein, the computer-readable storage medium includes computer instructions; when the computer instructions run on a computing device, the computing device executes the method according to any one of claims 1-11.

26. A computer program product, wherein, when the computer program product runs on a computing device, the computing device executes the method according to any one of claims 1-11.

Citation Information

Cited By

  • Handwritten Chinese character recognition and error correction method based on artificial intelligence

    CN120954019A