Text recognition method and apparatus, and electronic device and readable storage medium

By performing text branch processing and semantic processing on the original image, the misrecognition problem caused by overlapping or too close Chinese characters in multi-line handwritten text recognition is solved, and a higher recognition accuracy is achieved.

WO2025111923A1PCT designated stage expired Publication Date: 2025-06-05BOE TECHNOLOGY GROUP CO LTD +1

Patent Information

Application Number
PCT/CN2023/135398
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When the prior art recognizes multiple lines of handwritten text in an image, it is easy to overlap text or be too close to misrecognition, which reduces the recognition accuracy.

Method used

By processing the original image, a text branch image is obtained, and each text branch image is recognized and processed in sequence, the first text feature vector is obtained, and then the feature vector is semantic processed to obtain the second text feature vector, and finally the target prediction text is obtained based on these two feature vectors.

Benefits of technology

It effectively avoids mutual interference between two adjacent lines of text, improves the accuracy of text content recognition, and eliminates the problem of misidentification through semantic processing, improving the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023135398_05062025_PF_FP_ABST
    Figure CN2023135398_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a text recognition method and apparatus, and an electronic device and a readable storage medium. The method comprises: processing an original image, so as to obtain at least one text wrapping image, wherein the text wrapping image only includes one line of text content; sequentially performing recognition processing on each text wrapping image, so as to obtain a first text feature vector; performing semantic processing on the first text feature vector, so as to obtain a second text feature vector; and acquiring target predicted text on the basis of the second text feature vector and the first text feature vector. The present embodiment can eliminate the problem of recognition as a single character during image recognition due to misrecognition, the left and right structures not being compact or non-left and right structures being compact, and can achieve an effect of visual and semantic complementation, thereby facilitating an increase in the accuracy of text content recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Text recognition method, device, electronic device and readable storage medium Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a text recognition method, device, electronic device, and readable storage medium. Background Art

[0002] Related technologies typically use visual recognition models to identify text content in images. When the image to be recognized contains multiple lines of handwritten text, the characters of the adjacent existing text may overlap or be close to each other, resulting in misrecognition and reducing recognition accuracy.

[0003] Summary of the Invention

[0004] The present disclosure provides a text recognition method, apparatus, electronic device, and readable storage medium to address the deficiencies of related technologies.

[0005] According to a first aspect of an embodiment of the present disclosure, a text recognition method is provided, the method comprising:

[0006] Processing the original image to obtain at least one text line image; the text line image only contains one line of text content;

[0007] Perform recognition processing on each text line image in turn to obtain a first text feature vector;

[0008] Performing semantic processing on the first text feature vector to obtain a second text feature vector;

[0009] A target predicted text is obtained according to the second text feature vector and the first text feature vector.

[0010] Optionally, the original image is processed to obtain at least one text line image, including:

[0011] Preprocessing the original image to obtain a first image; the first image includes at least one line of text content, and the font size and line width of each character in the at least one line of text content are the same;

[0012] The at least one line of text content is processed into lines to obtain at least one text line image; the text line image only contains one line of text content.

[0013] Optionally, preprocessing the original image to obtain the first image includes:

[0014] determining an image resolution of the first image;

[0015] The position of each character in the first image is adjusted according to the image resolution, and the font size and line width of each character are adjusted to obtain a first image containing the same font size and line width.

[0016] Optionally, determining the image resolution of the first image includes:

[0017] Obtaining the length of each stroke in the text content in the original image, and using the maximum stroke length as the character height of a single character in the text content;

[0018] Obtaining the text height of the text content in the original image;

[0019] Acquire the number of text lines of the text content in the original image according to the text height and the character height;

[0020] The image height of the first image is acquired according to the number of text lines, and the image resolution corresponding to the image height is obtained as the image resolution of the first image.

[0021] Optionally, adjusting the position of each character in the first image and adjusting the font size and line width of each character according to the image resolution to obtain a first image having the same font size and line width includes:

[0022] Obtaining a zoom factor of the first image according to an image height of the first image and a height of the text;

[0023] Determining coordinate data of updated trajectory points according to the zoom factor and coordinate data of original stroke trajectory points;

[0024] The updated trajectory points of each stroke are connected with a preset line width to obtain a first image containing the same font size and line width.

[0025] Optionally, performing line-by-line processing on the at least one line of text content to obtain at least one text line-by-line image includes:

[0026] Inputting the first image into a single-word detection model to obtain coordinate data of a detection frame of each word in the first image; the coordinate data includes the horizontal coordinate and the vertical coordinate of the center point of the detection frame and the width and height of the detection frame;

[0027] Connecting the center points of each detection frame and its left adjacent detection frame in the first image to obtain a center point connection line;

[0028] Obtaining the angle between the center point line and the baseline, and determining that the detection frames whose angle is less than or equal to a preset angle threshold belong to the same text line;

[0029] Merge the text in the detection box of the same text line to obtain the line-by-line result;

[0030] Track point mapping is performed according to the line division result to obtain at least one text line division image.

[0031] Optionally, merge the text within the detection box of the same text line to obtain a line-by-line result, including:

[0032] sorting the detection frames of all single characters in the first image in ascending order according to the size of the horizontal coordinates of the center points, and storing them in a candidate text pool;

[0033] Calculate the angles between the center points of the detection boxes in the candidate text pool and the last detection box in each row and the baseline in sequence;

[0034] When the angle is less than or equal to a preset angle threshold, the single character corresponding to the angle is moved into the current candidate text row;

[0035] In response to traversing the detection frames in the candidate text pool and the candidate text pool is not empty, saving the single characters of the current candidate text line and clearing the current candidate text line, and sequentially calculating the angles between the lines connecting the center points of each detection frame in the candidate text pool and the last detection frame in each line and the baseline; in response to the candidate text pool being empty, obtaining the line branching result.

[0036] Optionally, each text line image is sequentially recognized to obtain a first text feature vector, including:

[0037] Get the image recognition model;

[0038] Each text line image is sequentially input into the image recognition model to obtain a first text feature vector of each text line image.

[0039] Optionally, the image recognition model includes an image encoder and an image decoder;

[0040] The image encoder is used to encode each text line image to obtain an image encoding vector;

[0041] The image decoder is used to decode the image coding vector to obtain a first text feature vector corresponding to the image coding vector, and the first text feature vector is passed through a linear normalization layer to obtain the first predicted text.

[0042] Optionally, performing semantic processing on the first text feature vector to obtain a second text feature vector includes:

[0043] Get the preset semantic model;

[0044] The first text feature vector is input into a preset semantic model to obtain a second text feature vector output by the preset semantic model; the second text feature vector represents the editing label of each word in the first text feature vector.

[0045] Optionally, the preset semantic model includes a semantic encoder and a semantic decoder;

[0046] The semantic encoder is used to obtain an embedding vector corresponding to the first text feature vector, obtain a position encoding vector corresponding to the character position in the embedding vector, and calculate self-attention on a synthetic vector synthesized from the embedding vector and the position encoding vector to obtain a semantic encoding vector;

[0047] The semantic decoder is used to determine a second text feature vector corresponding to the first text feature vector according to the semantic encoding vector.

[0048] Optionally, the semantic encoder includes at least one network unit connected in series; each network unit includes: a position encoding module, a feature synthesis module and a semantic encoding module; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the semantic encoding module; the semantic encoding module is connected to the semantic decoder;

[0049] The position encoding module is used to perform position encoding processing on the first text feature vector to obtain a position encoding vector;

[0050] The feature synthesis module is used to synthesize the embedding vector and the position encoding vector to obtain the feature synthesis vector;

[0051] The semantic encoding module is used to calculate self-attention on the feature synthesis vector to obtain the semantic encoding vector.

[0052] Optionally, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the semantic encoding module uses a Transformer network to implement it.

[0053] Optionally, the semantic decoder comprises at least one semantic network unit connected in series; the semantic network unit comprises a text position encoding module, a self-attention module, a semantic decoding module, a feedforward network module and a linear normalization module;

[0054] The text position encoding module is used to perform position encoding using a sinusoidal position encoding method. Its input is the previous predicted text, and its output is a position encoding vector.

[0055] The self-attention module is used to calculate the attention of the position encoding vector to obtain the attention vector;

[0056] The semantic encoding module is used to perform semantic encoding processing on the attention vector and the semantic encoding vector to obtain a weighted semantic encoding vector;

[0057] The feedforward network unit is used to perform nonlinear transformation processing on the weighted semantic coding vector to obtain a second text feature vector;

[0058] The linear normalization module is used to perform linear transformation and normalization processing on the second text feature vector to obtain a semantically predicted text, namely the first error-corrected text.

[0059] Optionally, the semantic encoding module includes:

[0060] The scaling point unit is used to calculate the attention weight of the semantic encoding vector and the attention vector to obtain the dot product matrix;

[0061] The activation layer unit is used to perform weight normalization on the dot product matrix to obtain the attention weight;

[0062] The matrix multiplication unit is used to perform matrix multiplication of the attention weight and the semantic encoding vector to obtain a weighted semantic encoding vector.

[0063] Optionally, the image recognition model and the preset semantic model are trained by the following steps, including:

[0064] Acquire a training sample set, wherein the training sample set includes a plurality of text line images;

[0065] Inputting each text line image into an image recognition model to obtain a first predicted text and a first text feature vector output by the image recognition model; obtaining the first predicted text after the first text feature vector passes through a linear normalization layer;

[0066] Inputting the first text feature vector into a preset semantic model to obtain a second text feature vector and a first error correction text output by the preset semantic model; the first error correction text is a predicted text content obtained by performing linear transformation and normalization processing on the second text feature vector;

[0067] Calculating the cross entropy of the first predicted text, the target predicted text, the first error-corrected text, and the annotated labels of the text line images, respectively, to obtain a first cross entropy, a second cross entropy, and a third cross entropy;

[0068] Calculate a current loss value based on the first cross entropy, the second cross entropy, and the third cross entropy;

[0069] In response to the current loss value being less than or equal to a preset loss threshold, it is determined that the image recognition model and the preset semantic model have completed training.

[0070] Optionally, obtaining a target predicted text according to the second text feature vector and the first text feature vector includes:

[0071] Inputting the first text feature vector and the second text feature vector into a connection layer to obtain a temporary text feature vector output by the connection layer;

[0072] The temporary text feature vector is input into a normalization layer, and the predicted text output by the normalization layer is obtained as the target predicted text.

[0073] According to a second aspect of an embodiment of the present disclosure, a text recognition device is provided, the device comprising:

[0074] A text line image acquisition module is used to process the original image to obtain at least one text line image; the text line image only contains one line of text content;

[0075] A first text vector acquisition module is used to sequentially perform recognition processing on each text line image to obtain a first text feature vector;

[0076] A second text vector acquisition module, configured to perform semantic processing on the first text feature vector to obtain a second text feature vector;

[0077] The target predicted text acquisition module is configured to acquire the target predicted text according to the second text feature vector and the first text feature vector.

[0078] Optionally, the text line image acquisition module includes:

[0079] A first image acquisition submodule is configured to preprocess the original image to obtain a first image; the first image includes at least one line of text content, and the font size and line width of each character in the at least one line of text content are the same;

[0080] The line-by-line image acquisition submodule is used to perform line-by-line processing on the at least one line of text content to obtain at least one text line-by-line image; the text line-by-line image only contains one line of text content.

[0081] Optionally, the first image acquisition submodule includes:

[0082] a resolution determining unit, configured to determine an image resolution of the first image;

[0083] The first image acquisition unit is used to adjust the position of each character in the first image and adjust the font size and line width of each character according to the image resolution to obtain a first image with the same font size and line width.

[0084] Optionally, the resolution determination unit includes:

[0085] a character height acquisition subunit, configured to acquire the length of each stroke in the text content in the original image, and use the maximum stroke length as the character height of a single character in the text content;

[0086] A text height acquisition subunit, used to acquire the text height of the text content in the original image;

[0087] A text line number acquisition subunit, configured to acquire the number of text lines of the text content in the original image according to the text height and the character height;

[0088] The resolution acquisition subunit is configured to acquire the image height of the first image according to the number of text lines, and obtain the image resolution corresponding to the image height as the image resolution of the first image.

[0089] Optionally, the first image acquisition unit includes:

[0090] a zoom factor obtaining subunit, configured to obtain a zoom factor of the first image according to an image height of the first image and a height of the text;

[0091] a trajectory coordinate determination subunit, configured to determine coordinate data of updated trajectory points according to the zoom factor and the coordinate data of the original stroke trajectory points;

[0092] The first image determination subunit is used to connect the updated trajectory points of each stroke with a preset line width to obtain a first image containing the same font size and line width.

[0093] Optionally, the line-by-line image acquisition submodule includes:

[0094] a detection frame coordinate acquisition unit, configured to input a first image of the at least one text line image into a single-word detection model to obtain coordinate data of a detection frame corresponding to each word in the first image; the coordinate data including the horizontal and vertical coordinates of the center point of the detection frame and the width and height of the detection frame;

[0095] a center line obtaining unit, configured to connect the center points of each detection frame in the first image and its adjacent detection frame on the left to obtain a center point line;

[0096] a text line determination unit, configured to obtain an angle between the center point line and the baseline, and determine that the detection frames whose angle is less than or equal to a preset angle threshold belong to the same text line;

[0097] A line branch result acquisition unit, used to merge the text within the detection box of the same text line to obtain a line branch result;

[0098] The line branch image determining unit is used to perform track point mapping according to the line branch result to obtain at least one text line branch image.

[0099] Optionally, the branch result obtaining unit includes:

[0100] a coordinate sorting subunit, configured to sort the detection frames of all single characters in the first image in ascending order according to the size of the horizontal coordinates of the center points, and store the candidate text pool;

[0101] An angle calculation subunit, configured to sequentially calculate the angles between the line connecting the center points of each detection box in the candidate text pool and the last detection box in each row and the baseline;

[0102] a text line calculation subunit, configured to move a single character corresponding to the angle into a current candidate text line when the angle is less than or equal to a preset angle threshold;

[0103] The line branch result determination subunit is configured to, in response to traversing the detection frames in the candidate text pool and the candidate text pool being not empty, save the individual characters of the current candidate text line and clear the current candidate text line, and sequentially calculate the angles between the line connecting the center points of each detection frame in the candidate text pool and the last detection frame in each line and the baseline; and in response to the candidate text pool being empty, obtain the line branch result.

[0104] Optionally, the first text vector acquisition module includes:

[0105] Recognition model acquisition submodule, used to obtain image recognition model;

[0106] The first text acquisition submodule is used to input each text line image into the image recognition model in sequence to obtain a first text feature vector of each text line image.

[0107] Optionally, the image recognition model includes an image encoder and an image decoder;

[0108] The image encoder is used to encode each text line image to obtain an image encoding vector;

[0109] The image decoder is used to decode the image coding vector to obtain a first text feature vector corresponding to the image coding vector, and the first text feature vector is passed through a linear normalization layer to obtain the first predicted text.

[0110] Optionally, the second text vector acquisition module includes:

[0111] A preset semantic model acquisition submodule is used to acquire a preset semantic model;

[0112] The second text vector acquisition module is configured to input the first text feature vector into a preset semantic model to obtain a second text feature vector output by the preset semantic model.

[0113] Optionally, the preset semantic model includes a semantic encoder and a semantic decoder;

[0114] The semantic encoder is used to obtain an embedding vector corresponding to the first text feature vector, obtain a position encoding vector corresponding to the character position in the embedding vector, and calculate self-attention on a synthetic vector synthesized from the embedding vector and the position encoding vector to obtain a semantic encoding vector;

[0115] The semantic decoder is used to determine a second text feature vector corresponding to the first text feature vector according to the semantic encoding vector.

[0116] Optionally, the semantic encoder includes at least one network unit connected in series; each network unit includes: a position encoding module, a feature synthesis module and a semantic encoding module; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the semantic encoding module; the semantic encoding module is connected to the semantic decoder;

[0117] The position encoding module is used to perform position encoding processing on the first text feature vector to obtain a position encoding vector;

[0118] The feature synthesis module is used to synthesize the embedding vector and the position encoding vector to obtain the feature synthesis vector;

[0119] The semantic encoding module is used to calculate self-attention on the feature synthesis vector to obtain the semantic encoding vector.

[0120] Optionally, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the semantic encoding module uses a Transformer network to implement it.

[0121] Optionally, the semantic decoder comprises at least one semantic network unit connected in series; the semantic network unit comprises a text position encoding module, a self-attention module, a semantic decoding module, a feedforward network module and a linear normalization module;

[0122] The text position encoding module is used to perform position encoding using a sinusoidal position encoding method. Its input is the previous predicted text, and its output is a position encoding vector.

[0123] The self-attention module is used to calculate the attention of the position encoding vector to obtain the attention vector;

[0124] The semantic encoding module is used to perform semantic encoding processing on the attention vector and the semantic encoding vector to obtain a weighted semantic encoding vector;

[0125] The feedforward network unit is used to perform nonlinear transformation processing on the weighted semantic coding vector to obtain a second text feature vector;

[0126] The linear normalization module is used to perform linear transformation and normalization processing on the second text feature vector to obtain a semantically predicted text, namely the first error-corrected text.

[0127] Optionally, the semantic encoding module includes:

[0128] The scaling point unit is used to calculate the attention weight of the semantic encoding vector and the attention vector to obtain the dot product matrix;

[0129] The activation layer unit is used to perform weight normalization on the dot product matrix to obtain the attention weight;

[0130] The matrix multiplication unit is used to perform matrix multiplication of the attention weight and the semantic encoding vector to obtain a weighted semantic encoding vector.

[0131] Optionally, the image recognition model and the preset semantic model are trained by the following steps, including:

[0132] Acquire a training sample set, wherein the training sample set includes a plurality of text line images;

[0133] Inputting each text line image into an image recognition model to obtain a first predicted text and a first text feature vector output by the image recognition model; obtaining the first predicted text after the first text feature vector passes through a linear normalization layer;

[0134] Inputting the first text feature vector into a preset semantic model to obtain a second text feature vector and a first error correction text output by the preset semantic model; the first error correction text is a predicted text content obtained by performing linear transformation and normalization processing on the second text feature vector;

[0135] Calculating the cross entropy of the first predicted text, the target predicted text, the first error-corrected text, and the annotated labels of the text line images, respectively, to obtain a first cross entropy, a second cross entropy, and a third cross entropy;

[0136] Calculate a current loss value based on the first cross entropy, the second cross entropy, and the third cross entropy;

[0137] In response to the current loss value being less than or equal to a preset loss threshold, it is determined that the image recognition model and the preset semantic model have completed training.

[0138] Optionally, the target predicted text acquisition module includes:

[0139] a temporary vector acquisition submodule, configured to input the first text feature vector and the second text feature vector into a connection layer to obtain a temporary text feature vector output by the connection layer;

[0140] The second preset text acquisition submodule is used to input the temporary text feature vector into the normalization layer to obtain the predicted text output by the normalization layer as the target predicted text.

[0141] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, comprising

[0142] a processor; a memory for storing a computer program executable by the processor;

[0143] The processor is configured to execute the computer program in the memory to implement the method as described in any one of the first aspects.

[0144] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which, when an executable computer program in the storage medium is executed by a processor, can implement the method described in any one of the first aspects.

[0145] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0146] As can be seen from the above embodiments, the solution provided by the embodiment of the present disclosure processes the original image to obtain at least one text line image; the text line image contains only one line of text content; then, each text line image is sequentially recognized and processed to obtain a first text feature vector; thereafter, the first text feature vector is semantically processed to obtain a second text feature vector; finally, the target predicted text is obtained based on the second text feature vector and the first text feature vector. In this way, in this embodiment, the first predicted text is obtained by performing text recognition using the text line image, which can avoid mutual interference between two adjacent lines of text and improve the accuracy of text content recognition; and, by performing semantic processing on the first text feature vector, the problem of being recognized as a single word due to misidentification, non-compact left and right structures, or non-compact left and right structures during image recognition can be eliminated, thereby achieving a visual and semantic complementary effect, which is conducive to improving the accuracy of text content recognition.

[0147] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0148] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0149] Fig. 1 is a flowchart showing a text recognition method according to an exemplary embodiment.

[0150] Fig. 2 is a flow chart showing a method of acquiring a first image according to an exemplary embodiment.

[0151] Fig. 3 is a schematic diagram showing a single line of text according to an exemplary embodiment.

[0152] Fig. 4 is a schematic diagram showing a multi-line text according to an exemplary embodiment.

[0153] Fig. 5 is a schematic diagram showing coordinates of mapped trajectory points according to an exemplary embodiment.

[0154] Fig. 6 is a flow chart showing a method of acquiring at least one text line image according to an exemplary embodiment.

[0155] Fig. 7 is a schematic diagram showing a text line image channel transformation according to an exemplary embodiment.

[0156] Fig. 8 is a schematic diagram showing filling and re-recognition of a text line image according to an exemplary embodiment.

[0157] Fig. 9 is a flowchart showing another method of acquiring at least one text line image according to an exemplary embodiment.

[0158] FIG10 is a schematic diagram showing single-word detection, text line segmentation, and line segmentation results according to an exemplary embodiment.

[0159] Fig. 11 is a flowchart showing a method of obtaining a first predicted text according to an exemplary embodiment.

[0160] FIG12 is a schematic structural diagram of an image recognition model according to an exemplary embodiment.

[0161] Fig. 13 is a schematic structural diagram of an encoder of an image recognition model according to an exemplary embodiment.

[0162] FIG14 is a schematic diagram of the structure of a multi-head attention unit according to an exemplary embodiment.

[0163] Fig. 15 is a schematic diagram showing visualization of a region of interest in an original image according to an exemplary embodiment.

[0164] Fig. 16 is a block diagram of a decoder of an image recognition model according to an exemplary embodiment.

[0165] Fig. 17 is a block diagram showing an image recognition model according to an exemplary embodiment.

[0166] FIG18 is a schematic diagram showing a training architecture of an image recognition model according to an exemplary embodiment.

[0167] Fig. 19 is a flow chart showing another text recognition method according to an exemplary embodiment.

[0168] FIG20 is a schematic structural diagram of a preset semantic model according to an exemplary embodiment.

[0169] FIG21 is a schematic structural diagram of an encoder of a preset semantic model according to an exemplary embodiment.

[0170] FIG22 is a schematic diagram showing a training architecture of an image recognition model and a semantic recognition model according to an exemplary embodiment.

[0171] Fig. 23 is a block diagram of a text recognition device according to an exemplary embodiment. DETAILED DESCRIPTION

[0172] Exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements, unless otherwise indicated. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims. It should be noted that, unless there is a conflict, the features of the following embodiments and implementations may be combined with each other.

[0173] To solve the above technical problems, the embodiments of the present disclosure provide a text recognition method, apparatus, electronic device, and readable storage medium. FIG1 is a flowchart of a text recognition method according to an exemplary embodiment. Referring to FIG1 , a text recognition method includes steps 11 to 14:

[0174] In step 11, the original image is processed to obtain at least one text line image; the text line image only contains one line of text content.

[0175] In this step, the electronic device can pre-process the original image to obtain a first image; the first image contains at least one line of text content and the font size and line width of each character in the at least one line of text content are the same; then, the at least one line of text content is divided into lines to obtain at least one text line-divided image; the text line-divided image only contains one line of text content.

[0176] In this step, the electronic device may pre-process the original image to obtain a first image, see FIG2 , step 21 and step 22 .

[0177] In step 21 , the image resolution of the first image is determined.

[0178] In one example, the electronic device can obtain the length of each stroke in the text content in the original image, for example, stroke_len = max(xmax-xmin+1,ymax-ymin+1), where xmin is the minimum x-coordinate value of the current stroke, xmax is the maximum x-coordinate value of the current stroke, ymin is the minimum y-coordinate value of the current stroke, and ymax is the maximum y-coordinate value of the current stroke.

[0179] Then, the electronic device may use the maximum value of the stroke length as the character height (max(stroke_len)) of a single character in the text content.

[0180] Afterwards, the electronic device can obtain the text height (trace_h) of the text content in the original image, where trace_h=Ymax-Ymin+1, where Ymin is the minimum y-coordinate value of all strokes and Ymax is the maximum y-coordinate value of all strokes.

[0181] Afterwards, the electronic device may obtain the number of text lines (raw_num) of the text content in the original image according to the text height and the character height, where raw_num=trace_h / stroke_len.

[0182] Finally, the electronic device can obtain the image height (input_h) of the first image according to the number of text lines, where the image height of the first image input_h=raw_num×80. In this embodiment, the image height of the first image is used as the resolution of the first image.

[0183] In one example, referring to FIG3 , it can be obtained that the number of rows of the original image is 1.18 and the image height is 94 pixels; referring to FIG4 , it can be obtained that the number of rows of the original image is 3.2 and the image height is 256 pixels.

[0184] In step 22, the position and line width of each character in the first image are adjusted according to the image resolution to obtain a first image containing the same font size and line width.

[0185] In one example, the electronic device may obtain a zoom factor of the first image based on the image height and text height of the first image, where zoom factor ratio = input_h / trace_h.

[0186] Then, the electronic device can determine the coordinate data of the updated trajectory point based on the zoom factor and the coordinate data of the original stroke trajectory point. For example, point_X = (point_x-xmin) × ratio, point_Y = (point_y-ymin) × ratio, where point_x and point_y represent the x-coordinate and y-coordinate of the original trajectory point before mapping, xmin and ymin represent the minimum x-coordinate and minimum y-coordinate of all original trajectory points, respectively, and point_Y and point_Y represent the x-coordinate and y-coordinate of the updated trajectory point after mapping, respectively. Referring to Figure 5, the left figure shows the position of the original trajectory point, which can be mapped to the position of the updated trajectory point in the middle figure.

[0187] It should be noted that the methods for obtaining the original stroke trajectory points in the original image may include: (1) trajectory points collected by the touch acquisition device when the user writes the text content; (2) offline trajectory points attached when the original image is uploaded, which are obtained by using method (1) on other devices; (3) sampling after detecting a single word after acquiring the original image. Those skilled in the art can select stroke trajectory points according to the specific scenario, and this is not limited here.

[0188] Finally, the electronic device can connect the updated trajectory points of each stroke with a preset line width to obtain a first image containing the same font size and line width. For example, taking the stroke as a unit, all the trajectory points in each stroke are connected in sequence with a line with a line width of 2 pixels (the trajectory points are updated after mapping) to obtain the first image. Continuing to refer to Figure 5, after connecting the strokes, the text shown in the right figure can be obtained. In this way, this embodiment can ensure that each text in different original images has the same font size through mapping transformation processing; and can ensure that each text has the same line width through line width connection processing, thereby ensuring that the handwriting of the text written by different users is the same, which is conducive to eliminating the influence of different handwriting on the recognition accuracy.

[0189] It should be noted that after determining the updated trajectory points for each stroke, the electronic device can also fit the updated trajectory points for each stroke and determine that the line width of the fitted stroke is 2 pixels, thereby also obtaining each character. Technicians can select a fitting method based on the specific scenario, such as linear interpolation or quadratic fitting. If it can obtain each stroke, the corresponding solution falls within the scope of protection of this disclosure.

[0190] In this step, the electronic device may perform line-by-line processing on the at least one line of text content to obtain at least one text line-by-line image; the text line-by-line image only contains one line of text content.

[0191] In this step, the electronic device may perform line-by-line processing on the at least one line of text content to obtain at least one text line-by-line image; the text line-by-line image contains only one line of text content, see FIG. 6 , including steps 61 to 65 .

[0192] In step 61, the electronic device can input the first image into a single-word detection model to obtain coordinate data of the detection box of each word in the first image; the coordinate data includes the horizontal and vertical coordinates of the center point of the detection box and the width and height of the detection box.

[0193] The single word detection model can be pre-stored in the electronic device. In one example, the single word detection model is implemented using the YOLOv5 network model. The YOLOv5 network model includes a backbone network (Backbone), a neck part (Neck) and a prediction part (Prediction). Among them, the backbone network includes a slice structure (Focus), a convolution module (Conv), a bottleneck layer (C3) and a spatial pyramid pooling (SPP), which are used to extract the features of the input first image; the neck part uses a feature-like pyramid structure to fuse high-level features and low-level features for enhanced feature representation; the prediction part adopts a multi-scale prediction form, so that the network is suitable for target detection of different scales and has stronger generalization ability.

[0194] The slicing structure performs a slicing operation on the first input. See Figure 7. The left image is the first image. Every other pixel in the first image is extracted to obtain four images on the right. These four images are complementary and no information is lost. This converts the width W and height H information into channel space, expanding the input channels by a factor of four. This means the spliced ​​image has 12 channels compared to the original RGB primary color image. Finally, the new image on the right undergoes a convolution operation to obtain a doubly downsampled feature map with no information loss. For example, if the original 640×640×3 image is input to the Focus structure, the slicing operation first converts it into a 320×320×12 feature map. Then, after another convolution operation, the feature map becomes a 320×320×32 feature map.

[0195] In this example, single-word detection on the first image includes: centering the first image and padding it up, down, left, and right, with the padding pixel values ​​all being (255, 255, 255), to obtain a 2560*2560 text image to be recognized, see Figure 8, and padding the left image as the middle image; then, adjusting the image resolution to 960*960 pixels by scaling the length and width proportionally, as shown in the right image in Figure 8; finally, sending it to the YOLOv5 network model for single-word detection, and the YOLOv5 network model outputs the center point x-coordinate, center point y-coordinate, detection frame width, and detection frame height of each single-word detection frame, that is, obtaining the coordinate data of the detection frame of a single word.

[0196] In step 62, a line is drawn between the center points of each detection frame in the first image and the center points of the adjacent detection frame on its left.

[0197] In this step, the electronic device can sort the detection frames of all single characters in the first image in ascending order according to the size of the horizontal coordinate of the center point, and store the candidate text pool; traverse all the detection frames in the candidate text pool in turn, and obtain the connection line between each single character detection frame and the center point of the last detection frame in each row.

[0198] In step 63, the angle between the center point line and the baseline is obtained, and it is determined that the detection frames whose angle is less than or equal to a preset angle threshold belong to the same text line.

[0199] In this step, the electronic device can obtain the angle between the center point line and the reference line (such as the horizontal line), and when the above angle is less than or equal to a preset angle threshold (such as 5 to 15 degrees), determine that the single-word detection box belongs to the same text line.

[0200] In step 64, the text within the detection box of the same text line is merged to obtain a line separation result.

[0201] In step 65, trajectory point mapping is performed based on the line division results to obtain at least one text line division image. For example, the trajectory points can be divided into lines based on the line division results. A cropping frame is then formed based on the maximum and minimum values ​​of the vertical coordinates of the trajectory points. The first image is cropped based on this cropping frame to obtain a horizontal strip of the first image, i.e., a text line division image. This cropping operation is repeated to obtain a text line division image corresponding to each line of text.

[0202] As shown in Figure 9, the touchscreen device of an electronic device can collect trajectory point data (x, y, isLeave) during the user's writing process, calculate the number of text lines raw_num, and then map the trajectory points into a first image with a text line width of 2 pixels and a height of h = raw_num * 70. The first image is then fed into the YOLOv5-960 network for single-word detection to obtain single-word detection boxes. The single-word detection boxes are then merged to obtain text line results. Finally, the first image is cropped by line to obtain a text line image.

[0203] The text line separation results are obtained by merging single-word detection boxes, including:

[0204] Initialize an empty row text library raw_dict={} to store the line division results, where the key value key of the row text library raw_dict={} is the text row identification code id (i.e., the text row number), and the value value is the single-word detection box (identification code) contained in each row, in the format of "text row number, coordinates of all single-word detection boxes contained in the row". Then, the electronic device can arrange all single-word detection boxes in ascending order according to the x-coordinate of the center point, for example, raw_dict{0}=[boxes[0]] (meaning that the first single-word detection box after sorting (all single-word detection boxes are arranged in ascending order according to the x-coordinate value of the box center point) is placed in the first row and recorded in the row text library), and traverse all single-word detection boxes after the first one. Afterwards, the electronic device can calculate the angle between the center point of the current single-word detection box and the line connecting the center point of the last box in each row in raw_dict (i.e., the line connecting the center points of the last detection boxes of each detection box in the current candidate text row) and the horizontal line, and store the angle in the horizontal angle library angles. Get the minimum value min(angles) in the horizontal angle library angles, and determine whether the above minimum value is less than or equal to the preset angle threshold (such as 10 degrees), that is, min(angles)<10°. When the above minimum value is less than or equal to the preset angle threshold, if the current single-word detection box id is angles.index(min(angles)) (it means that the current single-word detection box belongs to the text line where the id of the single-word detection box is angles.index(min(angles))': the current single-word detection box belongs to the single-word detection box with the smallest angle with it), add the current single-word detection box to raw_dict={}, that is, add it to the current candidate text line. When the above minimum value is greater than the preset angle threshold, the current single-word detection box id is a new text line, a new key is created, and it is stored in raw_dict={}, that is, the next candidate text line is added. Afterwards, determine whether the traversal of the single-word detection box is completed? When the traversal is not completed, jump to the next candidate text line and continue traversing the single-word detection frame for all detection frames that have not been branched (that is, other detection frames outside the current candidate text line) until the traversal is completed and the branch result is input.

[0205] Referring to Figure 10, the top figure shows the detection of individual characters with a detection box placed outside each character; the middle figure shows the effect of detecting and separating three lines of text; and the bottom figure shows the effect of using different colors to represent each line of text (the first and third lines are lighter in color, while the second line is darker to distinguish them). It can be seen that this embodiment can accurately separate multiple lines of text even when the text is tilted, the line spacing is small, or there is overlap, which helps improve recognition accuracy.

[0206] In step 12, each text line image is sequentially recognized to obtain a first text feature vector.

[0207] In this step, the electronic device may perform recognition processing on each text line image in sequence to obtain a first text feature vector, as shown in FIG11 , which includes step 111 and step 112 .

[0208] In step 111, an image recognition model is obtained.

[0209] In this step, the electronic device can obtain an image recognition model, which is pre-trained and stored in a designated location. The above-mentioned designated location may include but is not limited to local memory, cache or cloud, etc., which will not be repeated here.

[0210] In this step, referring to Figure 12 , the image recognition model includes an encoder 121 and a decoder 122. The image encoder 121 is used to encode each text line image to obtain an image encoding vector; the image decoder 122 is used to decode the image encoding vector to obtain the first predicted text corresponding to the image encoding vector.

[0211] In one example, referring to FIG13 , the encoder 121 may include a first image encoding module 131, a position encoding module 132, a feature synthesis module 133, and a second image encoding module 134. The first image encoding module 131 is connected to the position encoding module 132 and the feature synthesis module 133, respectively; the position encoding module 132 is connected to the feature synthesis module 133; the feature synthesis module 133 is connected to the second image encoding module 134; and the second image encoding module 134 is connected to the decoder 122.

[0212] The first image encoding module 131 is used to encode the text single line image to obtain the image feature vector corresponding to the text single line image;

[0213] The position coding module 132 is used to perform position coding processing on the pixel positions in the image feature vector to obtain a position coding vector;

[0214] The feature synthesis module 133 is used to synthesize the image feature vector and the position coding vector to obtain a feature synthesis vector;

[0215] The second image encoding module 134 is used to perform feature encoding on the feature synthesis vector image to enhance the pixels in the region where the text is located in the image feature vector to obtain an image encoding vector.

[0216] In one embodiment, the first image encoding module 131 in the encoder 121 is implemented using a DenseNet network, such as a DenseNet98 network. Referring to Table 1, the DenseNet98 network sequentially includes a convolution layer, a pooling layer, a dense block 1, a transition layer 1, a dense block 2, a transition layer 2, a dense block 3, and a linear transformation layer. Taking dense block 1 as an example, in the DenseNet98 network, dense block 1 includes 16 series units, each unit including 2 operations, namely, one convolution operation with a convolution kernel of 1*1 and a step size of 1 and one convolution operation with a convolution kernel of 3*3 and a step size of 1. The 16 series units are executed 16 times, thereby increasing the number of channels of the input data. The structures and input and output data of the other network layers of the DenseNet98 network can be found in Table 1 and will not be repeated here.

[0217] Table 1 DenseNet98 network structure

[0218] In one embodiment, the position encoding module 132 in the encoder 121 implements position encoding using a sinusoidal position encoding method, as shown in equations (1) and (2).

[0219] In formula (1) and formula (2), pos represents the position of the pixel to be encoded in the text single line image (when position encoding is performed on a two-dimensional image, pos_x and pos_y are calculated and then CatConcat is connected), i represents the position encoding dimension index, and d model Indicates the position encoding dimension (d is used for position encoding of a two-dimensional text single-line image). model Equal to half the number of channels of the image to be encoded, for example d model =128).

[0220] In this embodiment, when encoding text, each character is represented by a 256-dimensional vector through a word embedding layer (embedding layer), where the word embedding layer is used to construct a dictionary mapping table and construct a 256-dimensional vector for each character; then, position encoding is performed according to the above sinusoidal position encoding text, dmodel = 256, and a sinusoidal position encoding value is obtained, that is, the position encoding vector in subsequent embodiments.

[0221] In one embodiment, the second image encoding module 134 in the encoder 121 is implemented using a Transformer network. Continuing with FIG13 , the second image encoding module 134 includes a multi-head attention unit 1341, a first residual normalization unit 1342, a feedforward network unit 1343, and a second residual normalization unit 1344. The input and output of the multi-head attention unit 1341 are respectively connected to the first input and second input of the first residual normalization unit 1342. The output of the first residual normalization unit 1342 is respectively connected to the input of the feedforward network unit 1343 and the first input of the second residual normalization unit 1344. The output of the feedforward network unit 1343 is connected to the second input of the second residual normalization unit 1344.

[0222] The multi-head attention unit 1341 is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector;

[0223] The first residual normalization unit 1342 is used to add the feature synthesis vector and the weighted image feature vector and then perform normalization processing to obtain a normalized feature vector;

[0224] The feedforward network unit 1343 is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector;

[0225] The second residual normalization unit 1344 is used to calculate the residual data between the initial image coding vector and the normalized feature vector and perform normalization processing to obtain the image coding vector.

[0226] Referring to Figure 14, the multi-head attention unit 1341 includes a scale dot-product unit (Scale dot-product) 141, a first activation layer (softmax) 142, an attention refinement module (Attention Refinement Module, ARM) 143, a second activation layer (softmax) 144 and a matrix multiplication (Matmul) unit 145.

[0227] The scaled dot product unit 41 is used to calculate the attention weights for the predicted text (Q) and the image encoding vector (K) to obtain a dot product matrix E. The attention weights represent the correlation between each element in the predicted text Q and each element in K.

[0228] The first activation layer 142 is used to perform weight normalization processing on the dot product matrix E to obtain the attention weight A; wherein the attention weight A represents the correlation between the decoded single word sequence and the image encoding vector (or image feature map);

[0229] The attention refining unit 143 is used to obtain the corresponding area of ​​the decoded text in the image coding vector to obtain a refined matrix R; and obtain a difference matrix ER between the dot product matrix E and the refined matrix R; the difference matrix ER represents the unresolved area in the image feature vector (or image feature map), and the above-mentioned unresolved area refers to the corresponding area of ​​the undecoded text in the image coding vector (for example, the actual written content in the original image is 'x+y=1', and after three decodings, the decoded text is 'x+y', and the undecoded text is '=1'), that is, the decoder pays more attention to the unresolved area in the feature map;

[0230] The second activation layer 144 is used to normalize the difference matrix ER to obtain an adjustment matrix (ER)';

[0231] The matrix multiplication unit 145 is used to perform matrix multiplication on the adjustment matrix and the image feature vector (V) to obtain a weighted image feature vector. In the weighted image feature vector, a higher weight is assigned to the region that has not been decoded so that the decoder pays more attention to it.

[0232] In one example, the structure of the multi-head attention unit 1341 shown in FIG14 can be expressed using formula (3):

[0233] In formula (3), Q, K and V represent query, key and value respectively; K and V have the same value in the multi-head attention unit 1341; d k Indicates the dimension of K.

[0234] Among them, in Q, K and V in the multi-head attention unit 1341 shown in Figure 14, Q is the feature synthesis vector output by the feature synthesis module 133, and K and V are the image coding vectors after adding position coding, which are used to determine the resolved areas and unresolved areas in the image feature map, and increase the attention map weight of the resolved area.

[0235] In one embodiment, referring to FIG14 , the specific operation process of the attention refining unit 143 is shown in Formula (4), Formula (5) and Formula (6).

[0236] In equations (4) to (6), T = len(query); L = len(key) = h_0 × w_0, where h_0 and w_0 are the height and width of the encoded feature map output by DenseNet98, respectively; h is the multi-head attention head value (h = 16); and the time step t∈[0,T). C is accumulated by A. By accumulating the attention weights to detect resolved regions, the decoder can also focus on unresolved regions to avoid repeated recognition. It is the result of reshaping C. The dimension of C is T*L*h, where L=h_0*w_0, and the dimension is T*h_0*w_0*h. The convolutional layer generally acts on a one-dimensional plane of height*width, so the L in the C dimension must be expanded to h_0*w_0 through reshaping.

[0237] Referring to FIG15 , the attention refining unit 143 visualizes the results of the refined matrix during the recognition process. The resolved regions in each subgraph are darker in color, indicating that ARM will suppress the attention weights of these resolved regions, thereby encouraging the model to focus on unresolved regions.

[0238] In one embodiment, the feedforward network unit 1343 can be implemented using a fully connected neural network, such as a linear layer + a ReLU layer + a linear layer, as shown in equation (7). FFN(x) = max(0, xW1+b1)W2+b2; (7)

[0239] In formula (7), x represents the initial weight data output by the first residual normalization unit 1342; xW1+b1 respectively represent the weight values ​​adjusted by the first linear layer; max(0,xW1+b1) represents selecting the larger value from 0 and xW1+b1; max(0,xW1+b1)W2+b2 represents the weight value adjusted by the second linear layer.

[0240] In one embodiment, the first residual normalization unit 1342 and the second residual normalization unit 1344 use residual connections so that the output of each layer is added to the input to prevent gradient vanishing and explosion. The expressions of the first residual normalization unit 1342 and the second residual normalization unit 1344 are shown in Equations (8) and (9), respectively. output = layer_norm(Multi-Head Attention(input)+input); (8) output 2 = layer_norm(FFN(output)+output); (9)

[0241] In formulas (8), (9) and (10), Multi-Head Attention (input) represents the initial text vector output by the multi-head attention unit 1341, input represents the feature synthesis vector, layer_norm() represents the normalization process, output represents the initial weight data output by the first residual normalization unit 1342, outpu2 represents the image encoding vector output by the second residual normalization unit 1344; FFN() represents the operation as in formula (7); x is the input feature, μ and σ 2 are the mean and variance of x, respectively, ε=10 -6 .

[0242] In one embodiment, referring to FIG16 , the decoder 122 includes a text position encoding module 161 , a self-attention module 162 , a multi-head attention map module 163 , a feedforward network module 164 , and a linear normalization module 165 .

[0243] The text position coding module 161 can use sinusoidal position coding to implement position coding. Its input is the previous predicted text. Its value in the first prediction is the starting symbol. <sos>, then add each predicted word based on the starting symbol, for example, when predicting for the first time, the value of the previous predicted text is <sos>, the predicted result output is "Jing"; when making the second prediction, the value of the previous predicted text is <sos>Beijing, the prediction result output is "East"; the value of the previous prediction text in the third prediction is <sos>JD.com, and so on, until the prediction result output is the end character. <eos>End the prediction. Please refer to formula (1) and formula (2) for details, which will not be repeated here.

[0244] The self-attention module 162 can be implemented using a multi-head attention method, specifically referring to the structure shown in FIG14 , which is used to calculate the correlation between words in the predicted text (i.e., to obtain the semantic information contained in the predicted text sequence), and its expression is shown in formula (11).

[0245] In formula (11), Q, K and V are the predicted texts after single-word position encoding.

[0246] The implementation scheme of the multi-head attention module 163 is the same as the structure shown in Figure 14 and will not be repeated here.

[0247] The implementation scheme of the feedforward network module 164 is shown in formula (7), which will not be described here in detail.

[0248] Based on the above analysis, an example structure of the image recognition model in this embodiment is shown in FIG17 .

[0249] In one example, the dimension of the input data of the image recognition model, that is, the original image, is 80*w*1, where w represents the image width.

[0250] During the encoding stage, the single-line text image is encoded by the image encoding module 131 to obtain an image feature vector of 5*w / 16*256 dimensions; the position encoding module 132 performs position encoding on the image feature vector to obtain a position encoding vector of 5*w / 16*256 dimensions; the feature synthesis module 133 outputs a feature synthesis vector of l*256 dimensions, where l=5*w / 16, l represents the width of the feature synthesis vector after transformation; the second image encoding module 134 encodes the feature synthesis vector twice to obtain an image coding vector of l*256 dimensions.

[0251] During the decoding phase, the text position encoding module 161 in the encoder 122 can position-encode the previous predicted text of length (the length of the predicted text vector) * 1 dimension into a text sequence of length * 256 dimensions; the self-attention module 162 can output text of length * 256 dimensions. The input data of the multi-head attention module 163 also includes the image encoding vector of l * 256 dimensions input by the encoder 121, and its output is text of length * 256 dimensions; the input and output of the feedforward network module 164 are both text of length * 256 dimensions; the linear normalization module 165 transforms the text of length * 256 dimensions into a text vector of length * k dimensions, where k is the number of single words supported by the image recognition model, and the last single word is taken as the recognition result of this time. After multiple recognition processes, the predicted text can be obtained.

[0252] Considering the diversity of handwritten text styles, it is necessary to improve the robustness of the image recognition model. In this embodiment, the idea of ​​adversarial learning is used to train the image recognition model shown in Figure 17. The training framework is shown in Figure 18. Referring to Figure 18, two image recognition models with identical structures are used. The output data of the first image encoding module 131 is passed through a linear unit (linear) to calculate the consistency loss value; the predicted text vectors of the two image recognition models are respectively compared with the annotated labels of the training sample images (subsequent handwritten text images and printed text images) to calculate the cross-entropy loss value; finally, the total loss value of this training is calculated based on the consistency loss value and the two cross-entropy loss values ​​to determine whether the image recognition model has completed training.

[0253] It should be noted that the purpose of passing the output data of the first image encoding module 131 through the linear unit is to reduce the number of channels of the output data, thereby reducing the amount of calculation when calculating the consistency loss value, which is conducive to speeding up the training speed.

[0254] The steps for training the image recognition model in this embodiment include:

[0255] In this step, the electronic device can obtain a training sample data set. In one example, a handwritten text image and its annotated label can be obtained; the annotated label is used to represent the text in the handwritten text image. Then, a printed text image is generated based on the annotated label, and the printed text image and the handwritten text image constitute an image pair; the printed text image and the handwritten text image have the same annotated label. For example, a white image with the same resolution as the handwritten text image and a pixel value of 255 is initialized as the background image; then, the Pillow toolkit in Python is called to write the text content of the annotated label in the white background, and the text font size is set to be consistent with the text font size in the handwritten image to obtain a printed text image.

[0256] The electronic device can then input the handwritten text images in each pair of images into the previous image recognition model shown in FIG18 (hereinafter referred to as the first image recognition model), obtaining a first image feature vector and a first predicted text string output by the first image recognition model; the first image feature vector and the first predicted text string serve as the output data of the first image recognition model; and the printed text images in each pair of images can be input into the next image recognition model shown in FIG18 (hereinafter referred to as the second image recognition model), with the second image feature vector and the target predicted text string output by the second image recognition model serving as the output data of the second image recognition model. Finally, the total loss value of the image recognition model pair is obtained based on the output data of each image recognition model.

[0257] In this step, the electronic device can calculate the consistency loss value of the first image recognition model and the second image recognition model based on the first image feature vector and the second image feature vector, as shown in formula (12).

[0258] In formula (12), L CL represents the consistency loss value, F h and F p Represent the first image feature vector and the second image feature vector respectively, σ represents the ReLU activation function, F h '=σ(Wσ(F h )) and F p '=σ(Wσ(F p )) represent the normalized feature maps of the first image feature vector and the second image feature vector respectively.

[0259] It is understandable that the above consistency loss value can constrain the consistency of features recognized by the handwritten text model and the printed text model, so that the features extracted by the encoder are less affected by the written font as possible, that is, the written font noise is eliminated and the essential features of the text are directly extracted.

[0260] Then, the electronic device can obtain a first cross-entropy loss value of the first image recognition model based on the first predicted text string and the annotated label of the handwritten text image, and obtain a second cross-entropy loss value of the second image recognition model based on the target predicted text string and the annotated label of the printed text image, as shown in Formula (13).

[0261] In formula (13), exp represents the exponential function with the natural constant e as the base, exp(1) = e 1 =2.718…, τ represents temperature. In one example, τ=0.05.

[0262] It can be understood that the first cross entropy loss value and the second cross entropy loss value are used to constrain the divergence between handwritten text and printed text, so that the image recognition model has text recognition capabilities.

[0263] Afterwards, the electronic device can obtain the average of the first cross entropy loss value and the second cross entropy loss value to obtain the average cross entropy loss value; then, obtain the product of the consistency loss value and the preset weight value to obtain the weighted consistency loss value; finally, obtain the sum of the weighted consistency loss value and the average cross entropy loss value to obtain the total loss value, as shown in formula (14).

[0264] In formula (14), L total Represents the total loss value, the weight coefficient λ = 0.02, and gt represents the annotation label.

[0265] Finally, the electronic device can determine the relationship between the total loss value and a preset loss value threshold (e.g., 95.0% to 99.9%). When the total loss value is greater than the preset loss value threshold, it can be determined that the image recognition model has not completed training; when the total loss value is less than or equal to the preset loss value threshold, it can be determined that the image recognition model has completed training. At this time, the first image recognition model in the image recognition model pair can be regarded as the image recognition model that has completed training and stored in the electronic device.

[0266] In step 112, each text line image is sequentially input into the image recognition model to obtain a first text feature vector of each text line image.

[0267] In this step, the electronic device can sequentially input each text line image into the image recognition model to obtain a first text feature vector corresponding to each text line image. It is understood that after the first text feature vector passes through a linear normalization layer, a first predicted text can be obtained, i.e., the preset text obtained after visual recognition. In this embodiment, using text line images for text recognition to obtain the first predicted text can avoid mutual interference between adjacent lines of text, thereby improving the accuracy of text content recognition.

[0268] In step 13, semantic processing is performed on the first text feature vector to obtain a second text feature vector.

[0269] In this step, the electronic device may perform semantic processing on the first text feature vector to obtain a second text feature vector, as shown in FIG. 19 , which includes steps 191 to 192 .

[0270] In step 191, a preset semantic model is obtained.

[0271] In this step, the electronic device can obtain the first text feature vector output by the image recognition model. Continuing with FIG16 , the first text feature vector refers to the first text feature vector output by the feedforward network module 164. The first text feature vector is processed by a linear normalization module. For example, the linear normalization module may include a Linear layer and a Softmax layer. The Linear layer is used to adjust the weight value of each word in the first text feature vector, and the Softmax layer is used to map each word to a text prediction vector that supports word recognition, thereby obtaining a first predicted text.

[0272] In this step, a preset semantic model is stored in the electronic device. Referring to FIG. 20 , the preset semantic model 200 includes a semantic encoder 201 and a semantic decoder 202. The semantic encoder 201 is configured to obtain an embedding vector corresponding to the first predicted text, obtain positional encoding vectors corresponding to character positions in the embedding vector, and calculate self-attention on a composite vector of the embedding vector and the positional encoding vector to obtain a semantic encoding vector. The semantic decoder 202 is configured to determine a second text feature vector corresponding to the first predicted text based on the semantic encoding vector.

[0273] In one example, referring to FIG21 , the semantic encoder 201 includes at least one network unit connected in series. FIG21 illustrates a scenario in which four network units are connected in series, represented by "4X." Each network unit includes a position encoding module 211, a feature synthesis module 212, and a semantic encoding module 213. The position encoding module 211 is connected to the feature synthesis module 212; the feature synthesis module 213 is connected to the semantic encoding module 213; and the semantic encoding module 213 is connected to the image decoder.

[0274] The position coding module 211 is used to perform position coding processing on the first text feature vector to obtain a position coding vector;

[0275] The feature synthesis module 212 is used to synthesize the embedding vector and the position encoding vector to obtain the feature synthesis vector;

[0276] The semantic encoding module 213 is used to calculate the self-attention of the feature synthesis vector to obtain the semantic encoding vector.

[0277] In one example, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the semantic encoding module uses a Transformer network to implement it.

[0278] It should be noted that the specific structure of the semantic encoder 201 can be the same as the specific structure of the encoder of the image coding model illustrated in FIG13 . For details, please refer to the analysis content of the encoder of the image coding model, which will not be repeated here.

[0279] In one example, the semantic decoder 202 may include at least one decoding network unit connected in series. In one example, the semantic decoder 202 is implemented using one decoding network unit. Continuing with FIG21 , the decoding network unit includes a text position encoding module 221, a self-attention module 222, a semantic decoding module 223, a feedforward network module 224, and a linear normalization module 225.

[0280] The text position coding module 221 is used to perform position coding using a sinusoidal position coding method, with its input being the previous predicted text and its output being a position coding vector;

[0281] The self-attention module 222 is used to calculate the attention of the position encoding vector to obtain an attention vector;

[0282] The semantic encoding module is used to perform semantic encoding processing on the attention vector and the semantic encoding vector, specifically including:

[0283] The scale dot-product unit calculates the attention weights for the semantic encoding vector K and the attention vector Q to obtain the dot product matrix E. The attention weights represent the correlation between each element in the attention vector Q and each element in K. The above elements can be understood as word embedding vectors.

[0284] The activation layer unit (softmax) performs weight normalization on the dot product matrix E to obtain the attention weight A;

[0285] The matrix multiplication unit is used to perform matrix multiplication of the attention weight A and the semantic encoding vector V to obtain a weighted semantic encoding vector. In the weighted semantic feature vector, higher weights are assigned to the word embedding vectors that have not yet been decoded, so that the decoder can focus on them.

[0286] The feedforward network unit 224 is used to perform nonlinear transformation processing on the weighted semantic coding vector to obtain a second text feature vector.

[0287] The linear normalization module 224 is used to perform linear transformation and normalization processing on the linear normalization module 225 to obtain a semantically predicted text, namely, a first error-corrected text.

[0288] In one embodiment, the preset semantic model can be implemented using a semantic encoder 20 and a linear normalization layer, that is, a set of linear layers and softmax layers are added to the output end of the semantic encoder 20, so that the semantic coding vector output by the semantic encoder 20 is used as the second text feature vector, and the second text feature vector is then linearly transformed and normalized into the first error-corrected text through the linear layer and the softmax layer.

[0289] In step 192, the first text feature vector is input into a preset semantic model to obtain a second text feature vector output by the preset semantic model; the second text feature vector represents the editing label of each word in the first text feature vector.

[0290] In this step, the electronic device can input the first predicted text into a preset semantic model, and the preset semantic model can output a second text feature vector; the second text feature vector represents the editing labels of each single character in the first text feature vector. The above editing labels are used to represent what kind of editing operations are performed on each character in the first preset text corresponding to the first text feature vector.

[0291] In one example, the editing labels can include but are not limited to keep (keep the current character), delete (delete the current character), keep|x (insert the target character x after the current character), delete|x (replace the current character with the target character x), etc., and the target character x can be any single character supported by the preset semantic model. Suppose the first predicted text output by the image recognition model is "For 3 to reward Shenying侍者's irrigation grace", and the decoder of the preset semantic model outputs the second text feature vector as "keep delete|了keep keep keep keep keep keep keep keep keep delete|之keep".

[0292] It can be understood that by adding a linear normalization unit based on the output of the preset semantic model, the second text feature vector can be converted into a preset text. For example, the first corrected text is "For the sake of rewarding Shenying侍者's irrigation grace". It should be noted that the first corrected text is only used in subsequent training of the preset semantic model, and only the conversion relationship between the second text feature vector and the first corrected text is described here.

[0293] The disclosed embodiment also provides a method for training an image recognition model and a preset semantic model, and its training structure is shown in Figure 22. Referring to Figure 22, the dimension of the text line image is h*w*1, and the dimension of the first text feature vector output by the image recognition model is length*256. Wherein, length represents the length of the first predicted text (equal to the length of the first error correction text and the length of the second preset text). The first text feature vector is linearly normalized to obtain a predicted vector of length*k, and then the maximum value of each of length k vectors is taken to obtain a dimension of length*1. Wherein k refers to the number of text categories supported by the network for recognition. The preset semantic model outputs a second text feature vector of dimension length*256, which is then transformed into a text vector of dimension lenth*(2*k+2) through a linear normalization layer (Linear&softmax), and the maximum value of each of length (2*k+2) vectors is taken to obtain the first error correction text (dimension length*1). The first text feature vector and the second text feature vector are spliced ​​to obtain a text splicing vector of dimension length*512. The text concatenation vector is then transformed into a text vector of dimension length*k through a linear normalization layer (Linear&softmax). The maximum value of each length k vector is taken to obtain the target predicted text.

[0294] In conjunction with the training structure shown in Figure 22, the training process includes:

[0295] Step 1: Obtain a training sample set, which includes multiple text line images. The method for obtaining the text line images can refer to the contents of steps 11 and 12, and will not be repeated here.

[0296] Step 2: Input each text line image into the image recognition model to obtain the first predicted text and the first text feature vector output by the image recognition model. Please refer to the content of step 13 for details, which will not be repeated here.

[0297] Step 3: Input the first text feature vector into a preset semantic model to obtain a second text feature vector and a first error correction text output by the preset semantic model; the first error correction text is the predicted text content obtained by linear transformation and normalization of the second text feature vector; please refer to the content of step 192 for details and will not be repeated here.

[0298] Step 4: Calculate the cross entropy of the first predicted text, the target predicted text, and the first error-corrected text for the annotated labels of each text line image to obtain a first cross entropy, a second cross entropy, and a third cross entropy. The cross entropy is calculated as shown in formula (13) and will not be repeated here.

[0299] Step 5, calculate the current loss value according to the first cross-entropy, the second cross-entropy, and the third cross-entropy, as shown in Equation (15).

[0300] In Equation (15), L total1 represents the current loss value, and L1, L2, and L3 respectively represent the first cross-entropy, the second cross-entropy, and the third cross-entropy.

[0301] Step 6, determine the magnitude relationship between the current loss value and the preset loss value. When the current loss value is greater than the preset loss threshold, jump to Step 2; when the current loss value is less than or equal to the preset loss threshold, it is determined that the image recognition model and the preset semantic model are completed with training.

[0302] It should be noted that during this training process, the image recognition model and the preset semantic model can be trained synchronously, and the image recognition model and / or the preset semantic model do not need to be trained separately, thereby improving the training efficiency.

[0303] In Step 14, obtain the target prediction text according to the second text feature vector and the first text feature vector.

[0304] In this step, the electronic device can obtain the composite vector of the first text feature vector and the second text feature vector to obtain the text splicing vector. For example, splice each channel of the first text feature vector with each channel of the second text feature vector to increase the length of each channel. Then, perform linear normalization processing on the text splicing vector and map it to the supported single-word vector to obtain the target prediction text. For example, the first prediction text is "To reward the divine attendant Ying for his irrigation grace", and the second text feature vector is "keep delete|with keep keep keep keep keep keep keep keep keep delete|of keep", then the target prediction text is "To reward the divine attendant Ying for his irrigation grace".

[0305] So far, in the solution provided by the embodiments of the present disclosure, semantic recognition processing can be performed on the first prediction text, which can eliminate the problem that a single character is recognized due to misrecognition, non-compact left-right structure, or non-left-right structure compactness during image recognition, achieving the effect of visual and semantic complementarity, and is beneficial to improving the accuracy of text content recognition.

[0306] Based on the text recognition method provided by the embodiments of the present disclosure, this embodiment also provides a text recognition device. Referring to FIG. 23, the device includes:

[0307] A text line image acquisition module 231, configured to process the original image to obtain at least one text line image; the text line image only contains one line of text content;

[0308] A first text vector acquisition module 232 is used to sequentially perform recognition processing on each text line image to obtain a first text feature vector;

[0309] A second text vector acquisition module 233 is configured to perform semantic processing on the first text feature vector to obtain a second text feature vector;

[0310] The target predicted text acquisition module 234 is configured to acquire a target predicted text according to the second text feature vector and the first text feature vector.

[0311] In one embodiment, the text line image acquisition module includes:

[0312] A first image acquisition submodule is configured to preprocess the original image to obtain a first image; the first image includes at least one line of text content, and the font size and line width of each character in the at least one line of text content are the same;

[0313] The line-by-line image acquisition submodule is used to perform line-by-line processing on the at least one line of text content to obtain at least one text line-by-line image; the text line-by-line image only contains one line of text content.

[0314] In one embodiment, the first image acquisition submodule includes:

[0315] a resolution determining unit, configured to determine an image resolution of the first image;

[0316] The first image acquisition unit is used to adjust the position of each character in the first image and adjust the font size and line width of each character according to the image resolution to obtain a first image with the same font size and line width.

[0317] In one embodiment, the resolution determination unit includes:

[0318] a character height acquisition subunit, configured to acquire the length of each stroke in the text content in the original image, and use the maximum stroke length as the character height of a single character in the text content;

[0319] A text height acquisition subunit, used to acquire the text height of the text content in the original image;

[0320] A text line number acquisition subunit, configured to acquire the number of text lines of the text content in the original image according to the text height and the character height;

[0321] The resolution acquisition subunit is configured to acquire the image height of the first image according to the number of text lines, and obtain the image resolution corresponding to the image height as the image resolution of the first image.

[0322] In one embodiment, the first image acquisition unit includes:

[0323] a zoom factor obtaining subunit, configured to obtain a zoom factor of the first image according to an image height of the first image and a height of the text;

[0324] a trajectory coordinate determination subunit, configured to determine coordinate data of updated trajectory points according to the zoom factor and the coordinate data of the original stroke trajectory points;

[0325] The first image determination subunit is used to connect the updated trajectory points of each stroke with a preset line width to obtain a first image containing the same font size and line width.

[0326] In one embodiment, the line-by-line image acquisition submodule includes:

[0327] a detection frame coordinate acquisition unit, configured to input a first image of the at least one text line image into a single-word detection model to obtain coordinate data of a detection frame corresponding to each word in the first image; the coordinate data including the horizontal and vertical coordinates of the center point of the detection frame and the width and height of the detection frame;

[0328] a center line obtaining unit, configured to connect the center points of each detection frame in the first image and its adjacent detection frame on the left to obtain a center point line;

[0329] a text line determination unit, configured to obtain an angle between the center point line and the baseline, and determine that the detection frames whose angle is less than or equal to a preset angle threshold belong to the same text line;

[0330] A line branch result acquisition unit, used to merge the text within the detection box of the same text line to obtain a line branch result;

[0331] The line branch image determining unit is used to perform track point mapping according to the line branch result to obtain at least one text line branch image.

[0332] In one embodiment, the branch result obtaining unit includes:

[0333] a coordinate sorting subunit, configured to sort the detection frames of all single characters in the first image in ascending order according to the size of the horizontal coordinates of the center points, and store the candidate text pool;

[0334] An angle calculation subunit, configured to sequentially calculate the angles between the line connecting the center points of each detection box in the candidate text pool and the last detection box in each row and the baseline;

[0335] a text line calculation subunit, configured to move a single character corresponding to the angle into a current candidate text line when the angle is less than or equal to a preset angle threshold;

[0336] The line branch result determination subunit is configured to, in response to traversing the detection frames in the candidate text pool and the candidate text pool being not empty, save the individual characters of the current candidate text line and clear the current candidate text line, and sequentially calculate the angles between the line connecting the center points of each detection frame in the candidate text pool and the last detection frame in each line and the baseline; and in response to the candidate text pool being empty, obtain the line branch result.

[0337] In one embodiment, the first text vector acquisition module includes:

[0338] Recognition model acquisition submodule, used to obtain image recognition model;

[0339] The first text acquisition submodule is used to input each text line image into the image recognition model in sequence to obtain a first text feature vector of each text line image.

[0340] In one embodiment, the image recognition model includes an image encoder and an image decoder;

[0341] The image encoder is used to encode each text line image to obtain an image encoding vector;

[0342] The image decoder is used to decode the image coding vector to obtain a first text feature vector corresponding to the image coding vector, and the first text feature vector is passed through a linear normalization layer to obtain the first predicted text.

[0343] In one embodiment, the second text vector acquisition module includes:

[0344] A preset semantic model acquisition submodule is used to acquire a preset semantic model;

[0345] The second text vector acquisition module is configured to input the first text feature vector into a preset semantic model to obtain a second text feature vector output by the preset semantic model.

[0346] In one embodiment, the preset semantic model includes a semantic encoder and a semantic decoder;

[0347] The semantic encoder is used to obtain an embedding vector corresponding to the first text feature vector, obtain a position encoding vector corresponding to the character position in the embedding vector, and calculate self-attention on a synthetic vector synthesized from the embedding vector and the position encoding vector to obtain a semantic encoding vector;

[0348] The semantic decoder is used to determine a second text feature vector corresponding to the first text feature vector according to the semantic encoding vector.

[0349] In one embodiment, the semantic encoder includes at least one network unit connected in series; each network unit includes: a position encoding module, a feature synthesis module and a semantic encoding module; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the semantic encoding module; the semantic encoding module is connected to the semantic decoder;

[0350] The position encoding module is used to perform position encoding processing on the first text feature vector to obtain a position encoding vector;

[0351] The feature synthesis module is used to synthesize the embedding vector and the position encoding vector to obtain the feature synthesis vector;

[0352] The semantic encoding module is used to calculate self-attention on the feature synthesis vector to obtain the semantic encoding vector.

[0353] In one embodiment, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the semantic encoding module uses a Transformer network to implement.

[0354] In one embodiment, the semantic decoder includes at least one semantic network unit connected in series; the semantic network unit includes a text position encoding module, a self-attention module, a semantic decoding module, a feedforward network module and a linear normalization module;

[0355] The text position encoding module is used to perform position encoding using a sinusoidal position encoding method. Its input is the previous predicted text, and its output is a position encoding vector.

[0356] The self-attention module is used to calculate the attention of the position encoding vector to obtain the attention vector;

[0357] The semantic encoding module is used to perform semantic encoding processing on the attention vector and the semantic encoding vector to obtain a weighted semantic encoding vector;

[0358] The feedforward network unit is used to perform nonlinear transformation processing on the weighted semantic coding vector to obtain a second text feature vector;

[0359] The linear normalization module is used to perform linear transformation and normalization processing on the second text feature vector to obtain a semantically predicted text, namely the first error-corrected text.

[0360] In one embodiment, the semantic encoding module includes:

[0361] The scaling point unit is used to calculate the attention weight of the semantic encoding vector and the attention vector to obtain the dot product matrix;

[0362] The activation layer unit is used to perform weight normalization on the dot product matrix to obtain the attention weight;

[0363] The matrix multiplication unit is used to perform matrix multiplication of the attention weight and the semantic encoding vector to obtain a weighted semantic encoding vector.

[0364] In one embodiment, the image recognition model and the preset semantic model are trained by the following steps, including:

[0365] Acquire a training sample set, wherein the training sample set includes a plurality of text line images;

[0366] Inputting each text line image into an image recognition model to obtain a first predicted text and a first text feature vector output by the image recognition model; obtaining the first predicted text after the first text feature vector passes through a linear normalization layer;

[0367] Inputting the first text feature vector into a preset semantic model to obtain a second text feature vector and a first error correction text output by the preset semantic model; the first error correction text is a predicted text content obtained by performing linear transformation and normalization processing on the second text feature vector;

[0368] Calculating the cross entropy of the first predicted text, the target predicted text, the first error-corrected text, and the annotated labels of the text line images, respectively, to obtain a first cross entropy, a second cross entropy, and a third cross entropy;

[0369] Calculate a current loss value based on the first cross entropy, the second cross entropy, and the third cross entropy;

[0370] In response to the current loss value being less than or equal to a preset loss threshold, it is determined that the image recognition model and the preset semantic model have completed training.

[0371] In one embodiment, the target predicted text acquisition module includes:

[0372] a temporary vector acquisition submodule, configured to input the first text feature vector and the second text feature vector into a connection layer to obtain a temporary text feature vector output by the connection layer;

[0373] The second preset text acquisition submodule is used to input the temporary text feature vector into the normalization layer to obtain the predicted text output by the normalization layer as the target predicted text.

[0374] It should be noted that the device embodiment shown in this embodiment matches the content of the above-mentioned method embodiment. You can refer to the content of the above-mentioned method embodiment and will not repeat it here.

[0375] In an exemplary embodiment, an electronic device is also provided, comprising

[0376] Display screen;

[0377] processor;

[0378] a memory for storing a computer program executable by the processor;

[0379] The processor is configured to execute the computer program in the memory to implement the above method.

[0380] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including an executable computer program. The executable computer program can be executed by a processor to implement the method of the above embodiment. The computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0381] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0382] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.< / eos> < / sos> < / sos> < / sos> < / sos>

Claims

1. A text recognition method, characterized in that, the method includes: Processing the original image to obtain at least one text line-separated image; the text line-separated image only contains one line of text content; Sequentially performing recognition processing on each text line-separated image to obtain a first text feature vector; Performing semantic processing on the first text feature vector to obtain a second text feature vector; Obtaining a target predicted text according to the second text feature vector and the first text feature vector.

2. The method according to claim 1, characterized in that, Processing the original image to obtain at least one text line-separated image includes: Preprocessing the original image to obtain a first image; the first image contains at least one line of text content and the font sizes and line widths of each character in the at least one line of text content are the same; Performing line separation processing on the at least one line of text content to obtain at least one text line-separated image; the text line-separated image only contains one line of text content.

3. The method according to claim 2, characterized in that, Preprocessing the original image to obtain a first image includes: Determining the image resolution of the first image; Adjusting the positions of each character in the first image and adjusting the font sizes and line widths of each character according to the image resolution to obtain a first image with the same font size and line width.

4. The method according to claim 3, characterized in that, Determining the image resolution of the first image includes: Obtaining the lengths of each stroke in the text content of the original image, and taking the maximum value of the stroke lengths as the character height of a single character in the text content; Obtaining the text height of the text content in the original image; Obtaining the number of text lines of the text content in the original image according to the text height and the character height; Obtaining the image height of the first image according to the number of text lines, and taking the image resolution corresponding to the image height as the image resolution of the first image.

5. The method according to claim 4, characterized in that, Adjusting the positions of each character in the first image and adjusting the font sizes and line widths of each character according to the image resolution to obtain a first image with the same font size and line width includes: Obtaining the scaling factor of the first image according to the image height of the first image and the text height; Determining the coordinate data of the updated trajectory points according to the scaling factor and the coordinate data of the original stroke trajectory points; Connecting the updated trajectory points of each stroke with a preset line width to obtain a first image with the same font size and line width.

6. The method according to claim 2, characterized in that, Performing line separation processing on the at least one line of text content to obtain at least one text line-separated image includes: Inputting the first image into a single-character detection model to obtain the coordinate data of the detection boxes of each character in the first image; the coordinate data includes the abscissa and ordinate of the center point of the detection box and the width and height of the detection box; Connecting the center points of each detection box in the first image with the center point of its adjacent detection box on the left to obtain a center point connection line; Obtain the included angle between the connection line of the center points and the reference line, and determine that the detection frames with the included angle less than or equal to the preset included angle threshold belong to the same text line; Merge the texts within the detection frames of the same text line to obtain the line-breaking result; Perform trajectory point mapping according to the line-breaking result to obtain at least one text line-breaking image.

7. The method according to claim 6, wherein, merging the texts within the detection frames of the same text line to obtain the line-breaking result includes: Sort the detection frames of all single texts in the first image in ascending order according to the abscissa of the center points, and store the candidate text pool; Calculate the included angle between the connection line of the center points of each detection frame in the candidate text pool and the last detection frame in each line and the reference line in turn; When the included angle is less than or equal to the preset included angle threshold, move the single text corresponding to the included angle into the current candidate text line; In response to traversing all the detection frames in the candidate text pool and the candidate text pool is not empty, save the single texts in the current candidate text line and clear the current candidate text line, and calculate the included angle between the connection line of the center points of each detection frame in the candidate text pool and the last detection frame in each line in turn; in response to the candidate text pool being empty, obtain the line-breaking result.

8. The method according to claim 2, wherein, Perform recognition processing on each text line-breaking image in turn to obtain the first text feature vector, including: Obtain an image recognition model; Input each text line-breaking image into the image recognition model in turn to obtain the first text feature vector of each text line-breaking image.

9. The method according to claim 8, wherein, the image recognition model includes an image encoder and an image decoder; the image encoder is used to perform encoding processing on each text line-breaking image to obtain an image encoding vector; the image decoder is used to perform decoding processing on the image encoding vector to obtain the first text feature vector corresponding to the image encoding vector, and the first text feature vector obtains the first predicted text after passing through a linear normalization layer.

10. The method according to any one of claims 1 to 9, wherein, perform semantic processing on the first text feature vector to obtain a second text feature vector, including: Obtain a preset semantic model; Input the first text feature vector into the preset semantic model to obtain the second text feature vector output by the preset semantic model; the second text feature vector represents the editing labels of each single character in the first text feature vector.

11. The method according to claim 10, wherein, the preset semantic model includes a semantic encoder and a semantic decoder; the semantic encoder is used to obtain the embedding vector corresponding to the first text feature vector, obtain the position encoding vector corresponding to the character position in the embedding vector, and calculate the self-attention on the combined vector of the embedding vector and the position encoding vector to obtain a semantic encoding vector; the semantic decoder is used to determine the second text feature vector corresponding to the first text feature vector according to the semantic encoding vector.

12. The method according to claim 11, wherein, The semantic encoder includes at least one network unit connected in series; each network unit includes: a position encoding module, a feature synthesis module, and a semantic encoding module; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the semantic encoding module; the semantic encoding module is connected to the semantic decoder; The position encoding module is configured to perform position encoding processing on the first text feature vector to obtain a position encoding vector; The feature synthesis module is configured to perform synthesis processing on the embedding vector and the position encoding vector to obtain the feature synthesis vector; The semantic encoding module is configured to calculate self-attention on the feature synthesis vector to obtain the semantic encoding vector.

13. According to the method described in claim 12, wherein, The position encoding module implements position encoding using a sine position encoding method, and the semantic encoding module is implemented using a Transformer network.

14. According to the method described in claim 11, wherein, The semantic decoder includes at least one semantic network unit connected in series; the semantic network unit includes a text position encoding module, a self-attention module, a semantic decoding module, a feed-forward network module, and a linear normalization module; The text position encoding module is configured to perform position encoding using a sine position encoding method, and its input is the previous predicted text, and its output is a position encoding vector; The self-attention module is configured to calculate attention on the position encoding vector to obtain an attention vector; The semantic encoding module is configured to perform semantic encoding processing on the attention vector and the semantic encoding vector to obtain a weighted semantic encoding vector; The feed-forward network unit is configured to perform non-linear transformation processing on the weighted semantic encoding vector to obtain a second text feature vector; The linear normalization module is configured to perform linear transformation and normalization processing on the second text feature vector to obtain a semantic prediction text, i.e., a first error-corrected text.

15. According to the method described in claim 14, wherein, The semantic encoding module includes: The scaling dot unit is configured to calculate attention weights on the semantic encoding vector and the attention vector to obtain a dot product matrix; The activation layer unit is configured to perform weight normalization processing on the dot product matrix to obtain attention weights; The matrix multiplication unit is configured to perform matrix multiplication on the attention weights and the semantic encoding vector to obtain a weighted semantic encoding vector.

16. According to the method described in claim 10, wherein, The image recognition model and the preset semantic model are trained through the following steps, including: Obtain a training sample set, where the training sample set includes multiple text line images; Input each text line image into the image recognition model to obtain a first predicted text and a first text feature vector output by the image recognition model; the first text feature vector obtains the first predicted text after passing through a linear normalization layer; Input the first text feature vector into the preset semantic model to obtain a second text feature vector and a first error-corrected text output by the preset semantic model; the first error-corrected text is the predicted text content obtained by performing linear transformation processing and normalization processing on the second text feature vector; Calculate the cross-entropies of the first predicted text, the target predicted text, and the first error-corrected text with the annotation labels of each text line image respectively, to obtain a first cross-entropy, a second cross-entropy, and a third cross-entropy; Calculate a current loss value according to the first cross-entropy, the second cross-entropy, and the third cross-entropy; In response to the current loss value being less than or equal to a preset loss threshold, determine that the image recognition model and the preset semantic model are completed with training.

17. According to the method of claim 10, wherein, obtaining a target predicted text according to the second text feature vector and the first text feature vector includes: inputting the first text feature vector and the second text feature vector into a connection layer to obtain a temporary text feature vector output by the connection layer; inputting the temporary text feature vector into a normalization layer to obtain the predicted text output by the normalization layer as the target predicted text.

18. A text recognition device, wherein, the device includes: a text line image acquisition module, configured to process an original image to obtain at least one text line image; the text line image only includes one line of text content; a first text vector acquisition module, configured to perform recognition processing on each text line image in sequence to obtain a first text feature vector; a second text vector acquisition module, configured to perform semantic processing on the first text feature vector to obtain a second text feature vector; a target predicted text acquisition module, configured to obtain a target predicted text according to the second text feature vector and the first text feature vector.

19. An electronic device, wherein, includes a processor; a memory for storing a computer program executable by the processor; wherein, the processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 17.

20. A computer-readable storage medium, wherein, when the executable computer program in the storage medium is executed by a processor, it can implement the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Image character recognition method, device and equipment and storage medium

    CN110569846A

  • Character recognition method and device, electronic equipment and storage medium

    CN111539410A

  • Handwritten character recognition method and device, storage medium and terminal

    CN112381057A

  • Text recognition method and device, electronic equipment and readable storage medium

    CN116861851A

  • Image recognition method, device, equipment, medium and program product

    CN117079289A

Cited By

  • Chinese writing error intelligent identification method and system based on deep neural network

    CN120808364A

  • Method and system for automatically generating machining program of numerical control machine tool

    CN121165620A