Mathematical formula recognition method and apparatus, and electronic device and readable storage medium

By adding a position encoding module to the encoder of the formula recognition model and using the Transformer network for feature encoding, the problem of difficult to obtain character correlation in handwritten formulas in the prior art is solved, and the accuracy of mathematical formula recognition is improved.

WO2025112994A1PCT designated stage expired Publication Date: 2025-06-05BOE TECHNOLOGY GROUP CO LTD +1

Patent Information

Application Number
PCT/CN2024/126481
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-10-22
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing mathematical formula recognition technology is difficult to effectively obtain the correlation between characters and characters in handwritten formulas, resulting in low recognition accuracy.

Method used

By adding a position encoding module to the encoder of the formula recognition model, the image feature vector and position encoding vector are synthesized, and the feature synthesis vector is quadratic feature encoding using the Transformer network to enhance the pixel correlation of the region where the formula is located.

Benefits of technology

Improve the accuracy of predicting character vectors, thereby improving the accuracy of restoring mathematical formulas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126481_05062025_PF_FP_ABST
    Figure CN2024126481_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a mathematical formula recognition method and apparatus, and an electronic device and a readable storage medium. The method comprises: acquiring an original image containing a mathematical formula; inputting the original image into a formula recognition model, so as to obtain a predicted character vector that is output by the formula recognition model; the formula recognition model performing enhancement processing on the basis of the correlation between pixels in a region where the formula in the original image is located, so as to obtain an image encoding vector, and performing recognition processing on the image encoding vector to obtain the predicted character vector; and generating the mathematical formula in the original image on the basis of the predicted character vector. In the present embodiment, by means of a formula recognition model, enhancement processing is performed on pixels in a region where a formula in an original image is located, such that the correlation between characters in the original image can be acquired, thereby facilitating an improvement in the accuracy of predicting a character vector, and further improving the accuracy of reducing a mathematical formula.
Need to check novelty before this filing date? Find Prior Art

Description

Mathematical formula recognition method, device, electronic device and readable storage medium Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a mathematical formula recognition method, device, electronic device, and readable storage medium. Background Art

[0002] Currently, handwritten formula recognition is typically implemented using an encoding and decoding architecture. For example, an image of a handwritten formula is fed into an encoder, which extracts an image feature map. This image feature map is then fed into a decoder, which performs character-by-character recognition on the handwritten formula, ultimately recognizing the formula.

[0003] In the existing scheme, the encoder in the encoding-decoding architecture is implemented using the DenseNet98 model, and the decoder can be implemented using the RNN model or the Transformer model.

[0004] Taking the Transformer model as an example, the Transformer decoder can learn the correlation between characters at different positions in the recognized LaTeX string, that is, the correlation between target characters; and it can learn the correlation between the encoding features of the LaTeX string and the original image, that is, the correlation between the source and the target, but it cannot obtain the correlation between sources.

[0005] Summary of the Invention

[0006] The present disclosure provides a mathematical formula recognition method, device, electronic device and readable storage medium to address the deficiencies of related technologies.

[0007] According to a first aspect of an embodiment of the present disclosure, a mathematical formula recognition method is provided, the method comprising:

[0008] Get the original image containing the mathematical formula;

[0009] Inputting the original image into a formula recognition model to obtain a predicted character vector output by the formula recognition model; an encoder of the formula recognition model enhances the weights of pixels in a region where a formula is located in the original image to obtain an image encoding vector, and performing recognition processing on the image encoding vector to obtain the predicted character vector;

[0010] A mathematical formula in the original image is generated according to the predicted character vector.

[0011] Optionally, the formula recognition model includes an encoder and a decoder;

[0012] The encoder is used to obtain an image feature vector corresponding to the original image, obtain a position coding vector corresponding to a pixel position in the image feature vector, and perform feature coding on a feature synthesis vector synthesized from the image feature vector and the position coding vector to obtain an image coding vector;

[0013] The decoder is used to determine the predicted character vector corresponding to the original image according to the image coding vector.

[0014] Optionally, the encoder includes a first image encoding module, a position encoding module, a feature synthesis module, and a second image encoding module; the first image encoding module is connected to the position encoding module and the feature synthesis module respectively; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the second image encoding module; and the second image encoding module is connected to the decoder;

[0015] The first image encoding module is used to encode the original image to obtain an image feature vector corresponding to the original image;

[0016] The position coding module is used to perform position coding processing corresponding to the pixel position in the image feature vector to obtain a position coding vector;

[0017] The feature synthesis module is used to synthesize the image feature vector and the position encoding vector to obtain the feature synthesis vector;

[0018] The second image encoding module is used to perform feature encoding on the feature synthesis vector image to perform enhancement processing based on the correlation between pixels in the area where the formula is located in the original image to obtain the image encoding vector.

[0019] Optionally, the first image encoding module is implemented using a DenseNet network, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the second image encoding module is implemented using a Transformer network.

[0020] Optionally, the second image encoding module includes a multi-head attention unit, a first residual normalization unit, a feedforward network unit and a second residual normalization unit; the input and output of the multi-head attention unit are respectively connected to the first input and second input of the first residual normalization unit; the output of the first residual normalization unit is respectively connected to the input of the feedforward network unit and the first input of the second residual normalization unit; the output of the feedforward network unit is connected to the second input of the second residual normalization unit;

[0021] The multi-head attention unit is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector;

[0022] The first residual normalization unit is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector and then perform normalization processing to obtain a normalized feature vector;

[0023] The feedforward network unit is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector;

[0024] The second residual normalization unit is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together, and then perform normalization processing to obtain the image coding vector.

[0025] Optionally, the formula recognition model is trained by the following steps, including:

[0026] Acquire a training sample data set; the training sample data set includes a plurality of image pairs; each image pair includes a handwritten formula image and a printed formula image, wherein the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image is in handwritten font and the formula of the printed formula image is in printed font;

[0027] Inputting each pair of images into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair;

[0028] Obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model;

[0029] In response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the formula recognition model that has completed training.

[0030] Optionally, inputting each pair of images into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair includes:

[0031] Inputting the handwritten formula image in each pair of images into a first formula recognition model to obtain a first image feature vector and a first predicted character string output by the first formula recognition model; the first image feature vector and the first predicted character string serve as output data of the first formula recognition model;

[0032] The printed formula image in each pair of images is input into the second formula recognition model, and the second formula recognition model outputs a second image feature vector and a second predicted string; the second image feature vector and the second predicted string serve as output data of the second formula recognition model.

[0033] Optionally, obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model includes:

[0034] Calculating a consistency loss value of the first formula recognition model and the second formula recognition model according to the first image feature vector and the second image feature vector;

[0035] Obtaining a first cross entropy loss value of the first formula recognition model according to the first predicted character string and the annotated label of the handwritten formula image;

[0036] Obtaining a second cross entropy loss value of the second formula recognition model according to the second predicted character string and the annotated label of the printed formula image;

[0037] The total loss value of this training is calculated according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value.

[0038] Optionally, calculating a total loss value of this training according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value includes:

[0039] Obtaining an average of the first cross entropy loss value and the second cross entropy loss value to obtain an average cross entropy loss value;

[0040] Obtaining a product of the consistency loss value and a preset weight value to obtain a weighted consistency loss value;

[0041] Obtain the sum of the weighted consistency loss value and the average cross entropy loss value to obtain the total loss value.

[0042] Optionally, obtaining a training sample data set includes:

[0043] Acquire a handwritten formula image and its annotated label; the annotated label is used to represent the formula in the handwritten formula image;

[0044] generating a printed formula image according to the annotated label, wherein the printed formula image and the handwritten formula image form an image pair; the annotated label of the printed formula image and the handwritten formula image is the same;

[0045] Repeat multiple times to obtain multiple pairs of images, and the multiple pairs of images constitute the training sample data set.

[0046] According to a second aspect of an embodiment of the present disclosure, a mathematical formula recognition device is provided, the device comprising:

[0047] An image acquisition module, used to acquire the original image containing the mathematical formula;

[0048] a vector acquisition module, configured to input the original image into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing based on the correlation between pixels in the region where the formula is located in the original image to obtain an image coding vector, and performs recognition processing on the image coding vector to obtain the predicted character vector;

[0049] A formula restoration module is used to generate a mathematical formula in the original image according to the predicted character vector.

[0050] Optionally, the formula recognition model includes an encoder and a decoder;

[0051] The encoder is used to obtain an image feature vector corresponding to the original image, obtain a position coding vector corresponding to a pixel position in the image feature vector, and perform feature coding on a feature synthesis vector synthesized from the image feature vector and the position coding vector to obtain an image coding vector;

[0052] The decoder is used to determine the predicted character vector corresponding to the original image according to the image coding vector.

[0053] Optionally, the encoder includes a first image encoding module, a position encoding module, a feature synthesis module, and a second image encoding module; the first image encoding module is connected to the position encoding module and the feature synthesis module respectively; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the second image encoding module; and the second image encoding module is connected to the decoder;

[0054] The first image encoding module is used to encode the original image to obtain an image feature vector corresponding to the original image;

[0055] The position coding module is used to perform position coding processing corresponding to the pixel position in the image feature vector to obtain a position coding vector;

[0056] The feature synthesis module is used to synthesize the image feature vector and the position encoding vector to obtain the feature synthesis vector;

[0057] The second image encoding module is used to perform feature encoding on the feature synthesis vector image to enhance the pixels in the area where the formula is located in the original image to obtain the image encoding vector.

[0058] Optionally, the first image encoding module is implemented using a DenseNet network, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the second image encoding module is implemented using a Transformer network.

[0059] Optionally, the second image encoding module includes a multi-head attention unit, a first residual normalization unit, a feedforward network unit and a second residual normalization unit; the input and output of the multi-head attention unit are respectively connected to the first input and second input of the first residual normalization unit; the output of the first residual normalization unit is respectively connected to the input of the feedforward network unit and the first input of the second residual normalization unit; the output of the feedforward network unit is connected to the second input of the second residual normalization unit;

[0060] The multi-head attention unit is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector;

[0061] The first residual normalization unit is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector and then perform normalization processing to obtain a normalized feature vector;

[0062] The feedforward network unit is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector;

[0063] The second residual normalization unit is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together, and then perform normalization processing to obtain the image coding vector.

[0064] Optionally, the formula recognition model is trained by the following steps, including:

[0065] Acquire a training sample data set; the training sample data set includes a plurality of image pairs; each image pair includes a handwritten formula image and a printed formula image, wherein the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image is in handwritten font and the formula of the printed formula image is in printed font;

[0066] Inputting each pair of images into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair;

[0067] Obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model;

[0068] In response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the formula recognition model that has completed training.

[0069] Optionally, inputting each pair of images into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair includes:

[0070] Inputting the handwritten formula image in each pair of images into a first formula recognition model to obtain a first image feature vector and a first predicted character string output by the first formula recognition model; the first image feature vector and the first predicted character string serve as output data of the first formula recognition model;

[0071] The printed formula image in each pair of images is input into the second formula recognition model, and the second formula recognition model outputs a second image feature vector and a second predicted string; the second image feature vector and the second predicted string serve as output data of the second formula recognition model.

[0072] Optionally, obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model includes:

[0073] Calculating a consistency loss value of the first formula recognition model and the second formula recognition model according to the first image feature vector and the second image feature vector;

[0074] Obtaining a first cross entropy loss value of the first formula recognition model according to the first predicted character string and the annotated label of the handwritten formula image;

[0075] Obtaining a second cross entropy loss value of the second formula recognition model according to the second predicted character string and the annotated label of the printed formula image;

[0076] The total loss value of this training is calculated according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value.

[0077] Optionally, calculating a total loss value of this training according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value includes:

[0078] Obtaining an average of the first cross entropy loss value and the second cross entropy loss value to obtain an average cross entropy loss value;

[0079] Obtaining a product of the consistency loss value and a preset weight value to obtain a weighted consistency loss value;

[0080] Obtain the sum of the weighted consistency loss value and the average cross entropy loss value to obtain the total loss value.

[0081] Optionally, obtaining a training sample data set includes:

[0082] Acquire a handwritten formula image and its annotated label; the annotated label is used to represent the formula in the handwritten formula image;

[0083] generating a printed formula image according to the annotated label, wherein the printed formula image and the handwritten formula image form an image pair; the annotated label of the printed formula image and the handwritten formula image is the same;

[0084] Repeat multiple times to obtain multiple pairs of images, and the multiple pairs of images constitute the training sample data set.

[0085] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, comprising

[0086] processor;

[0087] a memory for storing a computer program executable by the processor;

[0088] The processor is configured to execute the computer program in the memory to implement the method described in the first aspect.

[0089] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which, when an executable computer program in the storage medium is executed by a processor, can implement the method described in the first aspect.

[0090] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0091] As can be seen from the above embodiments, the solution provided by the embodiments of the present disclosure can obtain an original image containing a mathematical formula; then, the original image is input into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing based on the correlation between pixels in the region where the formula is located in the original image to obtain an image encoding vector, and performs recognition processing on the image encoding vector to obtain the predicted character vector; finally, the mathematical formula in the original image is generated based on the predicted character vector. In this way, this embodiment performs enhancement processing on the pixels in the region where the formula is located in the original image through the formula recognition model, and can obtain the correlation between characters in the original image, which helps to improve the accuracy of the predicted character vector, and thus improve the accuracy of restoring the mathematical formula.

[0092] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0094] Fig. 1 is a flow chart showing a method for recognizing a mathematical formula according to an exemplary embodiment.

[0095] Fig. 2 is a block diagram showing a formula recognition model according to an exemplary embodiment.

[0096] Fig. 3 is a block diagram of an encoder according to an exemplary embodiment.

[0097] FIG4 is a block diagram of a multi-head attention graph unit according to an exemplary embodiment.

[0098] Fig. 5 is a schematic diagram showing visualization of a region of interest in an original image according to an exemplary embodiment.

[0099] Fig. 6 is a block diagram showing a decoder according to an exemplary embodiment.

[0100] Fig. 7 is a block diagram showing a formula recognition model according to an exemplary embodiment.

[0101] FIG8 is a schematic diagram showing a training architecture of a formula recognition model according to an exemplary embodiment.

[0102] FIG9 is a flowchart showing a training formula recognition model according to an exemplary embodiment.

[0103] FIG10 is a schematic diagram showing the relationship between a mathematical formula, a preset formula format and a latex tag according to an exemplary embodiment.

[0104] Fig. 11 is a block diagram showing a mathematical formula recognition device according to an exemplary embodiment. DETAILED DESCRIPTION

[0105] Exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements, unless otherwise indicated. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims. It should be noted that, unless there is a conflict, the features of the following embodiments and implementations may be combined with each other.

[0106] To solve the above technical problems, the embodiments of the present disclosure provide a mathematical formula recognition method, apparatus, electronic device, and readable storage medium. The inventive concepts of each embodiment of the present disclosure include:

[0107] First, adjust the encoder of the formula recognition model and add a position encoding module (referred to as the second image encoding module in subsequent embodiments) to the encoder. The position encoding module continues to perform secondary feature encoding on the feature synthesis vector obtained by synthesizing the image feature vector and the position encoding vector, thereby enhancing the pixels in the focus area (i.e., the area where the formula is located) in the original image. For example, a higher weight value can be assigned to the pixels in the area where the formula is located. Enhancing the focus of the area where the formula is located means enhancing the correlation between the source data and the source data.

[0108] Second, the formula recognition model is trained by combining the adversarial learning concept, and the two formula recognition models are combined to obtain a formula recognition model pair. During the training process, the two formula recognition models (i.e., a formula recognition model pair) are trained simultaneously. That is, the handwritten formula image and the printed formula image are input into the formula recognition model pair, and the loss value is calculated using the output data of each pair of formula recognition models to determine whether the formula recognition model has completed training. After training is completed, the formula recognition model using the handwritten formula image is used in various embodiments of the present disclosure, making it applicable to handwritten formulas of different styles, which is conducive to improving recognition accuracy.

[0109] Based on the above-mentioned inventive concept, the embodiments of the present disclosure provide a mathematical formula recognition method that can be applied to electronic devices, which may include but are not limited to devices with processing capabilities such as smartphones, tablets, smart displays, or electronic whiteboards. Figure 1 is a flowchart of a mathematical formula recognition method according to an exemplary embodiment. Referring to Figure 1, a mathematical formula recognition method includes steps 11 to 13:

[0110] In step 11, an original image containing a mathematical formula is obtained.

[0111] In this embodiment, the electronic device can obtain an original image containing a mathematical formula. In one example, the electronic device can be provided with an image acquisition module, such as a camera or an image sensor. When there is a need to capture the original image, for example, when there is an object in the preview area of ​​the image acquisition module, for example, the object can be a piece of paper with a mathematical formula written on it, the electronic device can control the image acquisition module to shoot the object, thereby obtaining the original image. In another example, the electronic device can be provided with a communication module, such as a WiFi module, an infrared module, a USB module, etc., and the electronic device can communicate with the opposite communication module of other devices through the above communication module, and read the original image containing the mathematical formula from the other device.

[0112] It should be noted that when an electronic device captures an image, it is not certain whether the image contains a mathematical formula. Instead, it can pre-store a formula judgment model that determines whether the image contains mathematical symbols, such as equal signs, addition, subtraction, multiplication, and / or division symbols. When the electronic device determines that a formula exists in the image, it determines that an original image containing a data formula has been captured. Alternatively, the electronic device can assume that all captured images are original images containing mathematical formulas. For ease of description, the subsequent embodiments assume that the image contains mathematical formulas.

[0113] It should also be noted that in the subsequent embodiments, the original image can be a color image (i.e., an RGB image) or a grayscale image. In one example, the original image is implemented as a grayscale image, thereby reducing the data processing load of the formula recognition model and improving recognition efficiency.

[0114] In step 12, the original image is input into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing based on the correlation between pixels in the area where the formula is located in the original image to obtain an image coding vector, and performs recognition processing on the image coding vector to obtain the predicted character vector.

[0115] In this embodiment, the electronic device may store a formula recognition model, which may include an encoder and a decoder.

[0116] Referring to Figure 2, the encoder 21 is used to obtain the image feature vector corresponding to the original image, obtain the position coding vector corresponding to the pixel position in the image feature vector, and perform feature coding on the feature synthesis vector synthesized by the image feature vector and the position coding vector (such as adding the data at the same position of the image feature vector and the position coding vector, or weighting and then summing them) to obtain the image coding vector; the decoder 22 is used to determine the predicted character vector corresponding to the original image based on the image coding vector.

[0117] In one example, referring to FIG3 , the encoder 21 may include a first image encoding module 31, a position encoding module 32, a feature synthesis module 33, and a second image encoding module 34. The first image encoding module 31 is connected to the position encoding module 32 and the feature synthesis module 33 respectively; the position encoding module 32 is connected to the feature synthesis module 33; the feature synthesis module 33 is connected to the second image encoding module 34; and the second image encoding module 34 is connected to the decoder 22.

[0118] The first image encoding module 31 is used to encode the original image to obtain an image feature vector corresponding to the original image;

[0119] The position coding module 32 is used for performing position coding processing on the pixel positions in the image feature vector to obtain a position coding vector;

[0120] The feature synthesis module 33 is used to synthesize the image feature vector and the position coding vector to obtain a feature synthesis vector;

[0121] The second image encoding module 34 is used to perform feature encoding on the feature synthesis vector image to enhance the pixels in the region where the formula in the image feature vector is located, thereby obtaining an image encoding vector.

[0122] In one embodiment, the first image encoding module 31 in the encoder 21 is implemented using a DenseNet network, such as a DenseNet98 network. Referring to Table 1, the DenseNet98 network sequentially includes a convolution layer, a pooling layer, a dense block 1, a transition layer 1, a dense block 2, a transition layer 2, a dense block 3, and a linear transformation layer. Taking dense block 1 as an example, in the DenseNet98 network, dense block 1 includes 16 series units, each unit including 2 operations, namely, one convolution operation with a convolution kernel of 1*1 and a step size of 1 and one convolution operation with a convolution kernel of 3*3 and a step size of 1. The 16 series units are executed 16 times, thereby increasing the number of channels of the input data. The structures and input and output data of the other network layers of the DenseNet98 network can be found in Table 1 and will not be repeated here.

[0123] Table 1 DenseNet98 network structure

[0124] In one embodiment, the position encoding module 32 in the encoder 21 implements position encoding using a sinusoidal position encoding method, as shown in formula (1).

[0125] In formula (1) and formula (2), pos represents the position of the pixel to be encoded in the original image (when position encoding is performed on a two-dimensional image, pos_x and pos_y are calculated and then concat-connected), i represents the position encoding dimension index, and d model Indicates the position encoding dimension (when encoding the original two-dimensional image, d model Equal to half the number of channels of the image to be encoded, for example d model =128).

[0126] In this embodiment, when encoding characters, the position encoding module 32 represents each character with a 256-dimensional vector through the word embedding layer, wherein the word embedding layer is used to construct a dictionary mapping table and construct a 256-dimensional vector for each character; then, position encoding is performed according to the above sine position encoding formula, d model =256, and a sinusoidal position coding value is obtained, namely, the position coding vector in the subsequent embodiments.

[0127] In one embodiment, the second image encoding module 34 in the encoder 21 is implemented using a Transformer network. Continuing with FIG3 , the second image encoding module 34 includes a multi-head attention unit 341, a first residual normalization unit 342, a feedforward network unit 343, and a second residual normalization unit 344. The input and output of the multi-head attention unit 341 are respectively connected to the first input and second input of the first residual normalization unit 342. The output of the first residual normalization unit 342 is respectively connected to the input of the feedforward network unit 343 and the first input of the second residual normalization unit 344. The output of the feedforward network unit 343 is connected to the second input of the second residual normalization unit 344.

[0128] The multi-head attention unit 341 is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector;

[0129] The first residual normalization unit 342 is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector and then perform normalization processing to obtain a normalized feature vector;

[0130] The feedforward network unit 343 is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector;

[0131] The second residual normalization unit 344 is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together, and then perform normalization processing to obtain the image coding vector.

[0132] Referring to Figure 4, there are multiple multi-head attention units 341, each of which has the same structure, and one of the attention units will be described later. The multi-head attention unit 341 includes a scale dot-product unit 41, a first activation layer (softmax) 42, an attention refinement module (ARM) 43, a second activation layer (softmax) 44, and a matrix multiplication (Matmul) unit 45. It can be understood that the multi-head attention unit 341 includes multiple attention units, each of which has the same structure, and one of the attention units will be described later.

[0133] The scaled dot product unit 41 is used to calculate the attention weights for the predicted string (Q) and the image encoding vector (K) to obtain a dot product matrix E. The attention weights represent the correlation between each element in the predicted string Q and each element in K.

[0134] The first activation layer 42 is used to perform weight normalization processing on the dot product matrix E to obtain the attention weight A; wherein the attention weight A represents the correlation between the decoded character sequence and the image encoding vector (or image feature map);

[0135] The attention refining unit 43 is used to obtain the corresponding area of ​​the decoded string in the image coding vector to obtain a refined matrix R; and obtain a difference matrix ER between the dot product matrix E and the refined matrix R; the difference matrix ER represents the unresolved area in the image feature vector (or image feature map), and the above-mentioned unresolved area refers to the corresponding area of ​​the undecoded string in the image coding vector (for example, the actual written content in the original image is 'x+y=1', and after three decodings, the decoded string is 'x+y', and the undecoded string is '=1'), that is, the decoder pays more attention to the unresolved area in the feature map;

[0136] The second activation layer 44 is used to normalize the difference matrix ER to obtain an adjustment matrix (ER)';

[0137] The matrix multiplication unit 45 is used to perform matrix multiplication on the adjustment matrix and the image feature vector (V) to obtain a weighted image feature vector. In the weighted image feature vector, a higher weight is assigned to the area that has not been decoded so that the decoder pays more attention to it.

[0138] In one example, the structure of the multi-head attention unit 341 shown in FIG4 can be expressed using formula (3):

[0139] In formula (3), Q, K and V represent query, key and value respectively; K and V have the same value in the multi-head attention unit 341; d k Indicates the dimension of K.

[0140] Among them, among Q, K and V in the multi-head attention unit 341 shown in Figure 4, Q is the feature synthesis vector output by the feature synthesis module 33, and K and V are the image coding vectors after adding position coding, which are used to determine the resolved areas and unresolved areas in the image feature map, and increase the attention map weight of the resolved area.

[0141] In one embodiment, referring to FIG4 , the specific operation process of the attention refining unit 43 is shown in Formula (4), Formula (5) and Formula (6).

[0142] In equations (4) to (6), T = len(query); L = len(key) = h_0 × w_0, where h_0 and w_0 are the height and width of the encoded feature map output by DenseNet98, respectively; h is the multi-head attention head value (h = 16); and the time step t∈[0,T). C is accumulated by A. By accumulating the attention weights to detect resolved regions, the decoder can also focus on unresolved regions to avoid repeated recognition. It is the result of reshaping C. The dimension of C is T*L*h, where L=h_0*w_0, and the dimension is T*h_0*w_0*h. The convolutional layer generally acts on a one-dimensional plane of height*width, so the L in the C dimension must be expanded to h_0*w_0 through reshaping.

[0143] Referring to FIG5 , the attention refining unit 43 visualizes the refined matrix during the recognition process. The resolved regions in each subgraph are darker in color, indicating that ARM will suppress the attention weights of these resolved regions, thereby encouraging the model to focus on unresolved regions.

[0144] In one embodiment, the feedforward network unit 343 can be implemented using a fully connected neural network, such as a linear layer + a ReLU layer + a linear layer, and its expression is shown in formula (7).

[0145] FFN(x)=max(0,xW1+b1)W2+b2; (7)

[0146] In formula (7), x represents the initial weight data output by the first residual normalization unit 342; xW1+b1 respectively represent the weight values ​​adjusted by the first linear layer; max(0,xW1+b1) represents selecting the larger value from 0 and xW1+b1; max(0,xW1+b1)W2+b2 represents the weight value adjusted by the second linear layer.

[0147] In one embodiment, the first residual normalization unit 342 and the second residual normalization unit 344 use a residual connection so that the output of each layer is added to the input to prevent gradient vanishing and explosion. The expressions of the first residual normalization unit 342 and the second residual normalization unit 344 are shown in Equations (8) and (9), respectively.

[0148] output=layer_norm(Multi-Head Attention(input)+input); (8)

[0149] output 2=layer_norm(FFN(output)+output); (9)

[0150] In formulas (8), (9) and (10), Multi-Head Attention (input) represents the initial string vector output by the multi-head attention unit 341, input represents the feature synthesis vector, layer_norm() represents the normalization process, output represents the initial weight data output by the first residual normalization unit 342, outpu2 represents the image encoding vector output by the second residual normalization unit 344; FFN() represents the operation as in formula (7); x is the input feature, μ and σ 2 are the mean and variance of x, respectively, ε=10 -6 .

[0151] In one embodiment, referring to FIG6 , the decoder 22 includes a character position encoding module 61 , a self-attention module 62 , a multi-head attention map module 63 , a feedforward network module 64 and a linear normalization module 65 .

[0152] The character position encoding module 61 can use sinusoidal position encoding to implement position encoding. Its input is the previous predicted string, and its value in the first prediction is the starting symbol. <sos>, then add each predicted character based on the starting symbol, for example, when predicting for the first time, the value of the previous predicted string is <sos>, the predicted result output is "Jing"; when making the second prediction, the value of the previous predicted string is <sos>Jing, the prediction result output is "Dong"; when the third prediction is made, the value of the previous prediction string is <sos>JD.com, and so on, until the prediction result output is the end character. <eos>End the prediction. Please refer to formula (1) and formula (2) for details, which will not be repeated here.

[0153] The self-attention module 62 can be implemented using a multi-head attention method, specifically referring to the structure shown in FIG4 , which is used to calculate the correlation between characters in the latex string (i.e., obtain the semantic information contained in the predicted string sequence), and its expression is shown in formula (11).

[0154] In formula (11), Q, K, and V are latex strings after character position encoding.

[0155] The implementation scheme of the multi-head attention module 63 is the same as the structure shown in Figure 4 and will not be repeated here.

[0156] The implementation scheme of the feedforward network module 64 is shown in equation (7), which will not be described in detail here.

[0157] Based on the above analysis, an example structure of the formula recognition model in this embodiment is shown in FIG7 .

[0158] In one example, the dimension of the input data of the formula recognition model, ie, the original image, is 80*w*1, where w represents the image width.

[0159] In the encoding stage, the original image is encoded by the image encoding module 31 to obtain an image feature vector of 5*w / 16*256 dimensions; the position encoding module 32 performs position encoding on the image feature vector to obtain a position encoding vector of 5*w / 16*256 dimensions; the feature synthesis module 33 adds the above-mentioned 5*w / 16*256-dimensional image feature vector and the 5*w / 16*256-dimensional position encoding vector to obtain a feature synthesis vector of L*256 dimensions, where L=5*w / 16, L represents the width of the feature synthesis vector after transformation; the second image encoding module 34 encodes the feature synthesis vector twice to obtain an image coding vector of L*256 dimensions.

[0160] During the decoding phase, the character position encoding module 61 in the encoder 22 can position-encode the previously predicted string of length (the length of the predicted string vector) * 1 into a string sequence of length * 256 dimensions. The self-attention module 62 can output a string of length * 256 dimensions. The input data of the multi-head attention module 63 also includes the L * 256-dimensional image encoding vector input by the encoder 21, and its output is a string of length * 256 dimensions. The input and output of the feedforward network module 64 are both strings of length * 256 dimensions. The linear normalization module 65 transforms the length * 256-dimensional string into a word prediction string vector of length * k dimensions, where k is the number of characters supported by the formula recognition model, and the last character is taken as the recognition result. After multiple recognition processes, a predicted string vector can be obtained.

[0161] Considering the diversity of handwritten formula writing styles and the need to improve the robustness of the formula recognition model, this embodiment adopts the idea of ​​adversarial learning to train the formula recognition model shown in Figure 7, and the training framework is shown in Figure 8. Referring to Figure 8, two formula recognition models with the same structure, wherein the output data of the first image encoding module 31 is passed through a linear unit (linear) to calculate the consistency loss value; the predicted character vectors of the two formula recognition models are respectively compared with the annotated labels of the training sample images (subsequent handwritten formula images and printed formula images) to calculate the cross entropy loss value; finally, the total loss value of this training is calculated based on the consistency loss value and the two cross entropy loss values ​​to determine whether the formula recognition model has completed training.

[0162] It should be noted that the purpose of passing the output data of the first image encoding module 31 through the linear unit is to reduce the number of channels of the output data, so as to reduce the amount of calculation when calculating the consistency loss value, which is conducive to speeding up the training speed.

[0163] 9 , the steps of identifying the formula recognition model in this embodiment include steps 91 to 94 .

[0164] In step 91, a training sample data set is obtained; the training sample data set includes multiple pairs of image pairs; each pair of image pairs includes a handwritten formula image and a printed formula image, and the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image adopts a handwritten font and the formula of the printed formula image adopts a printed font.

[0165] In this step, a training sample data set can be obtained. In one example, a handwritten formula image and its annotated label can be obtained; the annotated label is used to represent the formula in the handwritten formula image. Then, a printed formula image is generated based on the annotated label, and the printed formula image and the handwritten formula image constitute a pair of image pairs; the printed formula image and the handwritten formula image have the same annotated label. For example, the matplotlib toolkit in pathon is called, and the annotated label of the handwritten formula image is used as the input parameter of the matplotlib toolkit, and the image output by the matplotlib toolkit is the printed formula image. Repeat multiple times to obtain multiple pairs of image pairs, and the multiple pairs of image pairs constitute the training sample data set.

[0166] In step 92, each pair of images is input into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair.

[0167] In this step, the handwritten formula images in each pair of images can be input into the previous formula recognition model shown in Figure 8 (hereinafter referred to as the first formula recognition model) to obtain the first image feature vector and the first predicted string output by the first formula recognition model; the first image feature vector and the first predicted string serve as the output data of the first formula recognition model; and the printed formula images in each pair of images can be input into the next formula recognition model shown in Figure 8 (hereinafter referred to as the second formula recognition model), and the second image feature vector and the second predicted string output by the above-mentioned second formula recognition model; the second image feature vector and the second predicted string serve as the output data of the second formula recognition model.

[0168] In step 93, the total loss value of the formula recognition model pair is obtained based on the output data of each formula recognition model.

[0169] In this step, the consistency loss value of the first formula recognition model and the second formula recognition model is calculated based on the first image feature vector and the second image feature vector, as shown in formula (12).

[0170] In formula (12), L CL represents the consistency loss value, F h and F p Represent the first image feature vector and the second image feature vector respectively, σ represents the ReLU activation function, F h '=σ(Wσ(F h )) and F p '=σ(Wσ(F p )) represent the normalized feature maps of the first image feature vector and the second image feature vector respectively.

[0171] It is understandable that the above consistency loss value can constrain the consistency of features recognized by the handwritten formula model and the printed formula model, so that the features extracted by the encoder are less affected by the written font as possible, that is, the written font noise is eliminated and the essential features of the characters are directly extracted.

[0172] Then, the first cross-entropy loss value of the first formula recognition model can be obtained based on the first predicted character string and the annotated label of the handwritten formula image, and the second cross-entropy loss value of the second formula recognition model can be obtained based on the second predicted character string and the annotated label of the printed formula image, as shown in Formula (13).

[0173] In formula (13), exp represents the exponential function with the natural constant e as the base, exp(1) = e 1 =2.718…, τ represents temperature. In one example, τ=0.05.

[0174] It can be understood that the first cross entropy loss value and the second cross entropy loss value are used to constrain the divergence between handwritten formulas and printed formulas, so that the formula recognition model has text recognition capabilities.

[0175] Finally, the total loss value of this training is calculated based on the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value. For example, the average of the first cross entropy loss value and the second cross entropy loss value can be obtained to obtain the average cross entropy loss value; then, the product of the consistency loss value and the preset weight value is obtained to obtain the weighted consistency loss value; finally, the sum of the weighted consistency loss value and the average cross entropy loss value is obtained to obtain the total loss value, as shown in formula (14).

[0176] In formula (14), L total Represents the total loss value, the weight coefficient λ = 0.02, and gt represents the annotation label.

[0177] In step 94 , in response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the trained formula recognition model.

[0178] In this step, the relationship between the total loss value and a preset loss value threshold (e.g., 95.0% to 99.9%) can be determined. When the total loss value is greater than the preset loss value threshold, it can be determined that the formula recognition model has not been trained, and the process returns to step 92. When the total loss value is less than or equal to the preset loss value threshold, it can be determined that the formula recognition model has been trained. At this point, the first formula recognition model in the formula recognition model pair can be regarded as the trained formula recognition model and stored in the electronic device.

[0179] In one embodiment, the electronic device can input the original image into the above-mentioned formula recognition model, and the formula recognition model recognizes the mathematical formula in the original image, thereby obtaining a predicted character vector output by the formula recognition model. It is understandable that in the process of obtaining the predicted character vector, the formula recognition model performs one image feature encoding and two position encodings in the encoding stage to increase the correlation between pixels in the original image (i.e., the correlation between source data). For example, if the written content is the Chinese character "recognition", it can be obtained in the encoding stage that the pixels in the area where the character "recognition" is located have a strong correlation, and the pixels in this area have a weak correlation with other areas in the image; the pixels in the area where the character "bie" is located have a strong correlation, and the pixels in this area have a weak correlation with other areas in the image, thereby enabling the encoder to have a stronger feature extraction capability.

[0180] In step 13, a mathematical formula in the original image is generated according to the predicted character vector.

[0181] In this step, the electronic device can restore the mathematical formula in the original image based on the predicted character vector. For example, the electronic device can obtain the characters in the mathematical formula and restore the mathematical formula in a preset formula format, thereby obtaining the mathematical formula in the original image. Taking the mathematical formula including fractions as an example, see Figure 10 to illustrate the relationship between the character position in the mathematical formula and the latex tag. See Figure 11, the formula Above (above) is character a, below (below) is character b, and to the right (right) is character c; when the preset formula format is \frac{above}{below}right, the above formula It can be written as \frac{a}{b}c.

[0182] At this point, the solution provided by the embodiment of the present disclosure can obtain an original image containing a mathematical formula; then, the original image is input into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing based on the correlation between pixels in the region where the formula is located in the original image to obtain an image encoding vector, and performs recognition processing on the image encoding vector to obtain the predicted character vector; finally, the mathematical formula in the original image is generated based on the predicted character vector. In this way, this embodiment performs enhancement processing on the pixels in the region where the formula is located in the original image through the formula recognition model, and can obtain the correlation between characters in the original image, which helps to improve the accuracy of the predicted character vector, and thus improve the accuracy of restoring the mathematical formula.

[0183] Based on the mathematical formula recognition method provided in the embodiment of the present disclosure, this embodiment further provides a mathematical formula recognition device, as shown in FIG11 , the device includes:

[0184] An image acquisition module 111 is used to acquire an original image containing a mathematical formula;

[0185] The vector acquisition module 112 is configured to input the original image into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing based on the correlation between pixels in the region where the formula is located in the original image to obtain an image coding vector, and performs recognition processing on the image coding vector to obtain the predicted character vector;

[0186] The formula restoration module 113 is configured to generate the mathematical formula in the original image according to the predicted character vector.

[0187] In one embodiment, the formula recognition model includes an encoder and a decoder;

[0188] The encoder is used to obtain an image feature vector corresponding to the original image, obtain a position coding vector corresponding to a pixel position in the image feature vector, and perform feature coding on a feature synthesis vector synthesized from the image feature vector and the position coding vector to obtain an image coding vector;

[0189] The decoder is used to determine the predicted character vector corresponding to the original image according to the image coding vector.

[0190] In one embodiment, the encoder includes a first image encoding module, a position encoding module, a feature synthesis module, and a second image encoding module; the first image encoding module is connected to the position encoding module and the feature synthesis module respectively; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the second image encoding module; and the second image encoding module is connected to the decoder;

[0191] The first image encoding module is used to encode the original image to obtain an image feature vector corresponding to the original image;

[0192] The position coding module is used to perform position coding processing corresponding to the pixel position in the image feature vector to obtain a position coding vector;

[0193] The feature synthesis module is used to synthesize the image feature vector and the position encoding vector to obtain the feature synthesis vector;

[0194] The second image encoding module is used to perform feature encoding on the feature synthesis vector image to enhance the pixels in the area where the formula is located in the original image to obtain the image encoding vector.

[0195] Optionally, the first image encoding module is implemented using a DenseNet network, the position encoding module uses a sinusoidal position encoding method to implement position encoding, and the second image encoding module is implemented using a Transformer network.

[0196] In one embodiment, the second image encoding module includes a multi-head attention unit, a first residual normalization unit, a feedforward network unit, and a second residual normalization unit; the input and output of the multi-head attention unit are respectively connected to the first input and second input of the first residual normalization unit; the output of the first residual normalization unit is respectively connected to the input of the feedforward network unit and the first input of the second residual normalization unit; the output of the feedforward network unit is connected to the second input of the second residual normalization unit;

[0197] The multi-head attention unit is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector;

[0198] The first residual normalization unit is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector and then perform normalization processing to obtain a normalized feature vector;

[0199] The feedforward network unit is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector;

[0200] The second residual normalization unit is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together, and then perform normalization processing to obtain the image coding vector.

[0201] In one embodiment, the formula recognition model is trained by the following steps, including:

[0202] Acquire a training sample data set; the training sample data set includes a plurality of image pairs; each image pair includes a handwritten formula image and a printed formula image, wherein the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image is in handwritten font and the formula of the printed formula image is in printed font;

[0203] Inputting each pair of images into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair;

[0204] Obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model;

[0205] In response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the formula recognition model that has completed training.

[0206] In one embodiment, each pair of images is input into a formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair, including:

[0207] Inputting the handwritten formula image in each pair of images into a first formula recognition model to obtain a first image feature vector and a first predicted character string output by the first formula recognition model; the first image feature vector and the first predicted character string serve as output data of the first formula recognition model;

[0208] The printed formula image in each pair of images is input into the second formula recognition model, and the second formula recognition model outputs a second image feature vector and a second predicted string; the second image feature vector and the second predicted string serve as output data of the second formula recognition model.

[0209] In one embodiment, obtaining the total loss value of the formula recognition model pair according to the output data of each formula recognition model includes:

[0210] Calculating a consistency loss value of the first formula recognition model and the second formula recognition model according to the first image feature vector and the second image feature vector;

[0211] Obtaining a first cross entropy loss value of the first formula recognition model according to the first predicted character string and the annotated label of the handwritten formula image;

[0212] Obtaining a second cross entropy loss value of the second formula recognition model according to the second predicted character string and the annotated label of the printed formula image;

[0213] The total loss value of this training is calculated according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value.

[0214] In one embodiment, calculating a total loss value of this training according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value includes:

[0215] Obtaining an average of the first cross entropy loss value and the second cross entropy loss value to obtain an average cross entropy loss value;

[0216] Obtaining a product of the consistency loss value and a preset weight value to obtain a weighted consistency loss value;

[0217] Obtain the sum of the weighted consistency loss value and the average cross entropy loss value to obtain the total loss value.

[0218] In one embodiment, obtaining a training sample data set includes:

[0219] Acquire a handwritten formula image and its annotated label; the annotated label is used to represent the formula in the handwritten formula image;

[0220] generating a printed formula image according to the annotated label, wherein the printed formula image and the handwritten formula image form an image pair; the annotated label of the printed formula image and the handwritten formula image is the same;

[0221] Repeat multiple times to obtain multiple pairs of images, and the multiple pairs of images constitute the training sample data set.

[0222] It should be noted that the device embodiment shown in this embodiment matches the content of the above-mentioned method embodiment. You can refer to the content of the above-mentioned method embodiment and will not repeat it here.

[0223] In an exemplary embodiment, an electronic device is also provided, comprising

[0224] Display screen;

[0225] processor;

[0226] a memory for storing a computer program executable by the processor;

[0227] The processor is configured to execute the computer program in the memory to implement the above method.

[0228] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including an executable computer program. The executable computer program can be executed by a processor to implement the method of the above embodiment. The computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0229] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0230] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.< / eos> < / sos> < / sos> < / sos> < / sos>

Claims

1. A mathematical formula recognition method, characterized in that: The method comprises: Get the original image containing the mathematical formula; Inputting the original image into a formula recognition model to obtain a predicted character vector output by the formula recognition model; The formula recognition model performs enhancement processing according to the correlation between pixels in the region where the formula is located in the original image to obtain an image coding vector, and performs recognition processing on the image coding vector to obtain the predicted character vector; A mathematical formula in the original image is generated according to the predicted character vector.

2. The method according to claim 1, characterized in that The formula recognition model includes an encoder and a decoder; The encoder is used to obtain an image feature vector corresponding to the original image, obtain a position coding vector corresponding to a pixel position in the image feature vector, and perform feature coding on a feature synthesis vector synthesized from the image feature vector and the position coding vector to obtain an image coding vector; The decoder is used to determine the predicted character vector corresponding to the original image according to the image encoding vector.

3. The method according to claim 2, characterized in that The encoder comprises a first image encoding module, a position encoding module, a feature synthesis module and a second image encoding module; the first image encoding module is connected to the position encoding module and the feature synthesis module respectively; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the second image encoding module; the second image encoding module is connected to the decoder; The first image encoding module is used to encode the original image to obtain an image feature vector corresponding to the original image; The position coding module is used to process the position coding corresponding to the pixel position in the image feature vector to obtain a position coding vector; The feature synthesis module is used to synthesize the image feature vector and the position encoding vector to obtain the feature synthesis vector; The second image encoding module is used to perform feature encoding on the feature synthesis vector image to perform enhancement processing according to the correlation between pixels in the area where the formula is located in the original image, so as to obtain the image encoding vector.

4. The method according to claim 3, characterized in that The first image encoding module is implemented by using a DenseNet network, the position encoding module implements position encoding by using a sinusoidal position encoding method, and the second image encoding module is implemented by using a Transformer network.

5. The method according to claim 4, characterized in that The second image encoding module includes a multi-head attention unit, a first residual normalization unit, a feedforward network unit and a second residual normalization unit; the input end and the output end of the multi-head attention unit are respectively connected to the first input end and the second input end of the first residual normalization unit; the output end of the first residual normalization unit is respectively connected to the input end of the feedforward network unit and the first input end of the second residual normalization unit; the output end of the feedforward network unit is connected to the second input end of the second residual normalization unit; The multi-head attention unit is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector; The first residual normalization unit is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector, add them together and then perform normalization processing to obtain a normalized feature vector; The feedforward network unit is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image coding vector; The second residual normalization unit is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together and then perform normalization processing to obtain the image coding vector.

6. The method according to any one of claims 1 to 5, characterized in that: The formula recognition model is trained by the following steps, including: Acquire a training sample data set; the training sample data set includes a plurality of image pairs; each image pair includes a handwritten formula image and a printed formula image, and the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image is in handwritten font and the formula of the printed formula image is in printed font; Inputting each pair of images into a pair of formula recognition models to obtain output data of each formula recognition model in the pair of formula recognition models; Obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model; In response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the formula recognition model that has completed training.

7. The method according to claim 6, characterized in that Inputting each pair of images into the formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair includes: Inputting the handwritten formula image in each pair of images into a first formula recognition model to obtain a first image feature vector and a first predicted character string output by the first formula recognition model; the first image feature vector and the first predicted character string serve as output data of the first formula recognition model; The printed formula image in each pair of images is input into the second formula recognition model, and the second formula recognition model outputs a second image feature vector and a second predicted character string; the second image feature vector and the second predicted character string serve as output data of the second formula recognition model.

8. The method according to claim 7, characterized in that Obtaining the total loss value of the formula recognition model pair according to the output data of each formula recognition model includes: Calculating a consistency loss value of the first formula recognition model and the second formula recognition model according to the first image feature vector and the second image feature vector; Acquire a first cross entropy loss value of the first formula recognition model according to the first predicted character string and the annotated label of the handwritten formula image; Acquire a second cross entropy loss value of the second formula recognition model according to the second predicted character string and the annotated label of the printed formula image; The total loss value of this training is calculated according to the consistency loss value, the first cross entropy loss value and the second cross entropy loss value.

9. The method according to claim 8, characterized in that Calculating a total loss value of this training according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value includes: Obtaining an average of the first cross entropy loss value and the second cross entropy loss value to obtain an average cross entropy loss value; Obtaining the product of the consistency loss value and a preset weight value to obtain a weighted consistency loss value; The sum of the weighted consistency loss value and the average cross entropy loss value is obtained to obtain the total loss value.

10. The method according to claim 6, characterized in that Get the training sample data set, including: Acquire a handwritten formula image and its annotated label; the annotated label is used to represent the formula in the handwritten formula image; Generate a printed formula image according to the annotated label, wherein the printed formula image and the handwritten formula image form an image pair; the annotated labels of the printed formula image and the handwritten formula image are the same; The process is repeated multiple times to obtain multiple pairs of images, which constitute the training sample data set.

11. A mathematical formula recognition device, characterized in that: The device comprises: An image acquisition module, used for acquiring an original image containing a mathematical formula; A vector acquisition module, used for inputting the original image into a formula recognition model to obtain a predicted character vector output by the formula recognition model; the formula recognition model performs enhancement processing according to the correlation between pixels in the region where the formula is located in the original image to obtain an image coding vector, and performs recognition processing on the image coding vector to obtain the predicted character vector; A formula restoration module is used to generate a mathematical formula in the original image according to the predicted character vector.

12. The device according to claim 11, characterized in that The formula recognition model includes an encoder and a decoder; The encoder is used to obtain an image feature vector corresponding to the original image, obtain a position coding vector corresponding to a pixel position in the image feature vector, and perform feature coding on a feature synthesis vector synthesized from the image feature vector and the position coding vector to obtain an image coding vector; The decoder is used to determine the predicted character vector corresponding to the original image according to the image encoding vector.

13. The device according to claim 12, characterized in that The encoder comprises a first image encoding module, a position encoding module, a feature synthesis module and a second image encoding module; the first image encoding module is connected to the position encoding module and the feature synthesis module respectively; the position encoding module is connected to the feature synthesis module; the feature synthesis module is connected to the second image encoding module; the second image encoding module is connected to the decoder; The first image encoding module is used to encode the original image to obtain an image feature vector corresponding to the original image; The position coding module is used to process the position coding corresponding to the pixel position in the image feature vector to obtain a position coding vector; The feature synthesis module is used to synthesize the image feature vector and the position encoding vector to obtain the feature synthesis vector; The second image encoding module is used to perform feature encoding on the feature synthesis vector image so as to enhance the pixels in the area where the formula is located in the original image to obtain the image encoding vector.

14. The device according to claim 13, characterized in that The first image encoding module is implemented by using a DenseNet network, the position encoding module implements position encoding by using a sinusoidal position encoding method, and the second image encoding module is implemented by using a Transformer network.

15. The device according to claim 14, characterized in that The second image encoding module includes a multi-head attention unit, a first residual normalization unit, a feedforward network unit and a second residual normalization unit; the input end and the output end of the multi-head attention unit are respectively connected to the first input end and the second input end of the first residual normalization unit; the output end of the first residual normalization unit is respectively connected to the input end of the feedforward network unit and the first input end of the second residual normalization unit; the output end of the feedforward network unit is connected to the second input end of the second residual normalization unit; The multi-head attention unit is used to perform weighted processing on the feature synthesis vector image to obtain a weighted image feature vector; The first residual normalization unit is used to calculate the residual data of the feature synthesis vector and the weighted image feature vector, add them together and then perform normalization processing to obtain a normalized feature vector; The feedforward network unit is used to perform nonlinear transformation processing on the normalized feature vector to obtain an initial image Encoding vector; The second residual normalization unit is used to calculate the residual data of the initial image coding vector and the normalized feature vector, add them together and then perform normalization processing to obtain the image coding vector.

16. The device according to any one of claims 11 to 15, characterized in that: The formula recognition model is trained by the following steps, including: Acquire a training sample data set; the training sample data set includes a plurality of image pairs; each image pair includes a handwritten formula image and a printed formula image, and the handwritten formula image and the printed formula image contain the same formula, and the formula of the handwritten font image is in handwritten font and the formula of the printed formula image is in printed font; Inputting each pair of images into a pair of formula recognition models to obtain output data of each formula recognition model in the pair of formula recognition models; Obtaining a total loss value of the formula recognition model pair according to output data of each formula recognition model; In response to the total loss value being less than or equal to a preset loss value threshold, the first formula recognition model in the formula recognition model pair is determined as the formula recognition model that has completed training.

17. The device according to claim 16, characterized in that Inputting each pair of images into the formula recognition model pair to obtain output data of each formula recognition model in the formula recognition model pair includes: Inputting the handwritten formula image in each pair of images into a first formula recognition model to obtain a first image feature vector and a first predicted character string output by the first formula recognition model; the first image feature vector and the first predicted character string serve as output data of the first formula recognition model; The printed formula image in each pair of images is input into the second formula recognition model, and the second formula recognition model outputs a second image feature vector and a second predicted character string; the second image feature vector and the second predicted character string serve as output data of the second formula recognition model.

18. The device according to claim 17, characterized in that Obtaining the total loss value of the formula recognition model pair according to the output data of each formula recognition model includes: Calculating a consistency loss value of the first formula recognition model and the second formula recognition model according to the first image feature vector and the second image feature vector; Acquire a first cross entropy loss value of the first formula recognition model according to the first predicted character string and the annotated label of the handwritten formula image; Acquire a second cross entropy loss value of the second formula recognition model according to the second predicted character string and the annotated label of the printed formula image; The total loss value of this training is calculated according to the consistency loss value, the first cross entropy loss value and the second cross entropy loss value.

19. The device according to claim 18, characterized in that Calculating a total loss value of this training according to the consistency loss value, the first cross entropy loss value, and the second cross entropy loss value includes: Obtaining an average of the first cross entropy loss value and the second cross entropy loss value to obtain an average cross entropy loss value; Obtaining the product of the consistency loss value and a preset weight value to obtain a weighted consistency loss value; The sum of the weighted consistency loss value and the average cross entropy loss value is obtained to obtain the total loss value.

20. The device according to claim 16, characterized in that Get the training sample data set, including: Acquire a handwritten formula image and its annotated label; the annotated label is used to represent the formula in the handwritten formula image; Generate a printed formula image according to the annotated label, wherein the printed formula image and the handwritten formula image form an image pair; the annotated labels of the printed formula image and the handwritten formula image are the same; The process is repeated multiple times to obtain multiple pairs of images, which constitute the training sample data set.

21. An electronic device, characterized in that: include A processor; a memory for storing a computer program executable by the processor; The processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that: When the executable computer program in the storage medium is executed by a processor, the method according to any one of claims 1 to 10 can be implemented.

Citation Information

Patent Citations

  • Character recognition method and device, electronic equipment and storage medium

    CN110866529A

  • Time sequence attention mechanism scene image recognition method

    CN113688822A

  • Formula identification method, related device, equipment and storage medium

    CN114359925A

  • Mathematical formula identification method and device, electronic equipment and readable storage medium

    CN116597457A

  • Image data processing method and device, equipment and medium

    CN116977663A

Cited By

  • Formula identification and model training method and device, related equipment and program product

    CN120544225A