Image recognition model training method, image recognition method, device and equipment

By introducing a decoding mechanism of target semantic features and bidirectional attention weights into the formula recognition model, the problems of low accuracy and slow speed in the recognition of handwritten mathematical formulas in the existing technology are solved, and efficient and accurate recognition of complex mathematical formulas is achieved.

CN116071758BActive Publication Date: 2026-02-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310118570.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-31
Publication Date
2026-02-13
Estimated Expiration
2043-01-31

AI Technical Summary

Technical Problem

Existing handwritten mathematical formula recognition methods suffer from low accuracy and efficiency under complex structures and diverse writing styles, and slow decoding speed. In particular, deep learning-based methods suffer from slow decoding speed due to the randomness of attention movement trajectories leading to repeated or missed recognition of symbols.

Method used

By introducing target semantic features and attention weights of target symbols into the formula recognition model, bidirectional decoders (such as transformer models and GRU networks) are used for symbol recognition. The formula recognition model is trained by combining forward and reverse attention weights, and attention distillation is introduced to improve decoding accuracy.

Benefits of technology

It improves the accuracy and decoding speed of complex mathematical formulas, reduces repeated decoding, and enhances the overall performance of the formula recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071758B_ABST
    Figure CN116071758B_ABST
Patent Text Reader

Abstract

The present disclosure provides a kind of training method of image recognition model, image recognition method, device and equipment, it is related to artificial intelligence technical field, specifically is deep learning, image processing, computer vision technical field, can be applied to OCR etc. Scene.The specific implementation scheme is: according to the target semantic feature of target symbol in target formula image, the target attention weight of the target symbol is determined;According to the target semantic feature and the target attention weight, the recognition result of the target symbol is determined;According to the recognition result, the target attention weight and the label data of the target symbol, formula recognition model is trained.Through the above technical scheme, the accuracy of formula recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of deep learning, image processing and computer vision, and can be applied to the scene of OCR. BACKGROUND

[0002] Mathematical formulas exist widely in the fields of education, office and library archives. Unlike regular text, handwritten mathematical formulas have complex spatial structures and diversified writing styles. The complex spatial structures are mainly caused by the unique structures of mathematical formulas, such as fractions, superscripts, subscripts and radicals. The diversity of structures and writing styles leads to unsatisfactory recognition results of existing handwritten mathematical recognition methods. SUMMARY

[0003] The present disclosure provides a training method of an image recognition model, an image recognition method, an apparatus and a device.

[0004] According to an aspect of the present disclosure, a training method of an image recognition model is provided, which comprises:

[0005] According to a target semantic feature of a target symbol in a target formula image, a target attention weight of the target symbol is determined;

[0006] According to the target semantic feature and the target attention weight, a recognition result of the target symbol is determined;

[0007] According to the recognition result, the target attention weight and label data of the target symbol, a formula recognition model is trained.

[0008] According to another aspect of the present disclosure, an image recognition method is provided, which comprises:

[0009] An image to be recognized is obtained;

[0010] A formula recognition model is used to perform symbol recognition on the image to be recognized, to obtain a target recognition result of the image to be recognized; wherein the formula recognition model is trained according to the training method of the image recognition model provided in any embodiment of the present disclosure.

[0011] According to another aspect of the present disclosure, a training apparatus of an image recognition model is provided, which comprises:

[0012] A target weight determination module is configured to determine a target attention weight of a target symbol in a target formula image according to a target semantic feature of the target symbol;

[0013] A recognition result determination module is configured to determine a recognition result of the target symbol according to the target semantic feature and the target attention weight.

[0014] a model training module configured to train a formula recognition model according to the recognition result, the target attention weight, and label data of the target symbol.

[0015] According to another aspect of the present disclosure, an image recognition device is provided, comprising:

[0016] a to-be-recognized image acquisition module configured to acquire a to-be-recognized formula image;

[0017] a target recognition result determination module configured to perform symbol recognition on the to-be-recognized formula image by using a formula recognition model to obtain a target recognition result of the to-be-recognized formula; wherein the formula recognition model is trained according to the training method of the image recognition model provided in any of the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0019] at least one processor; and

[0020] a memory in communication with the at least one processor; wherein

[0021] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the training method of the image recognition model or the image recognition method provided in any of the embodiments of the present disclosure.

[0022] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the training method of the image recognition model or the image recognition method provided in any of the embodiments of the present disclosure.

[0023] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the training method of the image recognition model or the image recognition method provided in any of the embodiments of the present disclosure.

[0024] According to the technology of the present disclosure, the accuracy of formula recognition can be improved.

[0025] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0027] Figure 1 is a flowchart of a training method of an image recognition model according to an embodiment of the present disclosure;

[0028] Figure 2 is a flowchart of another training method of an image recognition model according to an embodiment of the present disclosure;

[0029] Figure 3 is a flowchart of yet another training method of an image recognition model according to an embodiment of the present disclosure;

[0030] Figure 4 is a flowchart of an image recognition method according to an embodiment of the present disclosure;

[0031] Figure 5 is a structural schematic diagram of a training device of an image recognition model according to an embodiment of the present disclosure;

[0032] Figure 6 is a structural schematic diagram of an image recognition device according to an embodiment of the present disclosure;

[0033] Figure 7 is a block diagram of an electronic device for implementing the training method of an image recognition model or the image recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, in which various details of the present disclosure are set forth to facilitate an understanding. However, it will be appreciated that various embodiments of the present disclosure can be practiced without such specific details. In other instances, well-known methods, structures and techniques have not been described in detail in order to avoid obscuring the present disclosure. The following detailed description is, therefore, not to be taken in a limiting sense, as the scope of the present disclosure is defined by the appended claims.

[0035] It is to be understood that the terms "first", "second", "target", etc. in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0036] In addition, it also needs to be explained that the collection, storage, use, processing, transmission, provision and disclosure of the target formula image and the to-be-identified formula image and the like involved in the technical solutions of the present application all comply with the relevant laws and regulations and do not violate public order and good customs.

[0037] Currently, the handwritten mathematical formula recognition method can be divided into three categories. The first category is the traditional handwritten mathematical formula recognition method, which generally includes three steps: symbol segmentation, symbol recognition, and structure analysis. In the symbol recognition step, the hidden Markov algorithm, the elastic matching algorithm, and the support vector machine algorithm are often used. The second category is the sequence-based handwritten mathematical formula recognition method based on deep learning, which regards the handwritten mathematical formula recognition as an image-to-sequence translation task and uses the attention mechanism to decode each symbol in the mathematical formula one by one. The third category is the structured handwritten mathematical formula recognition method based on deep learning, which regards each mathematical formula as a syntax tree and decodes from the parent node to the child node of the tree, and then converts the tree decoding result into a mathematical formula according to the pre-defined rules.

[0038] The existing several technical solutions have the following defects:

[0039] 1) The first scheme is limited by the single feature learning ability and the overly complex grammar rules, and has the problems of low recognition accuracy and low efficiency when dealing with real scene handwritten mathematical formulas;

[0040] 2) The attention trajectory in the image is random when the second scheme recognizes the mathematical formula. Therefore, when decoding a complex mathematical formula, the attention may not be accurate, resulting in repeated recognition of a symbol or missing recognition of a symbol;

[0041] 3) The third scheme needs to find the parent node of the current node when decoding each symbol, and combines the features of the parent node to predict the category of the current node, which results in slow decoding speed.

[0042] Figure 1 is a flowchart of a method for training an image recognition model according to an embodiment of the present disclosure. The method is applicable to the case of how to recognize a formula image, and is particularly applicable to the case of how to recognize a handwritten mathematical formula. The method can be performed by a training device of an image recognition model, which can be realized in software and / or hardware, and can be integrated in an electronic device that carries the training function of the image recognition model, such as a server. As shown in Figure 1 The method for training an image recognition model according to the present embodiment can include:

[0043] S101, determining a target attention weight of a target symbol according to a target semantic feature of the target symbol in a target formula image.

[0044] In this embodiment, the target formula image refers to an image containing a formula used for training the formula recognition model. The target symbol refers to a symbol in the target formula image, including but not limited to letters, operators, etc. It should be noted that the number of target symbols can be one or more.

[0045] The target semantic feature refers to an image feature used to represent the target symbol in the target formula image, which can be represented in the form of a matrix or a vector. In an optional manner, the target semantic feature of each target symbol in the target formula image can be obtained by feature encoding of the target formula image. The feature encoder can be a deep neural network, such as a ResNet-50 network. Specifically, the target formula image can be input into the ResNet-50 network to obtain a feature map output by each level of the network, and a high-level feature map can be selected as the target semantic feature of each target symbol in the target formula image. It should be noted that the high-level feature map contains more semantic information about the symbol.

[0046] The target attention weight refers to the regional weight of the target semantic feature corresponding to the target symbol in the image feature corresponding to the target formula image in the recognition process of the target symbol. That is, in the recognition process of the target symbol, the formula recognition model can focus on the region of the target symbol in the target formula image, and mainly use the features corresponding to the region for symbol recognition.

[0047] In an optional manner, the target semantic feature of the target symbol in the target formula image can be processed by a feature decoder to obtain the target attention weight of the target symbol. The feature decoder can be a transformer model.

[0048] In another optional manner, a bidirectional decoder can also be used to process the target semantic feature of the target symbol in the target formula image to obtain a forward target attention weight and a reverse attention weight. The bidirectional decoder can be two parallel transformer models.

[0049] In S102, the recognition result of the target symbol is determined according to the target semantic feature and the target attention weight.

[0050] In this embodiment, the recognition result refers to the result of recognizing the target symbol, which can be represented in a sequence form, i.e., a sequence after splitting according to the formula writing steps.

[0051] Specifically, the target semantic feature and the target attention weight can be fused based on a preset feature fusion manner to obtain a fusion feature of the target symbol. For example, the target semantic feature and the target attention weight can be multiplied to obtain the fusion feature of the target symbol; or the target semantic feature and the target attention weight can be spliced to obtain the fusion feature of the target symbol. Then, the fusion feature is recognized to obtain a recognition result of the target symbol.

[0052] S103, training the formula recognition model according to the recognition result, the target attention weight and label data of the target symbol.

[0053] In this embodiment, the label data of the target symbol refers to data for labeling the target symbol in the target formula image.

[0054] Specifically, the recognition result of the target symbol, the label data and the target attention weight can be loss calculated based on a preset loss function to obtain a training loss; the preset loss function is not limited, for example, it can be a cross-entropy loss function. Then, the formula recognition model is trained using the training loss until a training stop condition is met. The training stop condition can be that the number of training iterations reaches a number threshold, or the training loss is stable at a set threshold; the number threshold and the set threshold can be set by a person skilled in the art according to actual conditions.

[0055] The technical scheme provided by the embodiment of the present disclosure determines the target attention weight of the target symbol according to the target semantic feature of the target symbol in the target formula image, then determines the recognition result of the target symbol according to the target semantic feature and the target attention weight, and further trains the formula recognition model according to the recognition result, the target attention weight and the label data of the target symbol. The above technical scheme introduces attention weight in the training process of the formula recognition model, and the positioning ability of the formula recognition model for symbols in the decoding process can be improved by using attention weight, so as to improve the recognition accuracy of the formula recognition model for complex formulas (especially mathematical formulas).

[0056] Figure 2 is a flowchart of another image recognition model training method provided by the embodiment of the present disclosure. The embodiment further optimizes the determination of the target attention weight of the target symbol according to the target semantic feature of the target symbol in the target formula image based on the above-mentioned embodiment, and provides an optional implementation scheme. As shown in Figure 2 The image recognition model training method of the embodiment can include:

[0057] S201, forward decoding the target semantic feature of the target symbol to obtain a target forward attention weight of the target symbol.

[0058] In this embodiment, the target forward attention weight refers to the attention weight obtained by performing forward decoding on the target semantic feature of the target symbol.

[0059] In an optional manner, a forward decoder can be used to decode the target semantic feature of each symbol in the target formula image in a left-to-right order to obtain the target forward attention weight of the target symbol. Specifically, the forward decoder can be used to perform attention operation on the target semantic feature of the target symbol to obtain the target forward attention weight of the target symbol. The forward decoder can be a transformer model that decodes in a left-to-right order.

[0060] In another optional manner, a first semantic feature and a first forward attention weight of a first symbol located before the target symbol in the forward decoding direction are determined, a first fusion feature of the target symbol is determined according to the target semantic feature of the target symbol and the first semantic feature and the first forward attention weight of the first symbol, and attention operation is performed on the first fusion feature to obtain the target forward attention weight of the target symbol.

[0061] The first symbol refers to a symbol located before the target symbol in the forward decoding direction (i.e., from left to right) in the formula, which can be one or more symbols before the target formula. The first semantic feature refers to a feature obtained by processing the semantic feature of the first symbol. The first forward attention weight refers to an attention weight obtained by performing forward decoding on the semantic feature of the first symbol. The first fusion feature refers to a feature obtained by fusing the first semantic feature and the first forward attention weight.

[0062] Specifically, a Gated Recurrent Unit (GRU) network can be used to process the semantic feature of the first symbol to obtain the first semantic feature of the first symbol, and the first forward attention weight of the first symbol is determined. Then, a preset fusion manner can be used to fuse the target semantic feature of the target symbol and the first semantic feature and the first forward attention weight of the first symbol, for example, the target semantic feature of the target symbol and the first semantic feature and the first forward attention weight of the first symbol can be added, or the target semantic feature of the target symbol and the first semantic feature and the first forward attention weight of the first symbol can be spliced to obtain the first fusion feature of the target symbol. Then, attention operation is performed on the first fusion feature to obtain the target forward attention weight of the target symbol.

[0063] It can be understood that the GRU network has stronger ability to retain long-term sequence information, and the first semantic feature of the symbol before the target symbol is determined by using the GRU network, so that the semantic information generated by the symbol before the target symbol can be retained when the target symbol is recognized, so that the target forward attention weight of the target symbol is more accurate, thereby improving the prediction result of the target symbol; at the same time, when determining the target forward attention weight of the target symbol, the forward attention weight of the symbol before the target symbol is introduced, that is, the historical attention weight is introduced, which can avoid the formula model from remembering the region that has been paid attention to when recognizing the target symbol, thereby preventing the same region from being repeatedly decoded, and more conducive to the generation of the target forward attention weight of the target symbol; in this way, in the decoding and recognition process, the GRU and the attention mechanism are combined, which can better recognize longer and more complex formulas (especially mathematical formulas).

[0064] S202, the target semantic feature of the target symbol is reversely decoded to obtain the target reverse attention weight of the target symbol.

[0065] In this embodiment, the target reverse attention weight refers to the attention weight obtained by reversely decoding the target semantic feature of the target symbol.

[0066] An optional way can be to use a reverse decoder to decode the target semantic feature of each symbol in the target formula image in the order from right to left to obtain the target reverse attention weight of the target symbol. Specifically, the target semantic feature of the target symbol can be subjected to attention operation by the reverse decoder to obtain the target reverse attention weight of the target symbol. The reverse decoder can be a transformer model that decodes in the order from right to left.

[0067] An optional way is to determine the second semantic feature and the second reverse attention weight of the second symbol before the target symbol in the reverse decoding direction; determine the second fusion feature of the target symbol according to the target semantic feature of the target symbol, and the second semantic feature and the second reverse attention weight of the second symbol; and perform attention operation on the second fusion feature to obtain the target reverse attention weight of the target symbol.

[0068] The second symbol refers to the symbol before the target symbol in the reverse decoding direction (i.e. from right to left) in the formula, which can be one or more symbols before the target formula. The second semantic feature refers to the feature obtained by processing the semantic feature of the second symbol. The second reverse attention weight refers to the attention weight obtained by forward decoding the semantic feature of the second symbol. The second fusion feature refers to the feature obtained by fusing the second semantic feature and the second reverse attention weight.

[0069] Specifically, a Gated Recurrent Unit (GRU) model can be used to process the semantic features of the second symbol to obtain second semantic features of the second symbol; the second reverse attention weight of the second symbol is determined; then a preset fusion manner can be used to fuse the target semantic features of the target symbol, the second semantic features of the second symbol and the second reverse attention weight, for example, the target semantic features of the target symbol, the second semantic features of the second symbol and the second reverse attention weight can be added, or the target semantic features of the target symbol, the second semantic features of the second symbol and the second reverse attention weight can be spliced to obtain second fusion features of the target symbol; and then attention operation is performed on the second fusion features to obtain the target reverse attention weight of the target symbol.

[0070] It can be understood that the GRU network has stronger ability to retain long-term sequence information, and the second semantic features of the symbols before the target symbol are determined by using the GRU network, so that when the target symbol is recognized, the semantic information generated by the symbols before the target symbol can be retained, the target reverse attention weight of the target symbol is more accurate, thereby improving the prediction result of the target symbol; at the same time, when the target reverse attention weight of the target symbol is determined, the reverse attention weight of the symbol before the target symbol, i.e., the historical attention weight, is introduced, which can avoid the formula model from remembering the region that has been paid attention to when recognizing the target symbol, thereby preventing repeated decoding of the same region, and is more helpful for the generation of the target reverse attention weight of the target symbol; in this way, in the decoding and recognition process, the GRU and the attention mechanism are combined, which can better recognize longer and more complex formulas (especially mathematical formulas).

[0071] S203, determining a recognition result of the target symbol according to the target semantic features and the target attention weight.

[0072] Optionally, a forward fusion feature of the target symbol is determined according to the target semantic features and the target forward attention weight of the target symbol, and a forward recognition result of the target symbol is determined based on the forward fusion feature.

[0073] Optionally, a reverse fusion feature of the target symbol is determined according to the target semantic features and the target reverse attention weight of the target symbol, and a reverse recognition result of the target symbol is determined based on the reverse fusion feature.

[0074] S204, training the formula recognition model according to the recognition result, the target attention weight and the label data of the target symbol.

[0075] Specifically, the formula recognition model is trained according to the forward recognition result, the reverse recognition result, the label data, the target forward attention weight and the target reverse attention weight of the target symbol.

[0076] The technical scheme provided by the embodiments of the present disclosure is that the target semantic features of the target symbol are forward decoded to obtain target forward attention weights of the target symbol, and the target semantic features of the target symbol are backward decoded to obtain target backward attention weights of the target symbol, then the recognition result of the target symbol is determined according to the target semantic features and the target attention weights, and the formula recognition model is trained according to the recognition result, the target attention weights and the label data of the target symbol. The technical scheme can better recognize complex formulas by decoding the target symbol in two directions, i.e., from left to right and from right to left, and can improve the recognition accuracy of the formula recognition model by performing attention constraint on the formula recognition model.

[0077] Figure 3 is a flowchart of another method for training an image recognition model according to an embodiment of the present disclosure. The present embodiment further optimizes the training of the formula recognition model according to the recognition result, the target attention weights and the label data of the target symbol based on the above-mentioned embodiments, and provides an optional implementation scheme. As shown in Figure 3 The method for training an image recognition model according to the present embodiment can include the following steps:

[0078] S301, the target semantic features of the target symbol are forward decoded to obtain target forward attention weights of the target symbol.

[0079] S302, the target semantic features of the target symbol are backward decoded to obtain target backward attention weights of the target symbol.

[0080] S303, the recognition result of the target symbol is determined according to the target semantic features and the target attention weights.

[0081] Optionally, the forward fusion features of the target symbol are determined according to the target semantic features and the target forward attention weights of the target symbol, and the forward recognition result of the target symbol is determined based on the forward fusion features.

[0082] Optionally, the backward fusion features of the target symbol are determined according to the target semantic features and the target backward attention weights of the target symbol, and the backward recognition result of the target symbol is determined based on the backward fusion features.

[0083] S304, the attention loss is determined according to the target forward attention weights and the target backward attention weights.

[0084] In this embodiment, the attention loss refers to a loss determined based on the target forward attention weight and the target backward attention weight.

[0085] In an optional manner, the attention loss can be determined based on a preset loss function and the target forward attention weight and the target backward attention weight. The preset loss function is not limited in the present disclosure.

[0086] In another optional manner, divergence between the target forward attention weight and the target backward attention weight is determined, and the attention loss is determined based on the divergence.

[0087] Specifically, the KL divergence (Kullback-Leibler Divergence) between the target forward attention weight and the target backward attention weight can be calculated, and the KL divergence is taken as the attention loss.

[0088] It can be understood that although the target symbol is decoded in two directions, the attention weights of the target symbol, i.e., the symbols at the same position, should be the same. The introduction of the attention loss makes the target forward attention weight and the target backward attention weight of the target symbol as same as possible, i.e., attention distillation, so that the two decoding networks can learn from each other in the formula recognition model training process, thereby improving the performance of the formula recognition model.

[0089] S305, determining a recognition loss based on the recognition result and the label data of the target symbol.

[0090] Specifically, a first recognition loss can be determined based on the forward recognition result of the target symbol and the label data based on a preset loss function, and a second recognition loss can be determined based on the backward recognition result of the target symbol and the label data. Then, the recognition loss can be determined based on the first recognition loss and the second recognition loss based on a preset rule, for example, the first recognition loss and the second recognition loss can be added, and the addition result is taken as the recognition loss; or, the first recognition loss and the second recognition loss can be averaged, and the average result is taken as the recognition loss.

[0091] S306, determining a training loss based on the attention loss and the recognition loss.

[0092] Specifically, the training loss can be determined based on the attention loss and the recognition loss based on a preset rule, for example, the attention loss and the recognition loss can be added, and the addition result is taken as the training loss; or, the attention loss and the recognition loss can be averaged, and the average result is taken as the training loss.

[0093] S307, training the formula recognition model based on the training loss.

[0094] Specifically, the formula recognition model can be trained by using the training loss until a training stop condition is met.

[0095] The technical solution provided by the embodiments of the present disclosure is that the target forward attention weight of the target symbol is obtained by forward decoding the target semantic feature of the target symbol, the target reverse attention weight of the target symbol is obtained by reverse decoding the target semantic feature of the target symbol, then the attention loss is determined according to the target forward attention weight and the target reverse attention weight, the recognition loss is determined according to the recognition result and the label data of the target symbol, and then the training loss is determined according to the attention loss and the recognition loss, and the formula recognition model is trained according to the training loss. In the formula recognition model training process, the attention constraint is introduced, that is, the attention distillation is performed between the forward decoder and the reverse decoder, which can improve the performance of the formula recognition model.

[0096] Figure 4 is a flowchart of an image recognition method according to an embodiment of the present disclosure. The method is applicable to the case of how to recognize a formula image, and is particularly applicable to the case of how to recognize a handwritten mathematical formula. The method can be executed by an image recognition device, which can be implemented in software and / or hardware, and can be integrated into an electronic device that carries an image recognition function, such as a server. As shown in Figure 4 The image recognition method of the present embodiment can include:

[0097] S401, obtaining a formula image to be recognized.

[0098] In the present embodiment, the formula image to be recognized refers to an image that needs to be recognized, which can be a handwritten formula image (such as a mathematical formula) uploaded by image acquisition equipment, or a formula image obtained from the Internet.

[0099] S402, performing symbol recognition on the formula image to be recognized by using a formula recognition model to obtain a target recognition result of the formula to be recognized.

[0100] The formula recognition model is trained according to the training method of the image recognition model provided in any embodiment of the present disclosure. The target recognition result refers to the formula recognition result after recognizing the formula image to be recognized, which can be represented in a sequence form.

[0101] Specifically, the formula to be recognized image can be input into the formula recognition model, the formula recognition model is used for symbol recognition on the formula to be recognized image, a target forward recognition result and a target reverse recognition result are obtained, then one of the target forward recognition result and the target reverse recognition result is selected as the target recognition result of the formula to be recognized image. Preferably, the target forward recognition result is selected as the target recognition result of the formula to be recognized image.

[0102] The technical solution provided by the embodiments of the present disclosure comprises the following steps: obtaining a formula to be recognized image, then using a formula recognition model to perform symbol recognition on the formula to be recognized image, and obtaining a target recognition result of the formula to be recognized. The technical solution can improve the accuracy of formula recognition by using the formula recognition model to recognize the formula to be recognized image.

[0103] Figure 5 is a structural schematic diagram of a training device of an image recognition model according to the embodiments of the present disclosure. The embodiments are applicable to the case of how to recognize a formula image, and are particularly applicable to the case of how to recognize a handwritten mathematical formula. The device can be realized in the form of software and / or hardware, and can be integrated in an electronic device that carries the training function of the image recognition model, such as a server. As shown in Figure 5 The training device 500 of the image recognition model of the present embodiment can comprise:

[0104] A target weight determination module 501 is configured to determine a target attention weight of a target symbol according to a target semantic feature of the target symbol in a target formula image.

[0105] An identification result determination module 502 is configured to determine an identification result of the target symbol according to the target semantic feature and the target attention weight.

[0106] A model training module 503 is configured to train the formula recognition model according to the identification result, the target attention weight, and label data of the target symbol.

[0107] The technical solution provided by the embodiments of the present disclosure comprises the following steps: obtaining a formula to be recognized image, then using a formula recognition model to perform symbol recognition on the formula to be recognized image, and obtaining a target recognition result of the formula to be recognized. The technical solution can improve the accuracy of formula recognition by using the formula recognition model to recognize the formula to be recognized image.

[0108] Further, the target weight determination module 501 comprises:

[0109] a forward weight determination unit, configured to perform forward decoding on the target semantic feature of the target symbol to obtain a target forward attention weight of the target symbol;

[0110] a backward weight determination unit, configured to perform backward decoding on the target semantic feature of the target symbol to obtain a target backward attention weight of the target symbol.

[0111] Further, the model training module 503 comprises:

[0112] an attention loss determination unit, configured to determine an attention loss according to the target forward attention weight and the target backward attention weight;

[0113] a recognition loss determination unit, configured to determine a recognition loss according to the recognition result and label data of the target symbol;

[0114] a training loss determination unit, configured to determine a training loss according to the attention loss and the recognition loss;

[0115] a model training unit, configured to train the formula recognition model according to the training loss.

[0116] Further, the attention loss determination unit is specifically configured to:

[0117] determine divergence between the target forward attention weight and the target backward attention weight;

[0118] determine the attention loss according to the divergence.

[0119] Further, the forward weight determination unit is specifically configured to:

[0120] determine a first semantic feature of a first symbol before the target symbol in a forward decoding direction and a first forward attention weight;

[0121] determine a first fusion feature of the target symbol according to the target semantic feature of the target symbol and the first semantic feature of the first symbol and the first forward attention weight;

[0122] perform attention operation on the first fusion feature to obtain the target forward attention weight of the target symbol.

[0123] Further, the backward weight determination unit is specifically configured to:

[0124] determine a second semantic feature of a second symbol before the target symbol in a backward decoding direction and a second backward attention weight;

[0125] According to the target semantic feature of the target symbol, and the second semantic feature and the second reverse attention weight of the second symbol, a second fusion feature of the target symbol is determined;

[0126] The second fusion feature is subjected to attention operation to obtain a target reverse attention weight of the target symbol.

[0127] Figure 6 A structural schematic diagram of an image recognition device according to an embodiment of the present disclosure is provided. The embodiment is applicable to the case of how to recognize a formula image, and is particularly applicable to the case of how to recognize a handwritten mathematical formula. The device can be implemented in a software and / or hardware manner, and can be integrated in an electronic device that carries an image recognition function, such as a server. As shown in the figure, the image recognition device 600 of the embodiment can include: Figure 6

[0128] An image to be recognized acquisition module 601 is configured to acquire a formula image to be recognized;

[0129] A target recognition result determination module 602 is configured to perform symbol recognition on the formula image to be recognized by using a formula recognition model to obtain a target recognition result of the formula to be recognized; wherein the formula recognition model is trained according to the training method of the image recognition model provided in any embodiment of the present disclosure.

[0130] The technical solution provided in the embodiment of the present disclosure acquires a formula image to be recognized, and then performs symbol recognition on the formula image to be recognized by using a formula recognition model to obtain a target recognition result of the formula to be recognized. The above technical solution can improve the accuracy of formula recognition by using a formula recognition model to recognize the formula image to be recognized.

[0131] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0132] Figure 7 is a block diagram of an electronic device for implementing the training method of the image recognition model or the image recognition method of the embodiment of the present disclosure. Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit the implementations of the present disclosure described and / or claimed in this document. ​

[0133] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0134] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0135] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as image recognition model training methods or image recognition methods. For example, in some embodiments, the image recognition model training method or image recognition method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the image recognition model training method or image recognition method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform a training method for an image recognition model or an image recognition method.

[0136] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0137] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0138] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0139] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0140] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0141] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0142] Artificial intelligence is a discipline that studies enabling computers to simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning technology, big data processing technology, knowledge graph technology, etc. several major directions.

[0143] Cloud computing refers to a technology system that accesses an elastic scalable shared physical or virtual resource pool through a network, the resources can include servers, operating systems, networks, software, applications and storage devices, etc., and the resources can be deployed and managed in a demand self-service manner. Through cloud computing technology, efficient and powerful data processing capabilities can be provided for artificial intelligence, blockchain and other technical applications and model training.

[0144] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which is not limited herein.

[0145] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for training an image recognition model, comprising: The target attention weight of the target symbol is determined based on the target semantic features of the target symbol in the target formula image; The target attention weights include target positive attention weights and target negative attention weights; The recognition result of the target symbol is determined based on the target semantic features and the target attention weight; The attention loss is determined based on the target positive attention weight and the target negative attention weight; Based on the recognition results and the label data of the target symbol, determine the recognition loss; The training loss is determined based on the attention loss and the recognition loss. The formula recognition model is trained based on the training loss.

2. The method according to claim 1, wherein, The step of determining the target attention weight of the target symbol based on the target semantic features of the target symbol in the target formula image includes: The target semantic features of the target symbol are forward decoded to obtain the target positive attention weight of the target symbol; The target semantic features of the target symbol are reverse-decoded to obtain the target reverse attention weight of the target symbol.

3. The method according to claim 1, wherein, The step of determining the attention loss based on the target positive attention weight and the target negative attention weight includes: Determine the divergence between the target positive attention weights and the target negative attention weights; Based on the divergence, the attention loss is determined.

4. The method according to claim 2, wherein, The step of forward decoding the target semantic features of the target symbol to obtain the target positive attention weight of the target symbol includes: Determine the first semantic feature and the first positive attention weight of the first symbol preceding the target symbol in the forward decoding direction; Based on the target semantic features of the target symbol, and the first semantic features and first positive attention weight of the first symbol, the first fusion feature of the target symbol is determined; Attention operations are performed on the first fused feature to obtain the target positive attention weight of the target symbol.

5. The method according to claim 2, wherein, The step of reverse decoding the target semantic features of the target symbol to obtain the target reverse attention weight of the target symbol includes: Determine the second semantic features and second reverse attention weights of the second symbol located before the target symbol in the reverse decoding direction; Based on the target semantic features of the target symbol, the second semantic features of the second symbol, and the second reverse attention weight, the second fusion feature of the target symbol is determined; Attention operations are performed on the second fusion feature to obtain the target inverse attention weight of the target symbol.

6. An image recognition method, comprising: Obtain the image of the formula to be recognized; The formula recognition model is used to perform symbol recognition on the image of the formula to be recognized, and the target recognition result of the formula to be recognized is obtained; wherein, the formula recognition model is trained by the training method of the image recognition model according to any one of claims 1-5.

7. A training device for an image recognition model, comprising: The target weight determination module is used to determine the target attention weight of the target symbol based on the target semantic features of the target symbol in the target formula image; The target attention weights include target positive attention weights and target negative attention weights; The recognition result determination module is used to determine the recognition result of the target symbol based on the target semantic features and the target attention weight; The model training module includes: An attention loss determination unit is used to determine the attention loss based on the target positive attention weight and the target negative attention weight; A recognition loss determination unit is used to determine the recognition loss based on the recognition result and the label data of the target symbol; A training loss determination unit is used to determine the training loss based on the attention loss and the recognition loss; The model training unit is used to train the formula recognition model based on the training loss.

8. The apparatus according to claim 7, wherein, The target weight determination module includes: A positive weight determination unit is used to perform positive decoding on the target semantic features of the target symbol to obtain the target positive attention weight of the target symbol; The reverse weight determination unit is used to reverse decode the target semantic features of the target symbol to obtain the target reverse attention weight of the target symbol.

9. The apparatus according to claim 7, wherein, The attention loss determination unit is specifically used for: Determine the divergence between the target positive attention weights and the target negative attention weights; Based on the divergence, the attention loss is determined.

10. The apparatus according to claim 8, wherein, The positive weight determination unit is specifically used for: Determine the first semantic feature and the first positive attention weight of the first symbol preceding the target symbol in the forward decoding direction; Based on the target semantic features of the target symbol, and the first semantic features and first positive attention weight of the first symbol, the first fusion feature of the target symbol is determined; Attention operations are performed on the first fused feature to obtain the target positive attention weight of the target symbol.

11. The apparatus according to claim 8, wherein, The reverse weight determination unit is specifically used for: Determine the second semantic features and second reverse attention weights of the second symbol located before the target symbol in the reverse decoding direction; Based on the target semantic features of the target symbol, the second semantic features of the second symbol, and the second reverse attention weight, the second fusion feature of the target symbol is determined; Attention operations are performed on the second fusion feature to obtain the target inverse attention weight of the target symbol.

12. An image recognition device, comprising: The image acquisition module is used to acquire the image of the formula to be recognized; The target recognition result determination module is used to perform symbol recognition on the image of the formula to be recognized using a formula recognition model to obtain the target recognition result of the formula to be recognized; wherein, the formula recognition model is trained by the image recognition model training method according to any one of claims 1-5.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the training method of the image recognition model according to any one of claims 1-5, or the image recognition method according to claim 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the training method of the image recognition model according to any one of claims 1-5, or the image recognition method according to claim 6.

15. A computer program product comprising a computer program that, when executed by a processor, implements a training method for an image recognition model according to any one of claims 1-5, or an image recognition method according to claim 6.

Citation Information

Patent Citations

  • Image recognition model training method and device, image recognition method and device and electronic equipment

    CN113378833A

  • Text recognition method and device, equipment and medium

    CN113705313A