Mathematical expression recognition model construction method and device, mathematical expression recognition method and device and storage medium

By segmenting mathematical expressions and compiling them into index value sequences, the problem that LaTeX tag sequence cannot express spaces and dictionary expansion is solved, and higher model accuracy and computational efficiency are achieved.

CN120146047APending Publication Date: 2025-06-13TIANJIN ZHONGLIAN INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244502.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing method of splitting LaTeX tag sequences based on spaces cannot effectively express the necessary element of rendering spaces, and Byte Pair Encoding (BPE) encoding causes dictionary expansion, increasing the computational volume of the model.

Method used

By performing word segmentation on mathematical expressions, a dictionary is constructed and compiled into an index value sequence using the maximum forward matching algorithm, a LaTeX tag is generated, so that all elements in the formula are expressed accurately in model training.

Benefits of technology

It solves the problem that LaTeX tag sequence cannot express special characters (such as spaces), avoids dictionary expansion, reduces the calculation amount of the model, and improves the recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146047A_ABST
    Figure CN120146047A_ABST
Patent Text Reader

Abstract

The invention provides a mathematical expression recognition model construction method and device, a mathematical expression recognition method and device and a storage medium, and relates to the technical field of data processing.The mathematical expression recognition model construction method comprises the steps that a formula data set is obtained, word segmentation is conducted on each target mathematical expression, a dictionary is constructed based on the mapping relation between lexical elements and index values, and compiling rules are set; compiling the target mathematical expression into an index value sequence based on the dictionary and the compiling rule by using a maximum forward matching algorithm, generating at least one first target LaTex tag, generating a first target image based on the first target LaTex tag, and generating a second target image based on the second target LaTex tag. And inputting the first target LaTex tag and the first target image into a mathematical expression recognition model for training, thereby solving the technical problem that the LaTex tag sequence cannot express some special characters and dictionary expansion in the prior art, reducing the calculation amount of the model, and compared with an encoder-decoder of an attention mechanism, improving the recognition efficiency of the model. And the precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a method, device, and storage medium for constructing and recognizing a mathematical expression recognition model. Background Art

[0002] Mathematical expressions are important sources of information. Recognizing content from structural images is a highly challenging task because it not only requires recognizing content objects but also precisely extracting spatial feature information. In recent years, deep learning methods based on the encoder-decoder with attention mechanisms have become the mainstream recognition solutions. However, the existing method of splitting LaTeX tag sequences based on spaces can retain the semantic information of the original LaTeX characters but ignores that spaces themselves are also necessary elements for rendering. In other words, LaTeX tag sequences cannot express spaces. If Byte Pair Encoding (BPE) is used, it will lead to dictionary expansion and increase the computational complexity of the model. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method, device, and storage medium for constructing and recognizing a mathematical expression recognition model in order to solve the above problems.

[0004] In a first aspect, an embodiment of the present application provides a method for constructing a mathematical expression recognition model, which is applied to construct a mathematical expression recognition model. The method includes: Obtain a formula data set, where the formula data set includes at least one target mathematical expression; Perform word segmentation on each of the target mathematical expressions, and each of the target mathematical expressions is split into at least one token; Construct a dictionary based on the mapping relationship between tokens and index values, and set compilation rules. Use the maximum forward matching algorithm to compile the target mathematical expressions into a sequence of index values based on the dictionary and the compilation rules, and generate at least one first target LaTex tag. Among them, one target mathematical expression is compiled into one first target LaTex tag; Generate a first target image based on the first target LaTex tag, where one of the first target LaTex tags corresponds to one or more of the first target images; Input the first target LaTex tag and the first target image into a mathematical expression recognition model for training.

[0005] In the embodiments provided by the present application, by splitting each mathematical expression and compiling all the tokens based on a dictionary, all elements in the formula can be accurately expressed when generating LaTex tags, solving the technical problems in the prior art that LaTeX tag sequences cannot express some special characters (such as spaces) and dictionary inflation, reducing the computational amount of the model, and improving the accuracy of the model compared with the encoder-decoder of the attention mechanism.

[0006] One possible way is that each of the first target LaTex tags includes: a sequence of index values of the corresponding target mathematical expression; The mathematical expression recognition model is configured to: recognize the first target image and output a target sequence of index values, and recognize the formula corresponding to the target index value based on the target index value.

[0007] One possible way is that in the step of generating the first target image based on the first target LaTex tag, one first target LaTex tag corresponds to multiple first target images; In the step of generating the first target image based on the first target LaTex tag, the following method is used to generate multiple first target images corresponding to one first target LaTex tag in the first target LaTex tag: Perform data augmentation on the one first target LaTex tag to generate multiple tags generated for the one first target LaTex tag; Generate multiple first target images corresponding to the one first target LaTex tag based on the multiple tags generated for the one first target LaTex tag, where the multiple tags generated for the one first target LaTex tag correspond one-to-one to the multiple first target images corresponding to the first target LaTex tag.

[0008] One possible way is that the lengths of the index value sequences of each first target LaTex tag in the first target LaTex tags are the same.

[0009] One possible way is that the method further includes: Perform one-hot encoding on the first target LaTex tag to generate a one-hot encoded first target LaTex tag; Construct a cross-loss entropy function based on the one-hot encoded first target LaTex tag; Verify the mathematical expression recognition model based on the loss entropy function.

[0010] One possible way is that in the step of constructing the loss entropy function based on the one-hot encoded first target LaTex tag, the following formula is used to construct the loss entropy function: ; — Cross entropy loss function; — One-hot encoded vector of the -th category at the -th time step; — Softmax function of the -th category; where ; — Logical value of the -th category at the -th time step; — Parameter; — Number of tokens in the dictionary table.

[0011] In a second aspect, the present application provides a method for identifying a mathematical expression, including: Obtain a second target image, input the second target image into the mathematical expression recognition model described in the first aspect, and output a second target LaTex label; Identify the mathematical expression corresponding to the second target image based on the second target LaTex label.

[0012] In a third aspect, an embodiment of the present application provides a device for constructing a mathematical expression recognition model, which is applied to construct a mathematical expression recognition model, including: An acquisition module: used to acquire a formula data set, where the formula data set includes at least one target mathematical expression; An analysis module: used to tokenize each of the target mathematical expressions, and each of the target mathematical expressions is split into at least one token; A compilation module: used to construct a dictionary based on the mapping relationship between tokens and index values, set compilation rules, and use the maximum forward matching algorithm to compile the target mathematical expressions into a sequence of index values based on the dictionary and the compilation rules, generating at least one first target LaTex label, where one target mathematical expression is compiled to generate one first target LaTex label; A generation module: used to generate a first target image based on the first target LaTex label, where one first target LaTex label corresponds to one or more first target images; A training module: used to input the first target LaTex label and the first target image into the mathematical expression recognition model for training.

[0013] Fourth aspect, the present application provides an electronic device, including: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the methods as described in the first aspect or the second aspect by invoking the program instructions.

[0014] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, and the computer instructions cause the computer to execute the method as described in the first or second aspect.

[0015] Other features and advantages of the present invention will be described in the subsequent description, and in part will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, the claims and the drawings.

[0016] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the drawings, makes the following detailed description. Description of the Drawings

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 Flowchart of a method for constructing a mathematical expression recognition model provided by an embodiment of the present invention; Figure 2 Another flowchart of a method for constructing a mathematical expression recognition model provided by an embodiment of the present invention; Figure 3 Flowchart of a method for recognizing a mathematical expression provided by an embodiment of the present invention; Figure 4 Structural diagram of a device for recognizing a mathematical expression provided by an embodiment of the present invention; Figure 5 Structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0020] Currently, mathematical expressions are important sources of information. Identifying content from structural images is an extremely challenging task because it not only requires identifying content objects but also precisely extracting spatial feature information. In recent years, deep learning methods based on the encoder-decoder with attention mechanisms have become the mainstream recognition solutions. However, the existing method of splitting LaTeX tag sequences based on spaces can preserve the semantic information of the original LaTeX characters but ignores that spaces themselves are also necessary elements for rendering. In other words, the LaTeX tag sequence cannot express spaces. If Byte Pair Encoding (BPE) is used, it will cause dictionary expansion, thus affecting the accuracy of formula expression recognition.

[0021] Based on this, the embodiments of the present invention provide a method, device, and storage medium for constructing and recognizing a mathematical expression recognition model to solve the above problems.

[0022] To facilitate the understanding of this embodiment, the method for constructing the mathematical expression recognition model disclosed in the embodiments of the present invention will be introduced in detail first.

[0023] This application provides a method for constructing a mathematical expression recognition model. Referring to Figure 1 as shown, it is specifically applied to constructing a mathematical expression recognition model and specifically includes the following steps: S101: Obtain a formula data set, where the formula data set includes at least one target mathematical expression.

[0024] Specifically, in the embodiments provided in this application, IM2LATEX-100K can be used as the basic data set, and the basic data set is processed to generate a formula data set.

[0025] This formula data set includes at least one target mathematical expression. It should be noted that the target mathematical expression is the expression used for training the model.

[0026] S102: Segment each target mathematical expression, and each target mathematical expression is split into at least one token.

[0027] In this step, each target mathematical expression is split. In this embodiment, one token corresponds to one Token. A token can be a number or a letter in a mathematical expression, or can refer to a symbol in a mathematical expression, such as a radical sign (e.g., square root, cube root), a fractional structure, or can also refer to the position where a symbol is located, such as a superscript, a subscript, a numerator, a denominator, etc.

[0028] S103: Construct a dictionary based on the mapping relationship between tokens and index values, and set compilation rules. Use the maximum matching algorithm to compile the target mathematical expression into a sequence of index values based on the dictionary and the compilation rules, and generate at least one first target LaTex label.

[0029] The so-called maximum matching algorithm scans the main string from left to right and tries to match the longest pattern string. If the match fails, it retreats one character and continues to try a shorter pattern string.

[0030] Specifically, in this step, one target mathematical expression is compiled to generate one first target LaTex label. In the dictionary, each token has a corresponding index value. For each target mathematical expression, based on the dictionary, its corresponding tokens are compiled into corresponding index values according to the indexing principle, and thus the first target LaTex label is generated. It should be noted that this first target LaTex label corresponds to the label used for model training later. For each label in the first target label, it includes the index value of the corresponding target mathematical expression.

[0031] Table 1:

[0032] Referring to Table 1, Table 1 shows a partial example of the dictionary. In this example, the tokens are on the left and the index values are on the right. For the formula: , in this example, each token is compiled into the corresponding index value one by one. Specifically, \(\begin{array}\) can be used as the start of the matrix, and \(\end{array}\) can be used as the end of the matrix.

[0033] In this example, The corresponding token is \(A=\left(\begin{array}{ccc}a_{11}&0&0\\0&a_{22}&0\\0&0&a_{33}\end{array}\right)\). Based on this, the maximum matching algorithm is used to compile the above tokens, and the index values of each token are determined one by one, and thus the first target LaTex label is generated.

[0034] It should be noted that spaces can also be used as a token. Specifically, for the compilation rules, it is [SPA]. At the same time, to mark the start and end of a mathematical expression, when splitting tokens, a special token [BOS] is added at the start, a special token [EOS] is added at the end, and a special token [PAD] is added in the same training batch to make them of equal length.

[0035] S104: Generate a first target image based on the first target LaTex label.

[0036] In this step, after the first target LaTex label is generated, each first target LaTex label is converted into an image form. Specifically, one first target LaTex label corresponds to one or more first target images.

[0037] Furthermore, in some embodiments, one first target LaTex label generates one first target image, and one first target LaTex label generates multiple first target images.

[0038] It should be noted that the first target image specifically refers to the image used to train the mathematical expression recognition model.

[0039] S105: Input the first target LaTex label and the first target image into the mathematical expression recognition model for training.

[0040] In this step, the aforementioned first target LaTex label and the first target image are input into the mathematical expression recognition model, thereby constructing the mathematical expression recognition model.

[0041] In the embodiments provided in the present application, by splitting each mathematical expression and compiling all tokens based on a dictionary, all elements in the formula can be accurately expressed when generating LaTex labels, solving the technical problems of the LaTeX label sequence in the prior art being unable to express some special characters (such as spaces) and dictionary inflation, reducing the computational amount of the model, and improving the accuracy of the model compared with the encoder-decoder of the attention mechanism.

[0042] The following elaborates on the mathematical expression recognition model provided in the embodiments of the present application: In the embodiments of the present application, based on the fact that each of the aforementioned first target LaTex labels includes the index value corresponding to the target mathematical expression, the mathematical expression recognition model recognizes the target index value and recognizes the formula corresponding to the target index value based on the target index value.

[0043] It should be noted that the target index value is the index value output after the mathematical expression recognition model recognizes the first target image, and it is also the index value based on which the mathematical expression recognition model outputs the expression.

[0044] Specifically, for a certain picture, taking Table 1 as an example, assuming that for a certain image, the index value recognized by the mathematical expression recognition model is 133372, then the mathematical expression output by the mathematical expression recognition model is: , in this example There are three spaces on the left.

[0045] In the above way, the mathematical expression recognition model can output the mathematical expression.

[0046] As a preferred embodiment, the mathematical expression recognition model includes: an encoder sub-model and a decoder sub-model.

[0047] The encoder sub-model first obtains the first target image and outputs a corresponding three-dimensional high-level semantic feature map for each first target image based on the first target image.

[0048] Specifically, the encoder sub-model specifically includes a backbone network, an attention aggregation unit, and a spatial reorganization layer.

[0049] In this example, for the selection of the backbone network, any convolutional neural network (Convolutional Neural Networks, abbreviated as CNN) can be specifically used. For example, DenseNet121 can be selected as the backbone network to extract a corresponding four-dimensional high-level semantic feature map for each first target image, and then the attention aggregation unit is used to reduce the dimension of the four-dimensional high-level semantic feature map extracted by the aforementioned backbone network. Exemplarily, the four-dimensional high-level semantic feature map with 1024 channels can be reduced to 256 channels, and by reducing the dimension of the feature map, the complexity of subsequent calculations can be reduced.

[0050] After the dimension of the four-dimensional high-level semantic feature map is reduced, the spatial reorganization layer is used to convert the spatial reorganization layer into a three-dimensional feature map, thereby generating a corresponding three-dimensional high-level semantic feature map for each first target image.

[0051] After generating the corresponding three-dimensional high-level semantic feature map for each first target image, the corresponding three-dimensional high-level semantic feature map for each first target image and the aforementioned first target LaTex label are input into the decoder sub-model. The decoder obtains the corresponding three-dimensional high-level semantic feature map for each first target image and the aforementioned first target LaTex label, and recognizes the mathematical expression based on the corresponding three-dimensional high-level semantic feature map for each first target image and the first target LaTex label.

[0052] Specifically, in the embodiments provided in the application, by way of example, the decoder sub-model has 6 standard Transformer decoding layers. Each layer includes a multi-head self-attention mechanism, a cross-attention mechanism, and a feed-forward network. The multi-head attention mechanism and the cross-attention mechanism are specifically used to capture long-range dependencies in the LaTeX tag sequence. In this example, the first target LaTeX tag is input into the decoder sub-model. The decoder sub-model captures the dependencies in the first target LaTeX tag sequence and combines the corresponding three-dimensional high-level semantic feature maps of each first target image, and then can output a sequence of index values, that is, the target index value. After the target index value is output, the mathematical expression corresponding to the target index value can be output using a dictionary.

[0053] Meanwhile, in this embodiment, there is no processing for the lexical position information in the Transformer architecture. Therefore, it is necessary to add position encoding after Embedding to make up for the lack of position information. In this embodiment, preferably, an absolute position encoding scheme based on trigonometric functions is adopted. The position encoding formula for a given position pos is: ; In the formula, is the dimension of the model, is the dimension index of the position encoding vector.

[0054] Regarding the setting of the decoding sub-model, it should be noted that the above decoder sub-model is only one implementable way and does not limit the decoder sub-model. Those skilled in the art can set the specific structure of the decoder according to the specific functions of the decoder sub-model.

[0055] In some embodiments, one first target LaTeX tag corresponds to multiple first target images. In other words, for a certain first target LaTeX tag, it can be converted into multiple first target images. Then, in this embodiment, for the convenience of description, one of the multiple first target images is taken as an example for elaboration.

[0056] Referring to Figure 2 , specifically, in this embodiment, first, execute S401a: perform data augmentation on one first target LaTeX tag to generate multiple tags generated for one first target LaTeX tag.

[0057] For the data augmentation method, specifically, fonts can be used for augmentation. Taking 17892 as an example, specifically for 17892, the sequence of this index value can be augmented. Specifically, 17892 can be respectively converted into 7 fonts, namely "Latin ModernMath", "Asana Math", "XITS Math", "NotoSans Math","DejaVu Math TeX Gyre", "Cambria Math", "STIX Two Math". Thus, data augmentation is performed on 17892, generating multiple labels for a first target LaTex label.

[0058] It should be noted that in this example, the multiple labels generated for a first target LaTex label specifically refer to: multiple labels generated based on a first target LaTex label. Specifically, in dictionary compilation, the Times NewRoman font is specified, and 17892 is respectively converted into multiple fonts such as "Latin Modern Math", "Asana Math", "XITS Math", "Noto SansMath","DejaVu Math TeX Gyre", "Cambria Math", "STIX Two Math". Then the multiple labels generated for a first target LaTex label include all the labels generated after the transformation of 17892 into "Latin Modern Math", "Asana Math", "XITS Math", "NotoSans Math","DejaVu Math TeX Gyre", "CambriaMath", "STIX Two Math".

[0059] In this example, 17892 corresponds to a first target LaTex label.

[0060] Then, S401b will be executed: generating multiple target images corresponding to a first target LaTex label based on the multiple labels generated for a first target LaTex label.

[0061] Based on the foregoing example, for the font of 17892 being "Latin Modern Math", "AsanaMath", "XITS Math", "NotoSans Math", "DejaVu Math TeX Gyre", "Cambria Math", "STIX Two Math", 17892 is converted into a first target image corresponding to one of the foregoing first target LaTex tags. For the font of 17892 being Latin Modern Math, 17892 is converted into a first target image corresponding to one of the foregoing first target LaTex tags, and so on. All the tags generated after the font transformation of 17892 can be converted into 7 images. It should be noted that, at the same time, for these 7 images, the index value sequence is all 17892, and here they can all be regarded as the first target images corresponding to a first target LaTex tag. In view of the data augmentation of a first target LaTex tag, therefore, the number of first target images corresponding to a first target LaTex tag is also multiple.

[0062] Combined with the foregoing, it can be seen that the multiple tags generated for a first target LaTex tag correspond one-to-one with the multiple first target images corresponding to the first target LaTex tag.

[0063] In this embodiment, in the form of data augmentation, the data volume is expanded and the recognition performance of the expression recognition model is enhanced.

[0064] It should be noted that the mathematical expression recognition model can only recognize the index value sequence of the length. For each first target LaTex tag, the length of the index value sequence is the same. Specifically, the length of the index value sequence of each first target LaTex tag is a fixed value. For the index value sequence of a certain first target LaTex tag, if its length is less than the preset value, it can be filled with [PAD], thereby ensuring the consistency of the length of the index value sequence.

[0065] Based on the foregoing embodiment, in order to verify the obtained mathematical expression recognition model, based on the foregoing embodiment, first, the first target LaTex tag is one-hot encoded to generate the one-hot encoded first target LaTex tag, and then, based on the one-hot encoded first target LaTex tag, a cross-loss entropy function is constructed, and the mathematical expression recognition model is verified based on the loss entropy function.

[0066] Specifically, assume that the result of tokenizing a certain target mathematical expression is: , the th time step of the tag one-hot encoded vector, 。

[0067] In the prior art, the following method can be specifically adopted to calculate the softmax function of the th category: ; — The logical value of the th time step for the th category; — The number of tokens in the dictionary table; In order to more flexibly control the randomness and diversity of the generated sequence, as a preferred embodiment, the following formula can be specifically adopted to calculate the softmax function of the th category: ; — The logical value of the th time step for the th category; — Parameter; — The number of tokens in the dictionary table.

[0068] By introducing the parameter, when is less than 1, the distribution generated by the model tends to be more concentrated, and higher-probability words are selected, and the generated LaTeX sequence will be more certain; when is greater than 1, the distribution generated by the model tends to be more uniform, that is, the probability of high-probability words decreases, and the probability of low-probability words increases, thereby increasing the diversity of generation.

[0069] Preferably, selecting a smaller value for the parameter can make the model more conservative and more inclined to generate high-probability words or tokens, so that the generated LaTeX prediction sequence is more deterministic, so that the model can learn the potential unified rules.

[0070] Referring to Figure 3 , on the basis of the foregoing embodiment, the embodiment of the present application provides a method for recognizing a mathematical expression, specifically including: S201: Obtain a second target image, input the second target image into the mathematical expression recognition model, and output a second target LaTex label.

[0071] S202: Recognize the mathematical expression corresponding to the second target image based on the second target LaTex label.

[0072] Through the above method, the corresponding expression can be recognized by using the image.

[0073] Referring toFigure 4 , an embodiment of the present application provides a device for constructing a mathematical expression recognition model, which is applied to construct a mathematical expression recognition model and includes: An acquisition module: used to acquire a formula data set, where the formula data set includes at least one target mathematical expression; An analysis module: used to segment each target mathematical expression, and each target mathematical expression is split into at least one token; A compilation module: used to construct a dictionary based on the mapping relationship between tokens and index values, set compilation rules, and use the maximum forward matching algorithm to compile the target mathematical expression into a sequence of index values based on the dictionary and compilation rules, generating at least one first target LaTex label, where one target mathematical expression is compiled to generate one first target LaTex label; A generation module: used to generate a first target image based on the first target LaTex label, where one first target LaTex label corresponds to one or more first target images; A training module: used to input the first target LaTex label and the first target image into the mathematical expression recognition model for training.

[0074] For the device provided by the embodiment of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiment. For a brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.

[0075] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0077] Such as Figure 5As shown, the electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors 410, a memory 430, and a communication bus 440 that connects different system components (including the memory 430 and the processing unit 410).

[0078] The communication bus 440 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.

[0079] The electronic device typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.

[0080] The memory 430 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 430 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.

[0081] A program / utility having a set (at least one) of program modules may be stored in the memory 430. Such program modules include - but are not limited to - an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules generally perform the functions and / or methods described in the embodiments of the present invention.

[0082] The processor 410 executes various functional applications and data processing by running the programs stored in the memory 430, such as implementing the embodiments of the present inventionFigures 1 to 3 The method shown

[0083] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method shown in the embodiments of the present invention Figures 1 to 3 The method shown

[0084] The above computer-readable storage medium may be any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ReadOnly Memory; hereinafter referred to as: ROM), an erasable programmable read-only memory (Erasable Programmable ReadOnly Memory; hereinafter referred to as: EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device

[0085] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device

[0086] The program code contained on a computer-readable medium may be transmitted by any appropriate medium, including - but not limited to - wireless, wire, optical fiber, RF, etc., or any suitable combination of the above

[0087] The computer program code for performing the operations of the embodiments of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0088] The above describes specific embodiments of the embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0089] In the description of the embodiments of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In the embodiments of the present invention, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in the embodiments of the present invention and the features of different embodiments or examples.

[0090] Furthermore, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the embodiments of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0091] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the embodiments of the present invention includes additional implementations, where the functions can be performed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present invention pertain.

[0092] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0093] It should be noted that the terminals involved in the embodiments of the present invention may include, but are not limited to, personal computers (Personal Computer; hereinafter referred to as: PC), personal digital assistants (Personal Digital Assistant; hereinafter referred to as: PDA), wireless handheld devices, tablet computers, mobile phones, MP3 players, MP4 players, etc.

[0094] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0095] In addition, in each embodiment of the embodiments of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0096] The integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include a number of instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0097] The above are only the preferred embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall be included in the scope of protection of the embodiments of the present invention.

Claims

1. A method for constructing a mathematical expression recognition model, applied to constructing a mathematical expression recognition model, characterized in that: The method comprises: Acquire a formula data set, wherein the formula data set includes at least one target mathematical expression; Performing word segmentation on each of the target mathematical expressions, wherein each of the target mathematical expressions is split into at least one word unit; A dictionary is constructed based on the mapping relationship between word units and index values, and compilation rules are set. The target mathematical expression is compiled into an index value sequence based on the dictionary and the compilation rules using a maximum forward matching algorithm to generate at least one first target LaTex tag, wherein one target mathematical expression is compiled to generate one first target LaTex tag; Generate a first target image based on the first target LaTex tag, wherein one first target LaTex tag corresponds to one or more first target images; The first target LaTex tag and the first target image are input into a mathematical expression recognition model for training.

2. The method according to claim 1, characterized in that: Each of the first target LaTex tags includes: a sequence of index values ​​of the corresponding target mathematical expression; The mathematical expression recognition model is configured to: recognize the first target image, output a target index value sequence, and identify a formula corresponding to the target index value based on the target index value.

3. The method according to claim 2, characterized in that In the step of generating the first target image based on the first target LaTex tag, one first target LaTex tag corresponds to multiple first target images; In the step of generating the first target image based on the first target LaTex tag, a plurality of first target images corresponding to one first target LaTex tag in the first target LaTex tag are generated in the following manner: Performing data augmentation on the first target LaTex tag to generate multiple tags generated for the first target LaTex tag; Based on multiple tags generated for a first target LaTex tag, multiple first target images corresponding to the first target LaTex tag are generated, wherein the multiple tags generated for the first target LaTex tag correspond one-to-one to the multiple first target images corresponding to the first target LaTex tag.

4. The method according to claim 2, characterized in that: The length of the index value sequence of each first target LaTex tag in the first target LaTex tags is the same.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: One-hot encodes the first target LaTex tag to generate a one-hot encoded first target LaTex tag; Constructing a cross-loss entropy function based on the first target LaTex label after the one-hot encoding; The mathematical expression recognition model is verified based on the loss entropy function.

6. The method according to claim 2, characterized in that In the step of constructing a loss entropy function based on the first target LaTex tag after the one-hot encoding, the loss entropy function is constructed using the following formula: ; —Cross loss entropy function; —No. Time step One-hot encoded vector of categories; —No. The softmax function of categories; in, ; —No. Time step logical value of categories; -parameter; —The number of tokens in the dictionary table.

7. A mathematical expression recognition method, characterized in that: include: Acquire a second target image, input the second target image into a mathematical expression recognition model constructed by the method according to any one of claims 1 to 6, and output a second target LaTex tag; A mathematical expression corresponding to the second target image is identified based on the second target LaTex tag.

8. A mathematical expression recognition model construction device, used for constructing a mathematical expression recognition model, characterized in that: include: An acquisition module: used to acquire a formula data set, wherein the formula data set includes at least one target mathematical expression; Analysis module: used for performing word segmentation on each of the target mathematical expressions, wherein each of the target mathematical expressions is split into at least one word unit; Compilation module: used for constructing a dictionary based on the mapping relationship between word units and index values, setting compilation rules, and compiling the target mathematical expression into an index value sequence based on the dictionary and the compilation rules using a maximum forward matching algorithm to generate at least one first target LaTex tag, wherein one target mathematical expression is compiled to generate one first target LaTex tag; A generating module: used for generating a first target image based on the first target LaTex tag, wherein one first target LaTex tag corresponds to one or more first target images; Training module: used for inputting the first target LaTex tag and the first target image into the mathematical expression recognition model for training.

9. An electronic device, characterized in that: include: at least one processor; as well as at least one memory in communication with the processor, wherein: The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 6 by calling the program instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method according to any one of claims 1 to 6 or 7.