Transform-based handwritten chemical equation identification method and system

By introducing a covering attention and symbol counting mechanism in the Transformer model, the problem of insufficient robustness and accuracy in handwritten chemical equation recognition is solved, and a more efficient and interpretable handwritten chemical equation recognition system is achieved.

CN120148049APending Publication Date: 2025-06-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312340.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems of robustness, interpretability and insufficient cross-domain adaptability in identifying handwritten chemical equations, resulting in low recognition accuracy and efficiency.

Method used

An improved Transformer model based on the combination of coverage attention and symbol counting was designed. Through the global count weighting module and attention refinement module, the model's ability to recognize handwritten chemical equations is improved.

Benefits of technology

It realizes high accuracy recognition of handwritten chemical equations, improves the robustness and interpretability of the model, and is suitable for handwritten chemical formula recognition tasks in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148049A_ABST
    Figure CN120148049A_ABST
Patent Text Reader

Abstract

The invention discloses a handwritten chemical equation recognition method and system based on Transform, and the method comprises the steps: firstly carrying out the preprocessing of a handwritten chemical equation image, and constructing a handwritten chemical equation data set which can comprise 7395 samples of 100 symbol categories; dividing a data set into three parts, namely a training set, a verification set and a test set, which are used for training and testing the model provided by the invention; then obtaining a model prediction result, evaluating the prediction result, and verifying the accuracy of the model; and finally, applying the best model, namely parameters, to system identification, and performing real-time identification on an image input into a system. According to the method, the global counting weighting module and the attention refining module are respectively designed, the diversity of samples and the meticulousness of labels are contained, so that the high quality of data is ensured, an experiment carried out on the basis of a data set is more rigorous, and a solid foundation is provided for improving the robustness and accuracy of a recognition algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and natural language processing, and particularly to an improved Transformer handwritten chemical equation recognition system based on the combination of coverage attention and symbol counting. Background Art

[0002] In the context of the current digital transformation of education, computer online marking systems have significantly improved teaching efficiency by automatically processing objective questions, but their recognition ability for subjective questions (especially science handwritten formulas) still has obvious shortcomings. This technical bottleneck not only causes teachers to still need to invest a lot of time in manual marking, but also restricts the real-time nature and personalized development of teaching feedback. In a wider range of application scenarios, due to problems such as being difficult to modify, easy to damage, inconvenient to store, and high dissemination costs of paper archives, people convert paper archives into images by scanning, photographing, etc. and upload them to a computer, and then use optical character recognition technology to recognize the document images as electronic documents, thereby facilitating the editing and storage of paper documents. However, historical documents in the scenario of digitalizing archives often have problems such as fading and damage, and the professional symbol system in engineering drawings is significantly different from general mathematical symbols. This domain gap severely limits the generalization ability of existing models. In this context, how to construct a chemical handwritten formula recognition system with robustness, interpretability, and cross-domain adaptability has become the core topic in promoting the process of education intelligence and cultural heritage digitalization, which is the key scientific problem to be solved.

[0003] In the task of handwritten chemical formula recognition, it is divided into online recognition and offline recognition. In online recognition, based on the sensors of the input device, the stroke order of handwriting can be captured, and the corresponding coordinate parameters are input into the computer. Then, based on these data, recognition can be carried out better. In offline recognition, it is not restricted by the input method. The input is just an image without any additional information, and the corresponding formula sequence is obtained through image recognition. This method makes the recognition more difficult and the recognition result is only so-so. In addition, in terms of data labels, there is currently a lack of datasets that can directly support machine-readable formats, and this gap poses an obstacle to the development of intelligent marking systems. In the field of online education and intelligent marking, the realization of automatic recognition and evaluation of chemical equations and other chemical formulas depends on high-quality, structured datasets. Such datasets need to contain machine-readable annotations for handwritten chemical equations to enable the training algorithm to accurately understand and process complex chemical formulas. Given the limitations of existing resources, constructing a machine-readable dataset specifically for the chemical field has become a crucial step in promoting the development of intelligent education technology. This not only helps improve the accuracy and reliability of the automatic grading system but also accelerates the digitalization and automation process of chemical knowledge. Therefore, it is necessary to design an improved Transformer handwritten chemical equation recognition system based on the combination of coverage attention and symbol counting.

[0004] The encoder-decoder framework is an important model architecture for processing sequential data in the field of deep learning, and is widely used in fields such as natural language processing, speech recognition, and computer vision. Among them, the encoder-decoder framework for dealing with image-to-sequence problems combines the two major fields of computer vision and natural language processing, and is called a vision-language model. Its core idea is to convert the input image into an output sequence through two main components - the encoder and the decoder. The encoder is responsible for extracting feature representations from the input image, while the decoder generates corresponding sequence outputs based on these features. For the encoder part, the encoder based on convolutional neural network plays a crucial role in image-to-sequence tasks. Through its unique architecture design, CNN can effectively extract rich feature representations from the input image, which is crucial for subsequent task processing. The Transformer-based decoder is a powerful architecture for generating sequence outputs, and it has demonstrated excellent performance in many vision-language tasks. Different from traditional RNN- or LSTM-based decoders, the Transformer decoder relies entirely on self-attention mechanisms and feed-forward neural networks, which enables it to process input data in parallel and effectively solve the long-distance dependence problem. In chemical equation recognition, it can be divided into grammar-rule-based methods and deep-learning-based methods. In the first method, Yang et al. proposed a two-stage recognition algorithm, which uses structural information to distinguish different parts and substances, then segments them into isolated characters, recognizes the characters, and finally recombines each part into a structure with syntactic features. Wang et al. upgraded this algorithm and divided the structural information into three levels: the formula level, the molecular level, and the text level. Similarly, Cheng et al. proposed a framework consisting of three parts: symbol group, structure analysis, and semantic verification. In the second method, Kong et al. proposed to optimize on the basis of the CNN+RNN+CTC model, introduce prefix beam into CTC decoding and introduce two dictionaries during this process to help CTC improve performance. Wang et al. proposed a method that combines traditional segmentation methods with deep learning, compared the effects of CNN+RNN+CTC and CNN+RNN+Attention, and proved that the effect of the model based on CNN+RNN+Attention is better. Wang et al. used a lightweight CNN model on the basis of the CNN+RNN+CTC model to construct the LCRNN model, which greatly reduced the number of parameters during feature extraction and improved the model operation efficiency. Shen et al. used the CNN+GRU model to complete the recognition task, and the recognition effect on synthetic data is very good, but the performance on real datasets is unsatisfactory. Summary of the Invention

[0005] The objective of the present invention is to implement an improved Transformer handwritten chemical equation recognition system based on the combination of coverage attention and symbol counting, so as to convert a handwritten chemical equation image into a machine-readable LaTex string and solve the problems mentioned in the above background art.

[0006] To achieve the above objective, the present invention proposes an improved Transformer handwritten chemical equation recognition system based on the combination of coverage attention and symbol counting, and the solution is as follows:

[0007] First, preprocess the handwritten chemical equation image to construct a handwritten chemical equation dataset containing 7,395 samples with 100 symbol categories; then divide the dataset into three parts: a training set, a validation set, and a test set, which are used to train and test the model proposed by the present invention; then obtain the model prediction results and evaluate the prediction results to verify the accuracy of the model; finally, apply the best model, that is, the parameters, to the system recognition to perform real-time recognition on the images input into the system.

[0008] The specific steps of this technical solution are as follows:

[0009] Step S1: Design an encoder for feature extraction based on DenseNet of convolutional neural network to achieve feature extraction and position encoding for handwritten chemical equation images.

[0010] Step S2: Based on the feature extraction results of Step S1, design and construct a global counting weighted module based on the symbol counting mechanism. Figure 1 For the architecture of the global counting weighted module, use the multi-branch adaptive characteristics of the SKNet network to select weights of different convolutional sizes for feature enhancement, which helps the module better identify symbol features of different sizes, complex two-dimensional structures, and long character sequences in handwritten chemical equation images.

[0011] Step S3: Use the image features containing position information obtained in Step S1 as the input of the decoder end, and construct an attention refinement module based on coverage attention. By splitting the attention scores into optimized attention scores and coverage attention scores, the performance of the model is improved by making the decoder more evenly notice each part of the image, which helps the model identify more complex images.

[0012] Step S4: Use the encoder designed in Step S1, the global counting weighted module constructed in Step S2, and the decoder optimized by the attention refinement module obtained in Step S3 to jointly construct a handwritten chemical equation recognition model, which is the core component of the system. The architecture diagram is as Figure 2 shown.

[0013] Step S5: Use the model proposed in S4 to test and validate on the dataset to obtain the optimal parameter configuration. Verify the effectiveness and advancement of the two modules of the model by analyzing the results of counting vector visualization and attention score visualization, as Figure 3 , 4 shown.

[0014] The present invention has the following advantages:

[0015] 1. The present invention develops a brand-new handwritten chemical equation dataset containing inorganic and organic equations with 7395 samples, named HCED-IOC. This dataset not only covers handwritten samples with various styles, but also provides detailed LaTeX annotations to meet the machine-readable requirements. The diversity of the samples and the meticulousness of the annotations therein not only ensure the high quality of the data, but also make the experiments based on this dataset more rigorous, providing a solid foundation for improving the robustness and accuracy of the recognition algorithm.

[0016] 2. Aiming at the problem that the traditional CNN+Transformer recognition model has poor performance for handwritten chemical equations due to characters of different sizes, complex two-dimensional structures, and long label sequences, the present invention designs two modules, namely a global counting weighted module and an attention refinement module, to alleviate these problems. Based on this model, a handwritten chemical equation recognition system is designed. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the architecture diagram of the global counting weighted module.

[0018] Figure 2 is the architecture diagram of the handwritten chemical equation recognition model.

[0019] Figure 3 is the visualization display of the counting vector.

[0020] Figure 4 is the visualization display of the attention score.

[0021] Figure 5 is the display of the sample diversity of the dataset.

[0022] Figure 6 is the performance comparison of the model at various string lengths. DETAILED DESCRIPTION OF THE INVENTION

[0023] The following will clearly and completely describe the technical solutions and implementation details in the specific implementation process of the present invention in combination with the accompanying drawings of the present invention. Obviously, the described implementation examples are only a part of the implementation examples of the present invention, rather than all the embodiments; all other implementation examples obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0024] The implementation plan of this experiment is as follows:

[0025] Step S1: Construct a data set. HCER-IOC: The data set samples are collected in two ways. The first way is to directly download 2,200 pictures of homework or test papers containing multiple chemical equations uploaded by junior and senior high school students from online education websites, and then segment 4,743 samples of single-line chemical equations from the pictures; the second way is to select 213 common inorganic chemical equations and ionic equations, as well as 50 organic chemical equations as the chemical equation dictionary, and then invite multiple volunteers to randomly write different chemical equations in the canvas on the tablet computer with different thickness brushes using a stylus, and a total of 2,652 samples are collected. There are also a small number of clerical errors during the collection process. To ensure the diversity of the samples, this part of the samples is also retained. In summary, the sample data collected by the two methods totals 7,395. These images are divided into a training set, a validation set, and a test set in the ratio of 8:1:1.

[0026] Step S2: Preprocess the data set. Perform preprocessing operations on the 7,395 image samples collected, including equal-proportion scaling and binarization. The equal-proportion scaling operation sets the input image to a height not exceeding 150px, making the influence of each feature on the final prediction result more balanced. To be closer to the real usage scenario, the interference caused by different paper backgrounds in the picture is retained. The student handwritten pictures downloaded from the website have various characteristics such as multiple light and shadows, different colors, different writing tools, and complex backgrounds. Specific examples are Figure 3 shown. Directly performing binarization operations on these pictures results in a less-than-satisfactory segmentation effect. Therefore, adaptive binarization is adopted. This method can perform binarization according to the average gray value in different images, and this way can effectively handle the problem of poor segmentation effect.

[0027] Step S3: Construct a handwritten chemical equation recognition model based on Transformer, which mainly includes four parts: an encoder based on DenseNet, a global counting weighted module, a decoder based on coverage attention, and a result fusion prediction module.

[0028] (1) Encoder based on DenseNet

[0029] Feature extraction is performed through the DenseNet network. A DenseNet network contains multiple DenseBlocks. Assume a grayscale image is given. Feature maps are obtained through the DenseNet network. where d b represents the feature dimension of the Dense Block, and H and W represent the height and width of the feature map, and For the output x of the l-th layer l , it will receive all the outputs of the 0-th to l-1-th layers, and use the concatenation operation to connect them and then pass through the DenseBlock. The specific calculation formula is as follows:

[0030] x l = H l ([x 0 , x 1 , …, x l-1 )

[0031] where, H l (·) represents the internal operation of the Dense Block mentioned above.

[0032] In the encoder designed in this paper, DenseNet is stacked by three Dense Blocks. After DenseNet feature extraction, in order to align the feature map F with the overall model dimension d model , a convolutional layer with a kernel size of 1 is added at the end of the encoder to obtain the output image features After position encoding, it will be used by both the global counting weighted module and the Transformer Decoder. The global counting weighted module is used to predict the number of each symbol class and generate a one-dimensional counting vector representing the counting result. Then the image feature X O After combining with the position encoding again, it gets and will be input into the decoder together with the counting vector to obtain the predicted output result.

[0033] (2) Global counting weighted module

[0034] First, the feature map is subjected to feature extraction twice. This step includes three parts: convolution, batch normalization, and ReLU activation function. The kernel sizes used in the convolution part are 3 and 5 respectively. Different from the previous subsection, in order to further improve efficiency and reduce the number of parameters in this module, the traditional 5×5 kernel convolution is replaced by dilated convolution, that is, the kernel size is 3 and the dilation size is 2, to obtain SP, Next, in order to integrate information at different scales and enable the module to dynamically select the receptive field size in the subsequent operations, the results of these two branches are summed element-wise to obtain To capture global information and better understand the context information of the image, the statistics of each channel are obtained by performing global average pooling on FU Then, a fully connected layer is used to perform dimensionality reduction on Z to generate a more compact feature Through these steps, the basis for adaptive selection is obtained, enabling the module to flexibly focus on local or global information as needed. Then, based on the RE value, the soft attention weights SA and SA' of the two receptive fields can be obtained through softmax. It should be noted that SA + SA' = 1 here, so only one parameter needs to be calculated to obtain the value of the other parameter, further improving the computational efficiency of the module. After that, the feature information of the two scales is weighted and summed according to the adaptive selection weights to obtain the feature

[0035]

[0036] SA′ i = 1 - SA i

[0037] SE i = SA i ·SP

[0038] SE′ i = SA′ i ·SP′

[0039] Z i = SE i + SE′ i

[0040] Among them, represents the trainable parameter matrix, represents the i-th row of A and B, and SA i , SA' represent the i-th elements of SA and SB, represents the feature map of Z on the i-th channel. Therefore, After that, the feature is mapped to through a convolutional layer with a kernel size of 1×1, where C is the number of label symbol categories in the HCER-IOC dataset, that is, C = 100. Finally, each element value in the feature map is compressed to a distribution on (0, 1) through a sigmod function. For each it can reflect the position of the i-th symbol class to a certain extent. That is to say, each M iIn fact, they are all a density mapping:

[0041] M = σ(conv 1×1 (S))

[0042] Finally, the sum operator is used to sum all elements in the H and W dimensions to obtain the count vector on this branch

[0043] (3) Decoder based on coverage attention

[0044] Given c t , the coverage attention score can be directly generated by a coverage modeling function, decoupling the calculation of the coverage attention score from the original attention score calculation. Thus, the calculation order of the formula is adjusted, and the specific calculation method is as follows:

[0045] e t,i = v T tanh(W k K + W q q t + b attn + conv(c t ))

[0046] = v T tanh(W k K + W q q t + b attn ) + v T conv(c t )#

[0047] = v T tanh(W k K + W q q t + b attn ) + C t #

[0048] Among them, and are trainable parameter matrices, and b attn is a trainable parameter. is the extended hidden state sequence. is the sum of the attention scores a i along the image dimension T t , representing the degree of coverage obtained from the attention mechanism. conv(·) represents a 11×11 convolution operation, aiming to convert c t into a coverage matrix. C t ∈R LThat is the covered attention part. In this way, without affecting the parallelism of the Transformer, the attention distribution at the decoder end is adjusted, that is, the new attention weight D ∈ R T ×L×h can be obtained by subtracting the covered attention matrix C ∈ R T×L×h from the original attention matrix E ∈ R T×L×h to get:

[0049]

[0050] Among them, for the h-th attention head of the l-th layer, it can be expressed as the query key and value

[0051] (4) Fusion result prediction module

[0052] Through the above two modules, the counting vector and the output tensor of the Tranformer decoder are obtained respectively. First, the counting vector V is linearly transformed into Then, for each symbol j ∈ (0, l], the dropout regularization technique is applied to generate the word probability distribution at each position. For this purpose, a counting tensor is initialized for storage, and the specific formula is as follows:

[0053] W_P j = Dropout(V)

[0054] where Dropout(·) represents the regularization function to reduce the overfitting of the model.

[0055] Then the counting tensor W_P and the output tensor O td are combined by element-wise addition, and then normalized and linearly transformed to map the output result to the size of the vocabulary. Through the above operations, not only can the information from different sources be fully utilized, but also the generalization ability and prediction accuracy of the model are improved. The following are the specific steps and their corresponding formulas:

[0056] O = proj(norm(O td + W_P))

[0057] where norm(·) represents the normalization process, and proj(·) represents the linear transformation to map the features to the target space. The result O passes through the activation function sotfmax() to make the sum of all outputs equal to 1.

[0058] Step S4: Network training. In the recognition model of the present invention, the editor uses DenseNet to extract features from chemical equation images, where each dense layer contains 16 bottleneck layers, and each pair of dense layers contains a transition layer for sampling the feature image to the size while doubling the number of channels, with a growth rate k = 24 and a dropout rate = 0.3. In the Transformer decoder, the feature dimension d model = 256, the number of heads N h in multi-head attention = 8, and the hidden layer dimension d f in the feed-forward layer = 1024. For multi-scale feature extraction in GCWM, the convolutional kernel sizes of 3 and 5 are used, and the intermediate channel dimension d channel = 1024. During the training of this model, the stochastic gradient descent method (SGD) with a weight decay of 10-4 is adopted for training, the initial learning rate is set to 0.08, and dynamic learning rate adjustment is implemented. When the accuracy ExpRate does not improve for 10 consecutive times, the learning rate is reduced to half of the original, and the training batch size is 8. Scale augmentation

[67] is performed on the input images in the model, which helps the network adapt to different scales HCER where the uniform sampling scaling factor σ ∈ [0.7, 1.4]. In sequence prediction, beam search is used to maximize the prediction probability, the maximum sequence length is 200, and the beam width is 10. The recognition model is trained using the PtyTorch framework, and this model is trained on NVIDIA GeForce RTX 3090 with 24G of memory for each GPU.

[0059] Step S5: Model testing. To verify the effectiveness of the model of the present invention and the performance of the two modules, ablation experiments are carried out on the HCER-IOC dataset, which proves that whether the two modules act alone or synergistically on the model, they can bring performance improvements, and the effect is the best when they act together. In addition, by comparing with the models with the best results in other existing studies on HCER-IOC, it is also concluded that the model of the present invention has the best performance. To prove that ImHCER can be widely applied to the formula recognition task, the performance of the advanced model and the model of the present invention is also compared on other chemical formula datasets, chemical and mathematical mixed formula datasets, and the mathematical dataset CROHME. The superiority of the model of the present invention is shown through various data. Figure 6 This is the comparison of the accuracy of the model of the present invention with the benchmark model ABM, other advanced models CoMER, etc. in each string length interval, showing the advantages of the model in each length, especially when dealing with the recognition of ultra-long sequence images.

[0060] Step S6: System construction. The system construction includes three layers, namely the client, the server and the computing power platform. The client is built with the Vue framework, and the file system and network module are implemented with javaScript. It is mainly responsible for providing interactive pages for users and interacting with the server through network requests. The server is built with SpringBoot, which can make the server development more focused on business logic. The server mainly performs data processing, including pre-processing and post-processing operations on the data. In addition, the server is also responsible for providing interfaces for the client and dispatching the services provided by the computing power platform through network requests. Finally, the computing power platform is the handwritten chemical equation recognition model obtained according to the above steps, and the model is encapsulated twice with Flask, and the trip network service is provided to the server. Therefore, the computing power platform is mainly responsible for the deployment and operation of the deep learning model and provides services for the server. The specific process is to first upload the test paper picture through the front-end page, and then the system converts the picture into a base64 string and sends it to the back-end through an http request. The back-end parses the JSON data, restores the binary data of the picture, and adaptively binarizes the picture, and then sends it to the recognition model for recognition to obtain the LaTex string. Finally, the LaTex string is post-processed and fed back to the front-end page for presentation to the user. The entire process is designed to improve the accuracy and efficiency of chemistry test paper recognition and provide strong technical support for educational evaluation.

Claims

1. A Transformer-based handwritten chemical equation recognition method, characterized in that: Firstly, the handwritten chemical equation images are preprocessed to construct a handwritten chemical equation dataset containing 7395 samples of 100 symbol categories. Then the dataset is divided into three parts: training set, validation set and test set, which are used to train and test the handwritten chemical equation recognition model. Then the prediction results of the handwritten chemical equation recognition model are obtained, and the prediction results are evaluated to verify the accuracy of the model; finally, the best handwritten chemical equation recognition model, i.e., the parameters, are applied to the system recognition to perform real-time recognition of the images input into the system.

2. The Transformer-based handwritten chemical equation recognition method according to claim 1, characterized in that: The specific steps of this method are as follows: Step S1: design a feature extraction encoder based on the convolutional neural network DenseNet to achieve feature extraction and position encoding of the handwritten chemical equation image; Step S2: Based on the feature extraction results of step S1, a global counting weighted module based on the symbol counting mechanism is designed and constructed; the global counting weighted module uses the multi-branch adaptive characteristics of the SKNet network to select weights of different convolution sizes for feature enhancement, so as to better recognize symbol features of different sizes and complex two-dimensional structures as well as longer character sequences in the handwritten chemical equation images; Step S3: Using the image features containing position information obtained in step S1 as the input of the decoder, an attention refinement module is constructed based on coverage attention; by splitting the attention score into an optimized attention score and a coverage attention score, the performance of the model is improved by making the decoder pay more attention to every part of the image more evenly, helping the model to recognize more complex images; Step S4: construct a handwritten chemical equation book recognition model using the encoder designed in step S1, the global count weighted module constructed in step S2, and the decoder optimized by the attention refinement module obtained in step S3; Step S5: Use the model proposed in S4 to test and verify it on the dataset to obtain the optimal parameter configuration. The effectiveness and advancement of the two modules of the model are verified by analyzing the results of count vector visualization and attention score visualization.

3. The Transformer-based handwritten chemical equation recognition method according to claim 2, characterized in that: In step S1, the HCER-IOC: dataset samples were collected in two ways. The first way was to download online 2,200 pictures of homework or test papers containing multiple chemical equations uploaded by junior and senior high school students, and then segmented 4,743 samples of single-line chemical equations from the pictures; the second way was to select 213 inorganic chemical equations and ionic equations, and 50 organic chemical equations as the chemical equation dictionary, and then use a stylus to randomly write different chemical equations using brushes of different thicknesses on the canvas on the tablet, and collect a total of 2,652 samples; the sample data collected by the two methods totaled 7,395, and these images were divided into training set, validation set and test set in the form of 8:1:

1.

4. The Transformer-based handwritten chemical equation recognition method according to claim 2, characterized in that: In step S2, the collected 7395 image samples are preprocessed, including proportional scaling and binarization; the proportional scaling operation sets the input image to a height not exceeding 150px, so that the influence of each feature on the final prediction result is more balanced.

5. The Transformer-based handwritten chemical equation recognition method according to claim 2, characterized in that: In step S3, a Transformer-based Shouxie chemical equation recognition model is constructed, including: an encoder based on DenseNet, a global count weighting module, a decoder based on coverage attention, and a result fusion prediction module; (1) Encoder based on DenseNet; Feature extraction is performed through the DenseNet network. A DenseNet network contains multiple DenseBlocks. Given a grayscale image Get the feature map through the DenseNet network where d b represents the feature dimension of the DenseBlock block, H and W represent the height and width of the feature map, and For the output x of layer l l , will accept all the outputs from layer 0 to layer l-1, and connect them using a cascade operation and then pass them through the Dense Block block. The specific calculation formula is as follows: x l =H l ([x0,x1,…,x l-1 ]) Among them, H l (·) indicates the internal operation of Dense Block mentioned above; In the encoder, DenseNet consists of three Dense Blocks stacked together; after DenseNet feature extraction, in order to separate the feature map F from the overall model dimension d model Alignment, add a convolution layer with a kernel size of 1 at the end of the encoder to get the output image features After position encoding, it will be used by the global count weighting module and Transformer Decoder at the same time; the global count weighting module is used to predict the number of each symbol class and generate a one-dimensional count vector representing the count result; then the image feature X O Combined with the position encoding, we get The sum count vector is input into the decoder to get the predicted output result; (2) Global counting weighted module; First, the feature map Two feature extractions were performed, including convolution, batch normalization, and ReLU activation function. The kernel sizes used in the convolution part were 3 and 5 respectively. The 5×5 kernel convolution was replaced by a dilated convolution, that is, the kernel size was 3 and the dilation size was 2, and SP was obtained. Next, in order to integrate information of different scales and enable the module to dynamically select the receptive field size in the next operation, the results of these two branches are summed element by element to obtain In order to capture global information and better understand the contextual information of the image, the FU is pooled globally to obtain the statistics of each channel. Then use the fully connected layer to reduce the dimension of Z to generate a more compact feature Get the basis for adaptive selection; then according to the RE value, the soft attention weights SA and SA' of the two receptive fields can be obtained through softmax; according to the adaptive selection weight, the feature information of the two scales is weighted and summed to obtain the feature SA'i = 1-SA i THIS i =THIS i ·SP IT'S i =SA′ i ·SP′ FROM i =SE i +SE′ i in, represents the trainable parameter matrix, represents the i-th row of A and B, SA i ,SA' represents the i-th element of SA,SB, represents the feature map of Z on the i-th channel, After that, the features are transformed into Mapping Where C is the number of label symbol categories in the HCER-IOC dataset, that is, C = 100; Finally, a sigmoid function is used to compress each element value in the feature map to a distribution on (0, 1). To some extent, it reflects the position of the i-th symbol class. Each M i In fact, it is a density map: M=σ(conv 1×1 (S)) Finally, the sum pool operator is used to sum all elements in the H and W dimensions to obtain the count vector on this branch. (3) Decoder based on coverage attention; Given c t , the specific calculation method is as follows: e t,i =v T tanh(W k K+W q q t +b attn +conv(c t )) =v T tanh(W k K+W q q t +b attn )+v T conv(c t ) =v T tanh(W k K+W q q t +b attn )+C t # in, and is the trainable parameter matrix, b attn is a trainable parameter; is the expanded hidden state sequence; is the attention score a over all past decoder time steps i Along the image dimension T t The sum of , indicating the coverage obtained from the attention mechanism; conv(·) represents an 11×11 convolution operation, which converts c t Transformed into a covering matrix; C t ∈R L That is, to cover the attention part; without affecting the parallelism of Transformer, adjust the attention distribution on the decoder side, that is, the new attention weight D∈R T×L×h With the original attention matrix E∈R T ×L×h Subtract the coverage attention matrix C∈R T×L×h get: Among them, the h-th attention head of the l-th layer is represented as the query key Sum (4) Fusion result prediction module; Through the above two modules, we can get the count vectors And Tranformer decoder output tensor First, the count vector V is linearly transformed into Then, for each symbol j∈(0,l], the dropout regularization technique is applied to generate the word probability distribution at each position; for this purpose, a count tensor is initialized For storage, the specific formula is as follows: W_P j =Dropout(V) Dropout(·) represents the regularization function to reduce the overfitting of the model; Then the count tensor W_P is combined with the output tensor O td The output result is combined by element-by-element addition, normalization and linear transformation. Mapping to vocabulary size; specific steps and corresponding formulas: O=proj(norm(O td +W_P)) Among them, norm(·) represents normalization processing, proj(·) represents linear transformation to map the features to the target space; the output result O is activated by the sotfmax() function so that the sum of all outputs is 1.

6. The Transformer-based handwritten chemical equation recognition method according to claim 2, characterized in that: In step S4, the editor uses DenseNet to extract features from the chemical equation image, where each dense layer contains 16 bottleneck layers and each pair of dense layers contains a transition layer for sampling the feature image to size, and at the same time double the number of channels, growth rate k = 24, dropout rate = 0.3; in the Transformer decoder, the feature dimension d model = 256, the number of heads N in multi-head attention h =8, the number of hidden layers in the feedforward layer is d f = 1024; for multi-scale feature extraction in GCWM, the convolution kernel sizes used are 3 and 5, and the intermediate channel dimension d channel =1024; during the training process, the stochastic gradient descent method SGD with a weight decay of 10-4 is used for training.

7. The Transformer-based handwritten chemical equation recognition method according to any one of claims 1 to 6, characterized in that: The system construction for implementing this method includes a three-layer structure, namely the client, the server and the computing power platform. The client is built with the Vue framework, and the file system and network module are implemented with javaScript, which is responsible for providing interactive pages for users and interacting with the server through network requests; the server is built with SpringBoot, which can make the server development more focused on business logic. The server performs data processing, including pre-processing and post-processing operations on the data. In addition, the server is also responsible for providing interfaces for the client and dispatching the services provided by the computing power platform through network requests; finally, the computing power platform uses Flask to encapsulate the handwritten chemical equation recognition model obtained according to the above steps, and provides the network service to the server. The computing power platform is responsible for the deployment and operation of the deep learning model and provides services for the server; the specific process is to first upload the test paper picture through the front-end page, and then convert the picture into a base64 string and send it to the back-end through an http request. The back-end parses the JSON data, restores the binary data of the picture, and adaptively binarizes the picture, and then sends it to the recognition model for recognition to obtain the LaTex string; finally, the LaTex string is fed back to the front-end page after post-processing to present it to the user.