Handwritten mathematical formula identification method based on dual feature fusion

By adopting a dual feature fusion method in handwritten mathematical formula recognition, combining cross-stage CSP and Transformer feature fusion, and parallel visual attention module, the problem of low accuracy of similar character recognition is solved, and higher recognition accuracy and robustness are achieved.

CN120147799APending Publication Date: 2025-06-13LIAONING NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510219011.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing handwritten mathematical formula recognition method has low recognition accuracy when dealing with similar characters, especially due to the difficulty of identifying similar characters due to the differences in writing styles of different people.

Method used

The handwritten mathematical formula recognition method based on dual feature fusion is adopted, and feature maps of different scales are extracted through the feature extraction module. The dual feature fusion module performs cross-stage feature fusion and self-attention mechanism feature fusion. Combining the cross-stage CSP unit and the Transformer feature fusion module, the detailed information and association relationship of mathematical symbols are captured, and the features of the parallel visual attention module are weighted to fusion, and attention is dynamically allocated to improve recognition accuracy.

Benefits of technology

The accuracy and robustness of handwritten mathematical formula recognition are improved, and the mathematical symbols with similar shapes can be more accurately distinguished, and the categories of symbols are judged through the global context, solving the problem of low accuracy in the recognition of similar characters in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147799A_ABST
    Figure CN120147799A_ABST
Patent Text Reader

Abstract

The invention discloses a handwritten mathematical formula identification method based on dual feature fusion. The method comprises the following steps: S1, acquiring a processed handwritten mathematical formula image; s2, a handwritten mathematical formula recognition model is established and trained, and a trained handwritten mathematical formula recognition model is obtained; wherein the established handwritten mathematical formula recognition model comprises a feature extraction module; the dual feature fusion module is used for performing feature fusion on the feature maps of different scales according to a cross-stage feature fusion mode and a mode based on a self-attention mechanism; the parallel visual attention module is used for performing weighted fusion on the fused feature map to obtain an aligned visual feature map; a bidirectional mutual learning decoding module; and S3, carrying out handwritten mathematical formula identification based on the trained handwritten mathematical formula identification model. The method focuses on the details of handwritten mathematical symbols, helps to distinguish symbols with similar forms and capture the relationship between the symbols, dynamically distributes attention according to the relative importance between different characters, helps a model to more accurately capture tiny but crucial differences, and improves the accuracy of similar character recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of handwritten mathematical formulas, and in particular, to a handwritten mathematical formula recognition method based on dual feature fusion. Background Art

[0002] Existing handwritten mathematical expression recognition (HMER) methods can be divided into traditional handwritten mathematical formula recognition methods and deep learning-based methods. Among them, deep learning-based methods have improved the performance of traditional handwritten mathematical formula recognition methods, and the mainstream framework uses an encoder-decoder network. Encoder network: It converts the input mathematical formula image into a high-dimensional feature vector or feature sequence, and usually uses a convolutional neural network (CNN) or a recurrent neural network (RNN) to extract the feature information of the image. Decoder network: It receives the feature representation generated by the encoder and outputs the structured result of the mathematical formula through step-by-step decoding. Models such as recurrent neural network (RNN), gated recurrent unit (GRU), or Transformer can be used to implement it. Through the encoder-decoder network architecture, the mapping relationship from the image to the mathematical formula can be directly learned end-to-end, and more complex and abstract feature representations and rules can be learned from a large amount of data, so as to achieve more accurate and robust performance in the tasks of mathematical formula recognition and understanding.

[0003] However, due to different writing styles of different people, some characters may look similar in appearance. For example, a and α, 0 and o. There is a problem of low recognition accuracy for similar characters in handwritten mathematical formula recognition using the encoder-decoder architecture. Therefore, deep learning methods still face challenges when dealing with mathematical symbols with similar appearances. Summary of the Invention

[0004] The present invention provides a handwritten mathematical formula recognition method based on dual feature fusion to overcome the technical problem of low recognition accuracy when recognizing similar characters in the prior art.

[0005] To achieve the above object, the technical solution of the present invention is:

[0006] A handwritten mathematical formula recognition method based on dual feature fusion, characterized in that the specific steps include:

[0007] S1: Obtain a handwritten mathematical formula image, and preprocess the handwritten mathematical formula image to obtain a processed handwritten mathematical formula image;

[0008] S2: Establish a handwritten mathematical formula recognition model, and train the handwritten mathematical formula recognition model based on the processed handwritten mathematical formula image to obtain a trained handwritten mathematical formula recognition model;

[0009] Among them, the established handwritten mathematical formula recognition model includes:

[0010] A feature extraction module, which is used to extract feature maps of different scales of the processed handwritten mathematical formula image;

[0011] A dual feature fusion module, which is used to perform feature fusion on the feature maps of different scales respectively in a cross-stage feature fusion manner and a self-attention mechanism-based manner, and finally obtain a fused feature map, where the fused feature map contains the detailed information of the mathematical symbols in the handwritten mathematical formula and the correlation relationship between the mathematical symbols;

[0012] A parallel visual attention module, which is used to perform weighted fusion on the fused feature map to obtain an aligned visual feature map;

[0013] A bidirectional mutual learning decoding module, which is used to perform bidirectional decoding on the aligned visual feature map to obtain the recognition result of the handwritten mathematical formula;

[0014] S3: Perform handwritten mathematical formula recognition based on the trained handwritten mathematical formula recognition model.

[0015] Furthermore, the dual feature fusion module includes:

[0016] A cross-stage feature fusion module, which performs feature fusion on the feature maps of different scales in a cross-stage feature fusion manner to obtain a first fused feature map;

[0017] The cross-stage feature fusion module includes several cross-stage CSP units, and enhances several feature maps in the feature fusion process through the several cross-stage CSP units;

[0018] A Transformer feature fusion module, which performs feature fusion on the feature maps of different scales based on the self-attention mechanism to obtain a second fused feature map;

[0019] The Transformer feature fusion module includes a position encoding module and a multi-head attention module, and pays attention to the correlation relationship of several feature maps in the feature fusion process through the position encoding module and the multi-head attention module;

[0020] A splicing module, which is used to splice the first fused feature map and the second fused feature map to obtain a fused feature map.

[0021] Furthermore, the feature extraction module includes a first convolutional layer, a first dense block, a second convolutional layer, a first pooling layer, a second dense block, a third convolutional layer, a second pooling layer, a third dense block, a third pooling layer, and a fully connected layer connected in sequence;

[0022] The first convolutional layer is used to perform a convolutional operation on the processed handwritten mathematical formula image to obtain an initial feature map and transmit it to the first dense block;

[0023] The first dense block is used to perform a convolutional operation on the initial feature map to obtain a number of first intermediate convolutional features, splice the number of intermediate convolutional features and the initial feature map according to the channel dimension, and perform a convolutional operation on the splicing result to obtain a first feature map and transmit the first feature map to the second convolutional layer;

[0024] The second convolutional layer is used to perform a convolutional operation on the first feature map to obtain an intermediate feature map and transmit it to the first pooling layer;

[0025] The first pooling layer is used to perform a pooling operation on the intermediate feature map to obtain a first dimensionality-reduced feature map and transmit it to the second dense block;

[0026] The second dense block is used to perform a convolutional operation on the first dimensionality-reduced feature map to obtain a number of second intermediate convolutional features, splice the number of first intermediate convolutional features, the number of second intermediate convolutional features, the initial feature map, and the first feature map according to the channel dimension, and perform a convolutional operation on the splicing result to obtain a second feature map and transmit the second feature map to the third convolutional layer;

[0027] The third convolutional layer is used to perform a convolutional operation on the second feature map to obtain a high-level feature map and transmit it to the second pooling layer;

[0028] The second pooling layer is used to perform a pooling operation on the high-level feature map to obtain a second dimensionality-reduced feature map and transmit it to the third dense block;

[0029] The third dense block is used to perform a convolutional operation on the second dimensionality-reduced feature map to obtain a number of third intermediate convolutional features, splice the number of first intermediate convolutional features, the number of second intermediate convolutional features, the number of third intermediate convolutional features, the initial feature map, the first feature map, the first dimensionality-reduced feature map, and the second feature map according to the channel dimension, and perform a convolutional operation on the splicing result to obtain a third feature map and transmit the third feature map to the third pooling layer;

[0030] The third pooling layer is used to perform a pooling operation on the third feature map to obtain a third dimensionality-reduced feature map and transmit it to the fully connected layer;

[0031] The fully connected layer is used to perform a fully connected operation on the third dimensionality-reduced feature map to obtain a fourth feature map.

[0032] Further, the cross-stage feature fusion module includes a first upsampling unit, a first CSP unit, a second upsampling unit, a second CSP unit, a first downsampling unit, a third CSP unit, a second downsampling unit, and a fourth CSP unit connected in sequence;

[0033] The cross-stage feature fusion module performs feature fusion on the feature maps of different scales according to the cross-stage feature fusion method to obtain a first fusion feature map, including:

[0034] The first upsampling unit is used to upsample the fourth feature map, splice the upsampling result and the second feature map along the channel dimension to obtain a first spliced feature map, and transmit the first spliced feature map to the first CSP unit;

[0035] The first CSP unit is used to perform an enhancement operation on the first spliced feature map to obtain a fifth feature map and transmit it to the second upsampling unit;

[0036] The second upsampling unit is used to upsample the fifth feature map, splice the upsampling result and the first feature map along the channel dimension to obtain a second spliced feature map, and transmit the second spliced feature map to the second CSP unit;

[0037] The second CSP unit is used to perform an enhancement operation on the second spliced feature map to obtain a sixth feature map and transmit it to the first downsampling unit;

[0038] The first downsampling unit is used to downsample the sixth feature map, splice the downsampling result and the fifth feature map along the channel dimension to obtain a third spliced feature map, and transmit the third spliced feature map to the third CSP unit;

[0039] The third CSP unit is used to perform an enhancement operation on the third spliced feature map to obtain a seventh feature map and transmit it to the second downsampling unit;

[0040] The second downsampling unit is used to downsample the seventh feature map, splice the downsampling result and the fourth feature map along the channel dimension to obtain a fourth spliced feature map, and transmit the fourth spliced feature map to the fourth CSP unit;

[0041] The fourth CSP unit is used to perform an enhancement operation on the fourth spliced feature map to obtain a first fusion feature map.

[0042] Further, each of the first CSP unit, the second CSP unit, the third CSP unit, and the fourth CSP unit includes: a first convolutional block, a channel separation unit, a Bottleneck unit, and a second convolutional block;

[0043] The first convolutional block is used to perform a convolution operation on the input concatenated feature map, perform batch normalization on the convolution result, and perform a non-linear transformation on the normalized result using the SiLU activation function to obtain an activated feature map, and transmit the activated feature map to the channel separation unit;

[0044] The channel separation unit is used to perform channel separation processing on the activated feature map to equally divide the number of channels of the activated feature map into two parts, obtain two independent feature maps, randomly transmit one part of the independent feature maps directly to the second convolutional block, and transmit the other part of the independent feature maps to the Bottleneck unit;

[0045] The Bottleneck unit is used to perform several compression, expansion, and recompression operations on the input independent feature map to obtain a compressed feature, and transmit it to the second convolutional block;

[0046] The second convolutional block is used to perform a convolution operation on the input independent feature map and the compressed feature, perform batch normalization on the convolution result, and perform a non-linear transformation on the normalized result using the SiLU activation function to obtain an enhanced feature map.

[0047] Further, the Transformer feature fusion module performs feature fusion on the feature maps of different scales based on the self-attention mechanism to obtain a second fused feature map, including:

[0048] Flatten the first feature map, the second feature map, and the fourth feature map into one-dimensional feature maps respectively through a flattening operation;

[0049] Unify the scales of the one-dimensional feature maps through a linear projection operation;

[0050] Concatenate the three one-dimensional feature maps with unified dimensions into a fifth concatenated feature map;

[0051] Perform position encoding on the fifth concatenated feature map through the position encoding module, and add the position-encoded fifth concatenated feature map to the original fifth concatenated feature map pixel by pixel to obtain an eighth feature map;

[0052] Process the eighth feature map through a multi-head attention module, including:

[0053] Calculate the query vector Q i corresponding to the element x i in the eighth feature map, the key vector Ki Sum vector V i , the mathematical expression is:

[0054]

[0055] where i = 1, 2, 3…, n; W Q is the query vector weight matrix; W K is the key vector weight matrix; W V is the value vector weight matrix;

[0056] Calculate the attention score for a pair of elements x i and x j , the mathematical expression is:

[0057]

[0058] where d k is the dimension of the key vector; K j is the key vector K j of the element x j ;

[0059] Convert the attention score value to a value between 0 and 1 and sum to 1 through the softmax function to obtain the attention weight ω ij , the mathematical expression is:

[0060] ω ij = softmax(score(Q i , K j )) (3)

[0061] Multiply the value vector V i of the element x i by its corresponding attention weight ω ij and sum to obtain the attention output z i , the mathematical expression is:

[0062]

[0063] Concatenate all the attention outputs to obtain the ninth feature map, the mathematical expression is:

[0064] Z = W O ·Concat(z 1 , z 2 ,..., z h ) (5)

[0065] where W O is the output weight matrix, and Concat is the concatenation operation;

[0066] Perform a fully connected operation on the ninth feature map to obtain a tenth feature map;

[0067] Add the tenth feature map and the eighth feature map pixel by pixel, and concatenate the addition result with the query vector Q i for i = 1, 2, 3…, n to obtain a second fused feature map.

[0068] Furthermore, the parallel visual attention module performs weighted fusion on the fused feature map to obtain an aligned visual feature map, including:

[0069] 1) Input the fused feature map into a parallel attention mechanism for processing to generate an attention map, including:

[0070] The mathematical expression of the parallel attention mechanism is:

[0071]

[0072] where e t,ij is the attention weight at position ij at time step t; W e , W o and W v are trainable weights, O t is the character reading order, and O t = 0, 1, 2,..., N - 1, N is the number of characters, f o is an embedding function; v ij is the visual feature at position ij in the fused feature map;

[0073] 2) Multiply the attention map and the fused feature map pixel by pixel to generate an aligned visual feature map g t , g t ∈R d , and the mathematical expression is:

[0074]

[0075] where g t is the aligned visual feature map at time step t;

[0076] Finally, the parallel visual attention module outputs the aligned visual feature maps G of all time steps in parallel.

[0077] Beneficial effects: By establishing a handwritten mathematical formula recognition model and introducing a dual feature fusion module, the present invention performs feature fusion on the feature maps of different scales, focusing on the details of handwritten mathematical symbols, such as the starting position, curvature, and thickness of strokes, helping to distinguish morphologically similar mathematical symbols and capture the relationships between mathematical symbols. At the same time, it recognizes the context and role of mathematical symbols in handwritten mathematical formulas, helping to distinguish the meanings of similar mathematical symbols in different contexts, enabling the handwritten mathematical formula recognition model to distinguish mathematical symbols from local details and determine the category of symbols through global context. By introducing a parallel visual attention module, it dynamically allocates attention according to the relative importance between different characters (such as local features or context information), preferentially focusing on those detailed features crucial for character recognition, helping the model to more accurately capture those tiny but crucial differences, and solving the problem that existing handwritten mathematical formula recognition methods have low recognition accuracy for similar characters due to the high structural similarity of some characters themselves, especially the tiny differences in strokes, shapes, or positions, and the differences in writing styles of different individuals (such as the thickness, spacing, and curvature of strokes) in the recognition process of handwritten mathematical formulas, resulting in very similar appearances of some characters. Brief Description of the Drawings

[0078] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following briefly introduces the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0079] Figure 1 It is a flowchart of a handwritten mathematical formula recognition method based on dual feature fusion in the present invention;

[0080] Figure 2 It is a schematic diagram of the result of the image rotation operation in the embodiment of the present invention;

[0081] Figure 3 It is a schematic diagram of the result of the image flipping operation in the embodiment of the present invention;

[0082] Figure 4 It is a schematic diagram of the structure of the feature extraction module in the embodiment of the present invention;

[0083] Figure 5 It is a schematic diagram of the structure of the dual feature fusion module in the embodiment of the present invention;

[0084] Figure 6 It is a schematic diagram of the structure of the cross-stage feature fusion module in the embodiment of the present invention;

[0085] Figure 7 It is a schematic structural diagram of the CSP unit in the embodiment of the present invention;

[0086] Figure 8 It is a schematic structural diagram of the Transformer feature fusion module in the embodiment of the present invention;

[0087] Figure 9 It is a schematic structural diagram of the parallel visual attention module in the embodiment of the present invention;

[0088] Figure 10 It is a schematic structural diagram of the bidirectional mutual learning decoding module in the embodiment of the present invention. Detailed implementation manners

[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0090] This embodiment provides a handwritten mathematical formula recognition method based on dual feature fusion, as Figure 1 shown, and the specific steps include:

[0091] S1: Obtain a handwritten mathematical formula image, and preprocess the handwritten mathematical formula image to obtain a processed handwritten mathematical formula image;

[0092] In this embodiment, in order to improve the generalization ability of the handwritten mathematical formula recognition model and enhance its robustness, data augmentation preprocessing is performed on the handwritten mathematical formula image, and the following method is adopted:

[0093] (1) Read the image I in sequence, and rotate the image I by 90 degrees, 180 degrees, and 270 degrees respectively to obtain the corresponding rotated images I 1 , I 2 and I 3 , as Figure 2 shown;

[0094] (2) Flip the image I horizontally to obtain the flipped image I 4 , as Figure 3 shown;

[0095] Specifically, in this embodiment, the handwritten mathematical formula image dataset is augmented by rotation, flipping, etc., and together with the original image I before processing, the number of the dataset is expanded, enabling the model to better adapt to handwritten mathematical formulas of various angles, postures, and styles, thereby enhancing its robustness and the effect in practical applications.

[0096] S2: Establish a handwritten mathematical formula recognition model, and train the handwritten mathematical formula recognition model based on the processed handwritten mathematical formula images to obtain a trained handwritten mathematical formula recognition model;

[0097] Among them, the established handwritten mathematical formula recognition model includes a feature extraction module, a dual feature fusion module, a parallel visual attention module, and a bidirectional mutual learning decoding module;

[0098] Among them, the feature extraction module is used to extract feature maps of different scales of the processed handwritten mathematical formula images;

[0099] Specifically, the feature extraction module adopts a DenseNet (Dense Connection Network). The Dense Connection Network is a deep convolutional neural network. By introducing dense connections in the network to enhance feature reuse and gradient flow, it can better extract multi-level and multi-scale features and is more suitable for processing images with complex structures, showing excellent performance in model accuracy and generalization ability compared with traditional neural networks.

[0100] Specifically, in this embodiment, as Figure 4 shown, the feature extraction module includes a first convolutional layer, a first Dense Block, a second convolutional layer, a first pooling layer, a second Dense Block, a third convolutional layer, a second pooling layer, a third Dense Block, a third pooling layer, and a fully connected layer connected in sequence;

[0101] The first convolutional layer is used to perform a convolution operation on the processed handwritten mathematical formula image to obtain an initial feature map and transmit it to the first Dense Block; specifically, the initial feature map contains primary features such as edges, textures, and simple shapes (such as straight lines, curves, etc.).

[0102] The first Dense Block is used to perform a convolution operation on the initial feature map to obtain a number of first intermediate convolutional features, splice the number of intermediate convolutional features and the initial feature map according to the channel dimension, perform a convolution operation on the splicing result to obtain a first feature map, and transmit the first feature map to the second convolutional layer;

[0103] Specifically, the first feature map A contains low-level local features (such as edges, textures, etc.) in the image, laying a foundation for subsequent extraction of more complex symbols, numbers, letters, and formula structures.

[0104] The second convolutional layer is used to perform a convolutional operation on the first feature map to obtain an intermediate feature map and transmit it to the first pooling layer;

[0105] The first pooling layer is used to perform a pooling operation on the intermediate feature map to obtain a first dimensionality-reduced feature map and transmit it to the second dense block;

[0106] The second dense block is used to perform a convolutional operation on the first dimensionality-reduced feature map to obtain a number of second intermediate convolutional features, concatenate a number of first intermediate convolutional features, a number of second intermediate convolutional features, the initial feature map, and the first feature map along the channel dimension, perform a convolutional operation on the concatenated result to obtain a second feature map, and transmit the second feature map to the third convolutional layer;

[0107] Specifically, the second feature map B contains high-level features such as the spatial relationship between symbols and complex combinations of mathematical symbols (such as fraction lines and radical symbols).

[0108] The third convolutional layer is used to perform a convolutional operation on the second feature map B to obtain a high-level feature map and transmit it to the second pooling layer;

[0109] The second pooling layer is used to perform a pooling operation on the high-level feature map to obtain a second dimensionality-reduced feature map and transmit it to the third dense block;

[0110] The third dense block is used to perform a convolutional operation on the second dimensionality-reduced feature map to obtain a number of third intermediate convolutional features, concatenate a number of first intermediate convolutional features, a number of second intermediate convolutional features, a number of third intermediate convolutional features, the initial feature map, the first feature map, the first dimensionality-reduced feature map, and the second feature map along the channel dimension, perform a convolutional operation on the concatenated result to obtain a third feature map, and transmit the third feature map to the third pooling layer;

[0111] The third pooling layer is used to perform a pooling operation on the third feature map to obtain a third dimensionality-reduced feature map and transmit it to the fully connected layer;

[0112] The fully connected layer is used to perform a fully connected operation on the third dimensionality-reduced feature map to obtain a fourth feature map C.

[0113] Specifically, in this embodiment, the processed handwritten mathematical formula image will obtain three feature maps with different scales after passing through the feature extraction module: the first feature map A (reduced by 8 times compared to the original size), the second feature map B (reduced by 16 times compared to the original size), and the fourth feature map C (reduced by 32 times compared to the original size).

[0114] Specifically, the dual feature fusion module is used to perform feature fusion on the feature maps of different scales respectively in a cross-stage feature fusion manner and a self-attention mechanism-based manner, and finally obtain a fused feature map, where the fused feature map contains the detailed information of the mathematical symbols in the handwritten mathematical formula and the correlation relationship between the mathematical symbols;

[0115] Specifically, as Figure 5 shown, the dual feature fusion consists of cross-stage partial CSP (Cross Stage Partial)-based feature fusion and Transformer-encoded feature fusion. The cross-stage feature fusion module and the Transformer feature fusion module respectively fuse the three feature maps of different scales output by the feature extraction, and the splicing module adds the feature maps output by the two fusion modules to obtain the final fused feature map F.

[0116] Specifically, in this embodiment, the dual feature fusion module includes:

[0117] A cross-stage feature fusion module that performs feature fusion on the feature maps of different scales in a cross-stage feature fusion manner to obtain a first fused feature map;

[0118] The cross-stage feature fusion module includes a number of cross-stage CSP units that enhance a number of feature maps during the feature fusion process;

[0119] Specifically, in this embodiment, as Figure 6 shown, the cross-stage feature fusion module includes a first upsampling unit, a first CSP unit, a second upsampling unit, a second CSP unit, a first downsampling unit, a third CSP unit, a second downsampling unit, and a fourth CSP unit connected in sequence;

[0120] The cross-stage feature fusion module performing feature fusion on the feature maps of different scales in a cross-stage feature fusion manner to obtain a first fused feature map includes:

[0121] The first upsampling unit is used to upsample the fourth feature map C, splice the upsampling result with the second feature map B in the channel dimension to obtain a first spliced feature map, and transmit the first spliced feature map to the first CSP unit;

[0122] The first CSP unit is used to perform an enhancement operation on the first spliced feature map to obtain a fifth feature map B 1 , and transmit it to the second upsampling unit;

[0123] The second upsampling unit is used to upsample the fifth feature map B 1Perform upsampling, concatenate the upsampling result with the first feature map A along the channel dimension to obtain a second concatenated feature map, and transmit the second concatenated feature map to the second CSP unit;

[0124] The second CSP unit is used to perform an enhancement operation on the second concatenated feature map to obtain a sixth feature map A 1 , and transmit it to the first downsampling unit;

[0125] The first downsampling unit is used to 1 perform downsampling on the sixth feature map A 1 and concatenate the downsampling result with the fifth feature map B along the channel dimension to obtain a third concatenated feature map, and transmit the third concatenated feature map to the third CSP unit;

[0126] The third CSP unit is used to perform an enhancement operation on the third concatenated feature map to obtain a seventh feature map C 1 , and transmit it to the second downsampling unit;

[0127] The second downsampling unit is used to 1 perform downsampling on the seventh feature map C

[0128] and concatenate the downsampling result with the fourth feature map C along the channel dimension to obtain a fourth concatenated feature map, and transmit the fourth concatenated feature map to the fourth CSP unit; 1 .

[0129] Specifically, in this embodiment, as Figure 7 shown, the first CSP unit, the second CSP unit, the third CSP unit, and the fourth CSP unit all include: a first convolutional block (Conv2D_BN_SiLU), a channel separation unit, a Bottleneck unit, and a second convolutional block;

[0130] The first convolutional block is used to perform a convolutional operation on the input concatenated feature map, perform batch normalization on the convolutional result, and perform a non-linear transformation on the normalized result using the SiLU activation function to obtain an activated feature map, and transmit the activated feature map to the channel separation unit;

[0131] The channel separation unit is used to perform channel separation processing on the activation feature map, so as to equally divide the number of channels of the activation feature map into two parts, obtaining two independent feature maps. Randomly, one part of the independent feature maps is directly transmitted to the second convolutional block, and the other part of the independent feature maps is transmitted to the Bottleneck unit; the Bottleneck unit performs several compression, expansion, and recompression operations on the input independent feature maps to obtain compressed features, and the compressed features are high-dimensional features and are transmitted to the second convolutional block;

[0132] Specifically, in this embodiment, the Bottleneck unit includes three Bottleneck structures. Therefore, the Bottleneck unit performs three compression, expansion, and recompression operations on the input independent feature maps in sequence.

[0133] The second convolutional block is used to perform a convolutional operation on the input independent feature maps and the compressed features, perform batch normalization processing on the convolutional result, and use the SiLU activation function to perform a non-linear transformation on the normalized result to obtain an enhanced feature map.

[0134] Specifically, both the first convolutional block and the second convolutional block include: a fourth convolutional layer (Conv2D), a batch normalization layer (Batch Normalization), and a SiLU activation function layer.

[0135] Specifically, the Transformer feature fusion module performs feature fusion on the feature maps of different scales based on the self-attention mechanism to obtain a second fused feature map;

[0136] The Transformer feature fusion module includes a position encoding module and a multi-head attention module, and pays attention to the correlation relationships of several feature maps during the feature fusion process through the position encoding module and the multi-head attention module;

[0137] Specifically, in this embodiment, as Figure 8 shown, the Transformer feature fusion module includes: a position encoding module (Position Encoding) and a multi-head attention module (Multi-Head Attention);

[0138] Specifically, the position encoding module enables the handwritten mathematical formula recognition model to understand the positions of individual characters in the input sequence, thereby capturing the order and structural information in the sequence; each attention head in the multi-head attention module can independently learn different attention weights, enabling the handwritten mathematical formula recognition model to simultaneously focus on information from different positions, so as to more meticulously capture the connections and differences in the sequence, fuse the visual information and semantic information of the picture, and have a better effect in processing irregular texts such as occlusion and blur.

[0139] Specifically, the Transformer feature fusion module performs feature fusion on the feature maps of different scales based on the self-attention mechanism to obtain the second fused feature map, including:

[0140] Flatten the first feature map A, the second feature map B, and the fourth feature map C into one-dimensional feature maps respectively through a flattening operation;

[0141] Unify the scales of the one-dimensional feature maps through a linear projection operation;

[0142] Concatenate the three one-dimensional feature maps with unified dimensions into a fifth concatenated feature map;

[0143] Perform position encoding on the fifth concatenated feature map through the position encoding module, and add the position-encoded fifth concatenated feature map to the original fifth concatenated feature map pixel by pixel to obtain the eighth feature map D, denoted as X = [x 1 , x 2 ,..., x n ;

[0144] Process the eighth feature map through the multi-head attention module, including:

[0145] Calculate the query vector Q i corresponding to the element x i , key vector K i , and value vector V i in the eighth feature map. The mathematical expression is:

[0146]

[0147] where i = 1, 2, 3,..., n; W Q is the query vector weight matrix; W K is the key vector weight matrix; W V is the value vector weight matrix;

[0148] Based on the scaled dot-product attention mechanism, calculate the attention score of a pair of elements x i and x j . The mathematical expression is:

[0149]

[0150] Among them, d k is the dimension of the key vector; K j is the key vector K j of the element x j ;

[0151] The attention score value is converted to a value between 0 and 1 and the sum is 1 through the softmax function to obtain the attention weight ω ij , and the mathematical expression is:

[0152] ω ij = softmax(score(Q i , K j )) (3)

[0153] Multiply the value vector V i of the element x i by its corresponding attention weight ω ij and sum them to obtain the attention output z i , and the mathematical expression is:

[0154]

[0155] The attention output z i not only contains the information of the current element but also contains the information of its context;

[0156] In the multi-head attention mechanism, the input data is divided into multiple "heads", and each "head" independently calculates the self-attention to obtain its own attention output;

[0157] Therefore, all the attention outputs are concatenated to obtain the ninth feature map, and the mathematical expression is:

[0158] Z = W O ·Concat(z 1 , z 2 ,..., z h ) (5)

[0159] Among them, W O is the output weight matrix, and Concat is the concatenation operation;

[0160] Perform a fully connected operation on the ninth feature map to obtain the tenth feature map;

[0161] Add the tenth feature map and the eighth feature map pixel by pixel, and concatenate the addition result with the query vector Q i (i = 1, 2, 3…, n) to obtain the second fusion feature map F 2 .

[0162] The splicing module is used to splice the first fused feature map F 1 and the second fused feature map F 2 to obtain the fused feature map F.

[0163] Specifically, in this embodiment, the dual feature fusion module can capture different feature information and optimize the combination of different feature vectors at the same time, thereby improving the classification accuracy and robustness of the model. The symbol writing in handwritten mathematical formulas has diversity and complexity, with a large number of variable symbols and many mathematical symbols, and some mathematical symbols are very similar in shape. The CSP unit can effectively extract local features, and the multi-scale processing method can focus on detail information of different sizes, while Transformer is good at capturing global context relationships. The combination of the two can enhance the understanding of local details and global context at the same time, solve the problems of regional information or overall structure that may be ignored by a single method. This dual feature fusion can significantly improve the accuracy and robustness of the model for handwritten mathematical formula recognition.

[0164] Specifically, the Parallel Visual Attention (PVA) module is used to perform weighted fusion on the fused feature map to obtain the aligned visual feature map;

[0165] Specifically, handwritten mathematical formula recognition not only requires recognizing the symbols themselves, but also capturing the two-dimensional spatial position relationship between the symbols. The spatial relationship between the symbols can be captured through the attention mechanism. Parallel Visual Attention is an improved attention mechanism. Existing attention-based methods are related to the time term, resulting in low efficiency. While Parallel Visual Attention uses the reading order as the query, making the calculation method independent of time, so it can be processed in parallel, and finally outputs the aligned visual features of all time steps in parallel, effectively improving the correct expression of the attention map for the visual region of the corresponding character, and being able to enhance and distinguish features for key regions more effectively. At the same time, it can dynamically "view" the output of the encoder during the decoding process, assign different weights to different parts of the input sequence, so that the model can focus on the most relevant information.

[0166] Specifically, in this embodiment, as Figure 9 shown, the Parallel Visual Attention module performs weighted fusion on the fused feature map to obtain the aligned visual feature map, including:

[0167] 1) Input the fused feature map F into the parallel attention mechanism for processing to generate an attention map, including:

[0168] The mathematical expression of the parallel attention mechanism is:

[0169]

[0170] Among them, e t,ij is the attention weight at position ij of time step t; W e , W o and W v are trainable weights, O t is the character reading order, and O t = 0, 1, 2,..., N - 1, where N is the number of characters, f o is the embedding function; v ij is the visual feature at position ij in the fused feature map;

[0171] Specifically, in parallel visual attention, the key-value set is the input two-dimensional feature (v ij , v ij ). To enable parallel computation, the reading order is used as the query. The reading order of the first character in the text is 0, the reading order of the second character is 1, and so on.

[0172] 2) Perform pixel-by-pixel multiplication of the attention map and the fused feature map to generate an aligned visual feature map g t , g t ∈R d , and the mathematical expression is:

[0173]

[0174] Among them, g t is the aligned visual feature map at time step t;

[0175] Since the calculation method is independent of time, all time steps of the aligned visual feature map G can be output in parallel through the parallel visual attention module.

[0176] Specifically, the bidirectional mutual learning decoding module is used to perform bidirectional decoding on the aligned visual feature map to obtain the recognition result of the handwritten mathematical formula; specifically, more abundant context information can be obtained through the bidirectional mutual learning decoding module, so as to obtain a more accurate recognition result.

[0177] Specifically, in this embodiment, as Figure 10 shown, the bidirectional mutual learning decoding module performs bidirectional decoding on the aligned visual feature map G, and the recognition result of the handwritten mathematical formula obtained includes:

[0178] The bidirectional mutual learning decoding module includes a first decoder and a second decoder. The first decoder is used to decode the aligned visual feature map from left to right (L2R), and the second decoder is used to decode the aligned visual feature map from right to left (R2L). The first decoder and the second decoder have the same architecture, only different in the decoding direction. The decoding process includes:

[0179] a. Respectively convert the aligned visual feature map G into LaTex sequences through the first decoder and the second decoder and which is expressed as:

[0180]

[0181] where <sos>Indicates the start of a LaTeX sequence; <eos>Indicates the end of the LaTeX sequence, Y 1 , Y 2 ,..., Y and Y T , Y T-1 ,..., Y are the target sequences of the L2R branch and the R2L branch respectively, and T is the length of the LaTeX sequence;

[0182] b. In each decoding step, the first decoder and the second decoder calculate the probability of the next character based on the current hidden state and the output of the previous character. The calculation formulas are as follows:

[0183]

[0184] where, h t and represent the current state and the predicted output of the t-th decoding step in the first decoder respectively, h t ′ and represent the current state and the predicted output of the t-th decoding step in the second decoder respectively; W o , W y , W h , W t , W o ′, W y ′, W h ′ and W t ′ are trainable weight matrices, and W o , W o ′ ∈ R K×d ; W y , W y ′ ∈ R d×n ; W h , W h ′ ∈ R d×n ; W t , W t ′ ∈ R d×D ; d, K, n, and D represent the attention dimension, the number of symbol categories, the dimension of the gated recurrent unit (GRU), and the dimension of the GRU hidden layer state respectively. E and E′ are embedding matrices, max represents the maxout activation function, and h t represents the hidden state generated by the gated recurrent unit and can be generated in the following way:

[0185]

[0186] where, f 1 and f 2 represent two one-way GRU models, and g t is the aligned visual feature map;

[0187] c. Sequence generation and optimization. During the model training process, the sequence generation method gradually generates each symbol in the handwritten mathematical formula.

[0188] The first and second decoders perform interactive learning through the KL loss function to minimize the probability distribution difference between the first and second decoders, including:

[0189] Calculate the probability of the first decoder, expressed as:

[0190]

[0191] Calculate the probability of the second decoder, expressed as:

[0192]

[0193] Where, is the predicted probability of the label symbol at the i-th step of decoding, and

[0194] Define the soft probability distribution of the first decoder as:

[0195]

[0196] Where k is the number of character categories, S is the temperature parameter for generating soft labels, represents the logical value of the i-th symbol in this sequence calculated by the decoder network,

[0197] Calculate and The KL distance between them is:

[0198]

[0199] Where S 2 is a constant factor used to scale the magnitude of the Kullback-Leibler (KL) divergence, ensuring that the true label and the probability distribution from the other branch can contribute to the model training to a considerable extent; and represent the logical values of the first decoder and the second decoder respectively.

[0200] Finally, the second decoder predicts each symbol based on the optimized probability distribution (i.e., obtained by minimizing the probability distribution difference between the first and second decoders), and generates the complete mathematical formula as the recognition result of the handwritten mathematical formula.

[0201] S3: Perform handwritten mathematical formula recognition based on the trained handwritten mathematical formula recognition model.

[0202] Specifically, to verify the effectiveness of the method proposed in this embodiment, the handwritten mathematical formula images after data augmentation are input into the constructed handwritten mathematical formula recognition model for offline training of the handwritten mathematical formula recognition model. The present invention uses the cross-entropy loss function, and the loss function is defined as:

[0203]

[0204] where and is the cross-entropy loss between the target label and the softmax probability between the two branches. λ is a hyperparameter used to balance the recognition loss and the KL divergence loss. Full-scale training is adopted, with a total of 240 training rounds, a batch size of 32, and the optimizer is Adadelta.

[0205] In summary, a handwritten mathematical formula recognition method based on dual feature fusion disclosed in this embodiment has the following beneficial effects:

[0206] The feature fusion adopted in this embodiment is a dual feature fusion based on the combination of Cross Stage Partial (CSP) and Transformer. CSP extracts local features in stages, focusing on the details of handwritten mathematical symbols (such as the starting position of strokes, curvature, thickness, etc.), which helps to distinguish morphologically similar symbols (such as "1" and "I"). Transformer models the global context through the self-attention mechanism, captures the relationships between symbols, identifies the context and role of symbols in handwritten mathematical formulas, and helps to distinguish the meanings of similar symbols (such as "1" and "I") in different contexts. The combination of the two enables the model to distinguish symbols from local details and judge the category of symbols through the global context.

[0207] (2) Parallel visual attention is introduced in this embodiment, and the attention weights at each scale can be adaptively adjusted. The module dynamically allocates attention according to the relative importance between different characters (such as local features or context information), and preferentially focuses on those detailed features that are crucial for character recognition. When dealing with similar characters, such dynamic weight allocation can help the model more accurately capture those tiny but crucial differences. By performing augmentation processing such as rotation and flipping on the handwritten mathematical formula image dataset, the model can better adapt to handwritten mathematical formulas with various angles, postures, and styles, thereby improving its robustness and the effect in practical applications.

[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.< / eos> < / sos>

Claims

1. A handwritten mathematical formula recognition method based on dual feature fusion, characterized in that: The specific steps include: S1: Acquire a handwritten mathematical formula image, and preprocess the handwritten mathematical formula image to obtain a processed handwritten mathematical formula image; S2: establishing a handwritten mathematical formula recognition model, and training the handwritten mathematical formula recognition model based on the processed handwritten mathematical formula image to obtain a trained handwritten mathematical formula recognition model; Among them, the handwritten mathematical formula recognition model established includes: A feature extraction module, used to extract feature maps of different scales of the processed handwritten mathematical formula image; A dual feature fusion module is used to perform feature fusion on the feature maps of different scales in a cross-stage feature fusion manner and in a manner based on a self-attention mechanism, and finally obtain a fused feature map, wherein the fused feature map contains detailed information of mathematical symbols in the handwritten mathematical formula and the correlation between the mathematical symbols; A parallel visual attention module, used for performing weighted fusion on the fused feature maps to obtain an aligned visual feature map; A bidirectional mutual learning decoding module, used for bidirectionally decoding the aligned visual feature map to obtain a recognition result of the handwritten mathematical formula; S3: Perform handwritten mathematical formula recognition based on the trained handwritten mathematical formula recognition model.

2. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 1 is characterized in that: The dual feature fusion module comprises: A cross-stage feature fusion module performs feature fusion on the feature maps of different scales according to a cross-stage feature fusion method to obtain a first fused feature map; The cross-stage feature fusion module includes a plurality of cross-stage CSP units, and the plurality of feature maps in the feature fusion process are enhanced by the plurality of cross-stage CSP units; A Transformer feature fusion module, which performs feature fusion on the feature maps of different scales based on a self-attention mechanism to obtain a second fused feature map; The Transformer feature fusion module includes a position encoding module and a multi-head attention module, and the position encoding module and the multi-head attention module are used to focus on the correlation relationship between several feature maps in the feature fusion process; The splicing module is used to splice the first fused feature map and the second fused feature map to obtain a fused feature map.

3. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 1 is characterized in that: The feature extraction module includes a first convolutional layer, a first dense block, a second convolutional layer, a first pooling layer, a second dense block, a third convolutional layer, a second pooling layer, a third dense block, a third pooling layer and a fully connected layer connected in sequence; The first convolution layer is used to perform a convolution operation on the processed handwritten mathematical formula image to obtain an initial feature map, and transmit it to the first dense block; The first dense block is used to perform a convolution operation on the initial feature map to obtain a plurality of first intermediate convolution features, concatenate the plurality of intermediate convolution features with the initial feature map according to a channel dimension, perform a convolution operation on the concatenated result to obtain a first feature map, and transmit the first feature map to the second convolution layer; The second convolutional layer is used to perform a convolution operation on the first feature map to obtain an intermediate feature map, and transmit the intermediate feature map to the first pooling layer; The first pooling layer is used to perform a pooling operation on the intermediate feature map to obtain a first dimension reduction feature map, and transmit it to the second dense block; The second dense block is used to perform a convolution operation on the first dimension reduction feature map to obtain a plurality of second intermediate convolution features, splice the plurality of first intermediate convolution features, the plurality of second intermediate convolution features, the initial feature map and the first feature map according to the channel dimension, and perform a convolution operation on the splicing result to obtain a second feature map, and transmit the second feature map to the third convolution layer; The third convolution layer is used to perform a convolution operation on the second feature map to obtain a high-level feature map, and transmit it to the second pooling layer; The second pooling layer is used to perform a pooling operation on the high-level feature map to obtain a second reduced-dimensional feature map, and transmit it to the third dense block; The third dense block is used to perform a convolution operation on the second dimension reduction feature map to obtain a plurality of third intermediate convolution features, splice the plurality of first intermediate convolution features, the plurality of second intermediate convolution features, the plurality of third intermediate convolution features, the initial feature map, the first feature map, the first dimension reduction feature map and the second feature map according to the channel dimension, and perform a convolution operation on the splicing result to obtain a third feature map, and transmit the third feature map to the third pooling layer; The third pooling layer is used to perform a pooling operation on the third feature map to obtain a third reduced-dimensional feature map, and transmit it to the fully connected layer; The fully connected layer is used to perform a fully connected operation on the third dimensionality reduction feature map to obtain a fourth feature map.

4. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 3 is characterized in that: The cross-stage feature fusion module includes a first upsampling unit, a first CSP unit, a second upsampling unit, a second CSP unit, a first downsampling unit, a third CSP unit, a second downsampling unit and a fourth CSP unit connected in sequence; The cross-stage feature fusion module performs feature fusion on the feature maps of different scales according to the cross-stage feature fusion method to obtain a first fused feature map, which includes: The first upsampling unit is used to upsample the fourth feature map, and splice the upsampling result with the second feature map according to the channel dimension to obtain a first spliced ​​feature map, and transmit the first spliced ​​feature map to the first CSP unit; The first CSP unit is used to perform an enhancement operation on the first spliced ​​feature map to obtain a fifth feature map, and transmit the fifth feature map to the second up-sampling unit; The second upsampling unit is used to upsample the fifth feature map, and splice the upsampling result with the first feature map according to the channel dimension to obtain a second spliced ​​feature map, and transmit the second spliced ​​feature map to the second CSP unit; The second CSP unit is used to perform an enhancement operation on the second spliced ​​feature map to obtain a sixth feature map, and transmit the sixth feature map to the first downsampling unit; The first downsampling unit is used to downsample the sixth feature map, and splice the downsampling result with the fifth feature map according to the channel dimension to obtain a third spliced ​​feature map, and transmit the third spliced ​​feature map to the third CSP unit; The third CSP unit is used to perform an enhancement operation on the third spliced ​​feature map to obtain a seventh feature map, and transmit the seventh feature map to the second down-sampling unit; The second down-sampling unit is used to down-sample the seventh feature map, and splice the down-sampling result with the fourth feature map according to the channel dimension to obtain a fourth spliced ​​feature map, and transmit the fourth spliced ​​feature map to the fourth CSP unit; The fourth CSP unit is used to perform an enhancement operation on the fourth spliced ​​feature map to obtain a first fused feature map.

5. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 4 is characterized in that: The first CSP unit, the second CSP unit, the third CSP unit and the fourth CSP unit each include: a first convolution block, a channel separation unit, a Bottleneck unit and a second convolution block; The first convolution block is used to perform a convolution operation on the input concatenated feature map, perform batch normalization processing on the convolution result, use the SiLU activation function to perform a nonlinear transformation on the normalized result to obtain an activation feature map, and transmit the activation feature map to the channel separation unit; The channel separation unit is used to perform channel separation processing on the activation feature map to equally divide the number of channels of the activation feature map into two parts to obtain two independent feature maps, and randomly transmit one part of the independent feature map directly to the second convolution block, and transmit the other part of the independent feature map to the Bottleneck unit; The Bottleneck unit performs compression, expansion and recompression operations on the input independent feature map for several times to obtain compressed features and transmit them to the second convolution block; The second convolution block is used to perform a convolution operation on the input independent feature map and the compressed feature, perform batch normalization on the convolution result, and use the SiLU activation function to perform a nonlinear transformation on the normalized result to obtain an enhanced feature map.

6. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 4 is characterized in that: Based on the Transformer feature fusion module, feature fusion is performed on the feature maps of different scales based on the self-attention mechanism to obtain a second fused feature map including: Flatten the first feature map, the second feature map, and the fourth feature map into one-dimensional feature maps respectively through a flattening operation; Unify the scales of each one-dimensional feature map through linear projection operation; The three one-dimensional feature maps after the unified dimension are concatenated into a fifth concatenated feature map; Position encoding the fifth splicing feature map by the position encoding module, and adding the position-encoded fifth splicing feature map to the original fifth splicing feature map pixel by pixel to obtain an eighth feature map; Processing the eighth feature map through a multi-head attention module includes: Calculate the element x in the eighth feature map i The corresponding query vector Q i , key vector K i Sum value vector V i , the mathematical expression is: Where i = 1, 2, 3…, n; W Q is the query vector weight matrix; W K is the key vector weight matrix; W V is the value vector weight matrix; Calculates a pair of elements x i and x j The attention score is expressed as: Among them, d k is the dimension of the key vector; K j is the element x j The key vector K j ; The softmax function is used to convert the attention score value to between 0 and 1, and the sum is 1, and the attention weight ω is obtained. ij , the mathematical expression is: ω ij =softmax(score(Q i ,K j )) (3) The element x i The value vector V i The corresponding attention weight ω ij Multiply and sum to get the attention output z i , the mathematical expression is: All attention outputs are concatenated to obtain the ninth feature map, the mathematical expression is: W=W O Concat(z1,z2,...,z h ) (5) Among them, W O is the output weight matrix, Concat is the concatenation operation; Performing a full connection operation on the ninth feature map to obtain a tenth feature map; The tenth feature map is added pixel by pixel to the eighth feature map, and the addition result is added to the query vector Q i Perform splicing, i=1,2,3…,n, and obtain the second fused feature map.

7. The handwritten mathematical formula recognition method based on dual feature fusion according to claim 1 is characterized in that: The parallel visual attention module performs weighted fusion on the fused feature maps to obtain an aligned visual feature map, including: 1) Inputting the fused feature map into the parallel attention mechanism for processing to generate an attention map, including: The mathematical expression of the parallel attention mechanism is: Among them, e t,ij is the attention weight of position ij at time step t; W e , W o and W v is a trainable weight, O t is the character reading order, and O t =0,1,2,...,N-1, N is the number of characters, f o is the embedding function; v ij It is the visual feature at position ij in the fused feature map; 2) Multiply the attention map by pixel and the fused feature map to generate the aligned visual feature map g t , g t ∈R d , the mathematical expression is: Among them, g t is the aligned visual feature map at time step t; Finally, the aligned visual feature maps G of all time steps are output in parallel through the parallel visual attention module.

Citation Information

Cited By

  • Power transmission and transformation equipment inspection infrared image super-resolution reconstruction method and system

    CN120976022A