A structured image content recognition method based on dynamic feature extraction
Through technical means such as dynamic feature selection and relative position coding, the problems of large calculation overhead and neglect of relative position information of existing structured image content recognition methods are solved, achieving higher recognition accuracy and better application prospects.
Patent Information
- Application Number
- CN202210415242.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The existing structured picture content recognition method has too much calculation overhead when processing large-size feature maps, and ignores the relative position information between characters, resulting in a decrease in recognition accuracy.
The dynamic feature selection mechanism is adopted to select some useful feature vectors from the large-size feature map, remove redundant features, and enhance the generalization ability of the model through dynamic offset. Relative position encoding and position environment information are introduced into the spatial relationship encoder to extract more complex character spatial relationships.
Without significantly increasing the computing overhead, the model's ability to acquire complex spatial relationships in structured pictures is improved, the recognition accuracy is significantly improved, and it has good and wide application prospects.
Smart Images

Figure CN115019319B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image text recognition, and in particular to a structured image content recognition method based on dynamic feature extraction. Background Art
[0002] Structured image content recognition refers to converting the content of structured images such as mathematical formulas, chemical formulas, and musical scores into text sequences to facilitate editing, retrieval, and other operations. It has a wide range of applications in artificial intelligence scenarios such as personalized recommendations, music retrieval, and automatic problem solving. Structured image content recognition is more challenging than traditional text recognition because it not only requires determining all the characters in the image, but also requires determining the spatial relationship between all the characters in the image.
[0003] Recently, encoder-decoder models based on deep learning have been widely used to solve the problem of structured image content recognition. These methods generally consist of three modules: a visual encoder (convolutional neural network) that extracts the semantic features of the input image, a spatial relationship encoder that extracts the spatial relationship between characters, and a text decoder (recurrent neural network) that is used to predict the output sequence. Compared with non-deep learning methods, the encoder-decoder model has achieved a certain improvement in the accuracy of structured image content recognition problems, but it still has the following problems:
[0004] 1) Structured image content recognition requires fine-grained visual features. Usually, most models use a convolutional neural network with a small receptive field to encode the input image. However, this will produce a large-size feature map, which brings huge computational overhead to the spatial relationship encoder. Especially for spatial relationship encoders with complex operations, the amount of computation is often unbearable. There are two solutions to this problem in existing methods. One method is to simplify the extraction of position features. In the model proposed by Deng et al., the computational overhead of the spatial relationship encoder is reduced by only considering the spatial relationship of characters in the same line. However, this method ignores the spatial relationship of characters across lines, which reduces the recognition accuracy. Another method is to reduce the size of the visual encoder feature map as much as possible while retaining fine-grained features. Fu et al. used a connected domain-based character cutting algorithm to extract character-level features of the input image, which reduced the number of features of the visual encoder compared to other methods. However, for structured images with complex backgrounds (such as music scores), the connected domain segmentation algorithm does not work properly and the recognition performance is significantly reduced.
[0005] 2) Structured images have very rich spatial position information. For the characters in the image, in addition to the absolute position in the entire image, there is also the relative position between the characters. For structured image content recognition, relative position information is easier to infer the spatial semantics of characters than absolute position information. In existing structured image content recognition methods, the spatial relationship encoder only uses absolute position information and hardly considers the relative position between characters. In addition, unlike one-dimensional text sequences, the arrangement of characters in two-dimensional space is not closely connected, and for a character, the position environment information of whether there are other characters around it is not considered. Summary of the invention
[0006] The purpose of the present invention is to design a structured image content recognition method based on dynamic feature extraction in response to the deficiencies of the prior art. A dynamic feature selection mechanism is adopted to select some useful feature vectors from a large-size feature map to reduce the number of features, remove redundant features in the feature map, and then dynamically offset them to enhance the generalization ability of the model. Relative position coding and position environment information are introduced into the spatial relationship encoder, and more complex character spatial relationships are extracted, further improving the accuracy of structured image content recognition. Without significantly increasing the computational overhead, the model's ability to acquire complex spatial relationships in structured images is greatly improved. The method is simple, has high accuracy, and has good and broad application prospects.
[0007] The object of the present invention is achieved as follows: a structured image content recognition method based on dynamic feature extraction, which is characterized by adopting a dynamic feature selection mechanism, selecting some useful feature vectors from a large-size feature map to remove redundant features in the feature map, dynamically offsetting them, and introducing relative position coding and position environment information into a spatial relationship encoder to extract more complex character spatial relationships. The recognition of structured image content specifically includes the following steps:
[0008] 1) Fine-grained visual feature extraction: Use a convolutional neural network with a small receptive field to extract the fine-grained visual features of the input structured image, calculate the absolute position encoding of the feature vector in the feature map, and fuse the absolute position encoding with the fine-grained visual features.
[0009] 2) Dynamic feature selection: Use a neural network to determine the character type represented by each feature vector in the fine-grained visual features. Define a loss function that can be used for feature selection, set the scale parameter of the selected features, and determine the coordinates of the valid features in the feature map. Define a dynamic offset distribution, dynamically offset the selected coordinates according to the distribution, and obtain the final feature vector.
[0010] 3) Spatial relationship extraction: In the selected features, the relative position encoding of each pair of feature vectors in the complete feature map is calculated. The position environment information of each feature vector in the complete feature map is calculated. The spatial relationship between feature vectors is extracted using a spatial relationship extractor that combines relative position encoding and position environment information.
[0011] 4) Text decoding: Use the decoding model for text generation to decode the text sequence of structured image content.
[0012] 5) Model training: First, use the optimizer to train the loss function in the dynamic feature selection step and update some of the relevant model parameters. Then define the total loss function of the model and use the optimizer to update all the parameters of the model to obtain the text sequence of the structured image content.
[0013] In the step of extracting fine-grained visual features, a convolutional neural network with a very small receptive field is used to extract all the character details in the image; in the generated large feature map, the two-dimensional coordinates of each feature vector are calculated, and its absolute position code is calculated using an embedding matrix, and the absolute position code is fused with the fine-grained visual features.
[0014] In the dynamic feature selection step, a fully connected neural network is used to determine the category of each feature vector in the vocabulary of fine-grained visual features, a feature selection loss function is defined, a scale parameter of feature selection is set, and the coordinates of the selected features in the large feature map are determined; a dynamic offset distribution is defined, and sampling is performed according to the dynamic offset distribution with each coordinate as the center to obtain the offset coordinates, and finally the selected feature vector is determined.
[0015] In the spatial relationship extraction step, an embedding matrix is used to calculate the row relative position coding and column relative position coding of each pair of selected feature vectors; a convolutional neural network is used to calculate the position environment information of each feature vector in a certain area, and the relative position coding and position environment information are introduced into the transformer model to extract the spatial relationship between the selected feature vectors.
[0016] In the text decoding step, a transformer model is used, and the result of the spatial relationship extraction step is taken as input, and further decoded to obtain the final text prediction result.
[0017] In the model training step, the Adam optimizer is used to first train the feature selection loss function in the dynamic feature selection step, and the corresponding partial model parameters are updated. After the training is completed, the total loss function of the model is defined, and the Adam optimizer is used to update all parameters of the model until the total loss function converges.
[0018] Compared with the prior art, the present invention has a simple method and high accuracy. Under the condition of low computational overhead, the model's ability to obtain complex spatial relationships in structured images is improved, and more complex character spatial relationships are extracted. On the basis of retaining the fine-grained visual features of structured images, the present invention uses a dynamic feature selection mechanism to remove redundant features in the feature map, greatly reducing the computational overhead in the spatial relationship encoder, and the dynamic offset mechanism improves the generalization ability of the model. On this basis, the relative position information between character features and the position environment information of the characters are introduced into the spatial relationship encoder, further improving the accuracy of structured image content recognition, and has good and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of structured image content recognition in Example 1. DETAILED DESCRIPTION
[0020] Based on fine-grained visual features, the present invention uses a dynamic feature selection mechanism to first select some useful feature vectors from large-scale feature maps to reduce the number of features, and then dynamically offset them to enhance the generalization ability of the model. Relative position encoding and position environment information are introduced into the spatial relationship encoder, which improves the model's ability to obtain complex spatial relationships in structured images without significantly increasing computational overhead.
[0021] The present invention will be further described in detail in conjunction with the following specific embodiments and drawings. Except for the contents specifically mentioned below, the processes, conditions, experimental methods, etc. for implementing the present invention are common knowledge and common common sense in the field and are not particularly limited by the present invention.
[0022] Example 1
[0023] See also Figure 1 The present invention recognizes the structured image content according to the following steps:
[0024] 1) Fine-grained visual feature extraction: Input the structured image into a convolutional neural network with a small receptive field, extract the fine-grained visual features of the structured image, calculate the absolute position code in the feature map, and fuse it with the fine-grained visual features. The specific steps are as follows:
[0025] Step 1: Use a convolutional neural network with a small receptive field to encode the input image and obtain the feature map V represented by the following formula (a):
[0026] V={v (i,j) |i=1,...,H; j=1...,W} (a);
[0027] Where: H and W are the height and width of the feature map respectively.
[0028] Step 2: Use two embedding matrices to transform each feature vector v (i,j) The row and column coordinates i and j are encoded as two d / 2-dimensional position vectors respectively, and the d-dimensional absolute position encoding vector p is obtained by concatenating them. (i,j) .
[0029] Step 3: Put v (i,j) and p (i,j) Add them together to obtain the fine-grained feature set E represented by the following formula (b):
[0030] E={e (i,j) |i=1,...,H; j=1,...,W} (b);
[0031] Among them, e (i,j) =v (i,j) +p (i,j) .
[0032] 2) Dynamic feature selection: Use a fully connected neural network to predict the character type (including background type) of each feature vector in the feature map in the dictionary, define the feature selection loss function, and set the scale parameter of feature selection; find the feature vector of non-background type based on the prediction results of the fully connected network and retain its coordinates; define a dynamic offset distribution, sample the coordinates according to the offset distribution, and use the offset coordinates to select feature vectors from the fine-grained feature set. The specific steps are as follows:
[0033] Step 4: Use a fully connected neural network to train e (i,j) Perform character prediction (including background type) and obtain the probability distribution a of character prediction (i,j) .
[0034] Step 5: Accumulate and normalize the probability distribution of all feature vector predictions by character category using the following formula (c):
[0035]
[0036] Among them, k represents the kth character.
[0037] Step 6: Count the number of occurrences of the characters in each image label by character category and normalize them using the following formula (d):
[0038]
[0039] Among them, k represents the kth character.
[0040] Step 7: Calculate the feature selection loss using the following formula (e):
[0041]
[0042] Among them: α is the proportion of selected features to all features; ∈ represents the type of background; C is the number of all characters in the dictionary.
[0043] Step 8: Select the feature vectors predicted to be non-background type and record their coordinate value set A by the following formula (f): loc ,A loc :
[0044] A loc ,A loc ={(h1,w1),(h2,w2),...,(h m ,w m )} (f).
[0045] Step 9: Define the probability distribution p(i,j) with coordinate (i,j) as the center point, sample according to p(i,j), and obtain the offset coordinates of the (i,j) coordinates by the following formulas (g) to (h):
[0046]
[0047]
[0048] Step 10: According to the offset coordinates obtained in step 9, the vectors in the set E obtained in step 3 are taken out to obtain the selected feature set A represented by the following formula (i):
[0049]
[0050] 3) Spatial relationship extraction: Calculate the relative position encoding of each pair of feature vectors in the selected feature set, calculate the position environment information of each feature vector, and use the spatial relationship extractor that combines the relative position encoding and position environment information to extract the spatial relationship between the feature vectors. The specific steps are as follows:
[0051] Step 11: Calculate the row and column relative positions of each pair of feature vector coordinates in set A, and use two embedding matrices to encode the row and column relative positions.
[0052] Step 12: Define a mask matrix with the same size as the original feature map, record the position of the vector coordinates in set A as 1, and record other positions as 0.
[0053] Step 13: Use a convolutional neural network to encode the mask matrix, output a location environment information feature map of the same size as the mask matrix, and select the corresponding vector in the feature map according to the coordinates of the feature vector in A to obtain the location environment information set S represented by the following formula (j):
[0054] S={s1,s2,...,s m} (j);
[0055] Among them, s i is a i The corresponding location environment information.
[0056] Step 14: Define the attention mechanism that combines relative position encoding and position context information by the following formula (k):
[0057]
[0058] Step 15: Use the transformer model and replace the original attention mechanism in the transformer with the attention mechanism defined in step 14 to encode the set A in step 10, and output the set U represented by the following formula (l):
[0059] U={u1,u2,...,u m} (l).
[0060] 4) Text decoding: Use the decoding model for text generation to decode the text sequence of the structured image content. The specific steps are as follows:
[0061] Step 16: Use the transformer model as a decoder to decode the feature vector output in step 15 by time step to generate a text sequence of structured image content.
[0062] 5) Model training: First, train the loss function in the dynamic feature selection step until the loss function converges, then define the total loss function of the model, and train the entire model until the total loss function converges. The specific steps are:
[0063] Step 17: Use the Adam optimizer to train the loss function L in step 7 sace , update the corresponding model parameters to the loss function L sace convergence.
[0064] Step 18: Define the total loss function of the model as L = L sace +L out .L sace is the loss function in step 7, L out is the cross entropy loss function defined using the output of step 16.
[0065] Step 19: Use the Adam optimizer to train the total loss function L in step 18, update all model parameters until the loss function L converges, and obtain the text sequence of the structured image content.
[0066] Compared with the existing structured image content recognition methods, the present invention has the following contributions:
[0067] 1) The present invention adopts a dynamic feature selection mechanism to select effective features for decoding from the fine-grained visual features generated by the convolutional neural network, which reduces the number of visual features with almost no loss of image details. In addition, the dynamic offset mechanism increases the generalization ability of the model.
[0068] 2) Under the action of the dynamic feature selection mechanism, only part of the features are input into the spatial relationship encoder, reducing the computational overhead of the spatial relationship encoder. On this basis, the present invention introduces the relative position features and position environment information between characters into the spatial relationship encoder, enabling the model to capture the complex spatial relationship between characters in the structured image.
[0069] 3) The present invention has conducted a large number of experiments on mathematical formula data sets, mathematical equation data sets and music score data sets. The experiments show that the present invention can achieve the best experimental performance on the three data sets by using only 40% to 60% of the feature vectors in the original feature map. The recognition accuracy on all data sets has obvious advantages over the existing structured image content recognition model.
[0070] The above is only a further explanation of the present invention and is not intended to limit this patent. Any equivalent implementation of the present invention should be included in the scope of the claims of this patent.
Claims
1. A structured image content recognition method based on dynamic feature selection, characterized in that A dynamic feature selection mechanism is used to select some useful feature vectors from the large-size feature map to remove redundant features in the feature map, dynamically offset them, and introduce relative position encoding and position environment information into the spatial relationship encoder to extract more complex character spatial relationships. The recognition of structured image content specifically includes the following steps: (I) Fine-grained visual feature extraction Use a convolutional neural network with a small receptive field to extract fine-grained visual features of the input structured image, calculate the absolute position encoding of the feature vector in the feature map, and fuse the absolute position encoding with the fine-grained visual features; The specific steps of extracting fine-grained visual features are as follows: Step 1: Use a convolutional neural network with a small receptive field to encode the input image and obtain the feature map V represented by the following formula (a): V={v (i,j) |i = 1,...,H;j = 1..., W} (a); Where: H and W are the height and width of the feature map respectively; Step 2: Use two embedding matrices to transform each feature vector v (i,j) The row and column coordinates i and j are encoded as two d / 2-dimensional position vectors respectively, and the d-dimensional absolute position encoding vector p is obtained by concatenating them. (i,j) ; Step 3: Put v (i,j) and p (i,j) Add them together to obtain the fine-grained feature set E represented by the following formula (b): Among them, e (i,j) =v (i,j) +p (i,j) ; (II) Dynamic feature selection Use a fully connected neural network to determine the character type represented by each feature vector in the fine-grained visual features, define a loss function that can be used for feature selection, set the scale parameter of the selected features, and determine the coordinates of the valid features in the feature map; define a dynamic offset distribution, dynamically offset the selected coordinates according to the distribution, and obtain the final feature vector; The specific steps of dynamic feature selection are as follows: Step 4: Use a fully connected neural network to train e (i,j) Perform character prediction and obtain the probability distribution a of character prediction (i,j) ; Step 5: Accumulate and normalize the probability distribution of all feature vector predictions by character category using the following formula (c): Among them, k represents the kth character; Step 6: Count the number of characters appearing in each image label by character category and normalize them using the following formula (d): Among them, k represents the kth character; Step 7: Calculate the feature selection loss using the following formula (e): Among them: α is the proportion of selected features to all features; ∈ represents the type of background; C is the number of all characters in the dictionary; Step 8: Select the feature vectors predicted to be non-background type and record their coordinate value set A by the following formula (f): loc : A loc = {(h1,w1), (h2, w2), ..., (h m ,w m )} (f); Step 9: Define the probability distribution p(i, j) with coordinate (i, j) as the center point, perform sampling according to p(i, j), and obtain the offset coordinates of the coordinate (i, j) by the following formulas (g) to (h): Step 10: According to the offset coordinates obtained in step 9, the vectors in the set E obtained in step 3 are taken out to obtain the selected feature set A represented by the following formula (i): (III) Spatial relationship extraction Among the selected features, the relative position encoding of each pair of feature vectors in the complete feature map is calculated, the position environment information of each feature vector in the complete feature map is calculated, and the spatial relationship between the feature vectors is extracted using a spatial relationship extractor that combines the relative position encoding and the position environment information; The specific steps of spatial relationship extraction are: Step 11: Calculate the row and column relative positions of each pair of feature vector coordinates in set A, and use two embedding matrices to encode the row and column relative positions; Step 12: Define a mask matrix with the same size as the original feature map, record the position of the vector coordinates in set A as 1, and the other positions as 0; Step 13: Use a convolutional neural network to encode the mask matrix, output a location environment information feature map of the same size as the mask matrix, and select the corresponding vector in the feature map according to the coordinates of the feature vector in A to obtain the location environment information set S represented by the following formula (j): S={s1, s2, ..., s m } (j); Among them, s i is a i Corresponding location environment information; Step 14: Define the attention mechanism that combines relative position encoding and position context information by the following formula (k): Step 15: Use the background model and replace the original attention mechanism in the background with the attention mechanism defined in step 14 to encode the set A in step 10, and output the set U represented by the following formula (l): U = {u1, u2, ..., u m } (l); (IV) Text decoding Use the decoding model for text generation to decode the text sequence of the structured image content; (V) Model training The optimizer is used to train the loss function in the dynamic feature selection step, update some related parameters, and then define the total loss function. The optimizer is used to update all parameters to obtain the text sequence of the structured image content.
2. The structured image content recognition method based on dynamic feature selection according to claim 1 is characterized in that In the step of extracting fine-grained visual features, a convolutional neural network with a small receptive field extracts all character details in the image, calculates the two-dimensional coordinates of each feature vector in the generated large feature map, and uses an embedding matrix to calculate its absolute position encoding, which is then fused with the fine-grained visual features.
3. The structured image content recognition method based on dynamic feature selection according to claim 1, characterized in that In the dynamic feature selection step, a fully connected neural network is used to determine the category of each feature vector in the vocabulary of fine-grained visual features, a feature selection loss function is defined, a scale parameter of feature selection is set, and the coordinates of the selected features in the large feature map are determined; a dynamic offset distribution is defined, and sampling is performed according to the dynamic offset distribution with each coordinate as the center to obtain the offset coordinates, and finally the selected feature vector is determined.
4. The structured image content recognition method based on dynamic feature selection according to claim 1 is characterized in that In the spatial relationship extraction step, the row relative position coding and column relative position coding of each pair of selected feature vectors are calculated using an embedding matrix, and the position environment information of each feature vector in a certain area is calculated using a convolutional neural network. The relative position coding and position environment information are introduced into the transformer model to extract the spatial relationship between the selected feature vectors.
5. The structured image content recognition method based on dynamic feature selection according to claim 1 is characterized in that In the text decoding step, a transformer model is used, and the result of the spatial relationship extraction step is taken as input, and further decoded to obtain the final text prediction result.
6. The structured image content recognition method based on dynamic feature selection according to claim 1 is characterized in that In the model training step, the Adam optimizer is used to first train the feature selection loss function in the dynamic feature selection step, and the corresponding partial model parameters are updated. After the training is completed, the total loss function of the model is defined, and the Adam optimizer is used to update all parameters of the model until the total loss function converges.