Text recognition method, device, computer equipment and computer-readable storage medium

HK40085627BActive Publication Date: 2026-09-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
HK42023073839
Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-09-04
Estimated Expiration
2041-08-16

AI Technical Summary

Technical Problem

Existing OCR models have poor recognition performance when faced with font deformation in different scenarios, and obtaining training samples requires a lot of manpower, resulting in high training difficulty.

Method used

The training effect of the feature extraction model is enhanced by training it with unlabeled text image samples. The DenseNet neural network and multi-head attention mechanism are used for image feature extraction and attention feature extraction. The training sample index is calculated and predicted by combining image attribute information.

Benefits of technology

It improves the recognition ability of OCR models in different scenarios, reduces the training difficulty, reduces the dependence on labeled samples, and enhances the generalization ability of feature extraction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a text recognition method and device, computer equipment and a computer readable storage medium. The method comprises: obtaining a text image sample; performing image index calculation according to image attribute information of the text image sample, and determining a reference sample index based on a calculation result; performing image feature extraction processing on the text image sample by using a feature extraction model to obtain image feature information; performing attention feature extraction based on the image feature information by using the feature extraction model to obtain attention feature information of a context information of interest; predicting a prediction sample index based on the attention feature information; and training the feature extraction model according to the prediction sample index and the corresponding reference sample index, so as to extract the attention feature information of a to-be-recognized text image by using the trained feature extraction model to perform image text recognition. The method can train the feature extraction model by using a large number of unlabeled text image samples, and can enhance the training effect of the feature extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, specifically to a text recognition method, apparatus, computer device, and computer-readable storage medium. Background Technology

[0002] Optical Character Recognition (OCR) refers to the process by which computer devices detect the shape of characters, such as characters printed on paper or characters contained in images, and then use character recognition methods to translate the detected shapes into computer text. In some application scenarios, such as advertising and promotional posters, fonts are often deformed, and these deformations are quite diverse. To improve recognition accuracy, it is necessary to obtain a large number of training samples in the corresponding scenarios, and to annotate these training samples. The model is then trained using the annotated training data to improve its ability to recognize characters.

[0003] However, when the trained model is applied to other scenarios, the recognition effect is poor due to the different deformation methods of the font. Furthermore, obtaining training samples in different scenarios and annotating a large number of training samples requires a lot of manpower, resulting in high difficulty in obtaining training samples and high difficulty in training the model. Summary of the Invention

[0004] This application provides a text recognition method, apparatus, computer device, and computer-readable storage medium, which can use unlabeled text image samples to train a feature extraction model, thereby enhancing the training effect of the feature extraction model.

[0005] This application provides a text recognition method, including:

[0006] Obtain text image samples;

[0007] Image indices are calculated based on the image attribute information of the text image samples, and reference sample indices are determined based on the calculation results.

[0008] The text image sample is processed by a feature extraction model to extract image features, thereby obtaining the image feature information of the text image sample.

[0009] Based on the image feature information, the feature extraction model extracts attention features from the text image sample to obtain attention feature information of the attention context of the text image sample.

[0010] Based on the attention feature information of the text image sample, predict the predicted sample index of the text image sample;

[0011] The feature extraction model is trained based on the predicted sample index and the corresponding reference sample index, so as to extract the attention feature information of the text image to be recognized and perform image text recognition through the trained feature extraction model.

[0012] Accordingly, embodiments of this application also provide a text recognition device, comprising:

[0013] The acquisition unit is used to acquire text image samples;

[0014] The calculation unit is used to calculate image indicators based on the image attribute information of the text image sample, and determine the reference sample indicators of the text image sample based on the calculation results;

[0015] The first feature extraction unit is used to perform image feature extraction processing on the text image sample through a feature extraction model to obtain the image feature information of the text image sample.

[0016] The second feature extraction unit is used to extract attention features from the text image sample based on the image feature information using the feature extraction model, so as to obtain attention feature information of the attention context information of the text image sample;

[0017] The prediction unit is used to predict the predicted sample index of the text image sample based on the attention feature information of the text image sample.

[0018] The training unit is used to train the feature extraction model based on the predicted sample indicators and the corresponding reference sample indicators, so as to extract the attention feature information of the text image to be recognized through the trained feature extraction model for image text recognition.

[0019] Accordingly, this application also provides a computer device including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any of the text recognition methods provided in this application.

[0020] Accordingly, embodiments of this application also provide a computer-readable storage medium for storing a computer program, which is loaded by a processor to execute any of the text recognition methods provided in embodiments of this application.

[0021] This application embodiment acquires text image samples; calculates image metrics based on the image attribute information of the text image samples, and determines reference sample metrics based on the calculation results; performs image feature extraction processing on the text image samples using a feature extraction model to obtain image feature information of the text image samples; extracts attention features from the text image samples based on the image feature information using the feature extraction model to obtain attention feature information of the attention context of the text image samples; predicts predicted sample metrics of the text image samples based on the attention feature information of the text image samples; and trains the feature extraction model based on the predicted sample metrics and the corresponding reference sample metrics to extract attention feature information of the text image to be recognized for image text recognition. This scheme trains the feature extraction model using reference sample metrics and predicted sample metrics, and can utilize a large number of unlabeled text image samples to train the feature extraction model, thereby enhancing the training effect of the feature extraction model. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a scene diagram of the text recognition method provided in the embodiments of this application;

[0024] Figure 2 This is a flowchart of the text recognition method provided in the embodiments of this application;

[0025] Figure 3 This is a flowchart of the image restoration process provided in the embodiments of this application;

[0026] Figure 4 This is another flowchart of the text recognition method provided in the embodiments of this application;

[0027] Figure 5 This is a schematic diagram of the feature extraction network structure provided in an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of the model structure provided in the embodiments of this application;

[0029] Figure 7 This is a schematic diagram of the text recognition device provided in the embodiments of this application;

[0030] Figure 8This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] This application provides a text recognition method, apparatus, computer device, and computer-readable storage medium. The text recognition apparatus can be integrated into a computer device, which may be a server or a terminal, etc.

[0033] The terminal may include mobile phones, wearable smart devices, tablets, laptops, personal computers (PCs), and in-vehicle computers, etc.

[0034] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0035] For example, such as Figure 1 As shown, text recognition methods can include upstream tasks, downstream tasks, and feature extraction tasks. In the feature extraction task, the computer device can acquire text image samples. The DenseNet neural network of the feature extraction model performs image feature extraction processing on the text image samples to obtain image feature information of the text image samples. After the feature extraction model performs random masking on the image feature information, it performs attention feature extraction on the masked image feature information based on a multi-head attention feature mechanism to obtain the initial attention feature information of the text image samples. The initial attention feature information is then normalized by the first Batch Normalization (BN) layer, and then normalized again by the feedforward network and the second BN layer to obtain the attention feature information. The BN layer is used to normalize the data in the attention feature information into a data range, reducing the degree of data divergence and reducing the training difficulty of the feature extraction model.

[0036] The upstream task is used to train the feature extraction model, improving its feature extraction capabilities. The computer can calculate image metrics based on the image attribute information of the text image samples. For example, based on the color, contour, shape, and texture features of the text image samples, image metrics such as color histogram, boundary features, homogeneity, contrast, and entropy are calculated. Image metrics with matching expression types are merged to obtain reference sample metrics for the text image sample pairs. Predicted sample metrics are then predicted based on attention feature information using different fully connected layers. For example, image restoration and dimensionality transformation are used to obtain predicted sample metrics. The feature extraction model is trained based on the error between the reference and predicted sample metrics to obtain a pre-trained feature extraction model.

[0037] The downstream task involves extracting attention feature information from a small number of labeled target image samples using a pre-trained feature extraction model, and then predicting the target image samples based on this attention feature information using a text recognition model. The parameters of both the pre-trained feature extraction model and the text recognition model are adjusted based on the prediction results and sample labels to obtain a trained feature extraction model and a trained text recognition model. These models are then used to perform image-to-text recognition on the text image to be recognized. This approach trains the feature extraction model using reference sample metrics and predicted sample metrics, and can utilize a large number of unlabeled text image samples for training, thus enhancing the training effect of the feature extraction model.

[0038] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0039] This embodiment will be described from the perspective of a text recognition device, which can be integrated into a computer device, such as a server or a terminal. Figure 2 As shown, the specific process of the text recognition method is as follows:

[0040] 101. Obtain text image samples.

[0041] Among them, text image samples can be training samples used to train the feature extraction model. Text image samples can contain text and characters, and the text can be text from different languages, as well as various forms of artistic fonts. Text image samples can be text image samples without labels.

[0042] For example, text image samples can be obtained from a database or blockchain, or the text image to be recognized uploaded by the user on the terminal can be used as a text image sample.

[0043] 102. Calculate image indices based on the image attribute information of the text image samples, and determine the reference sample indices of the text image samples based on the calculation results.

[0044] Among them, image attribute information can be information that characterizes the properties of text image samples, such as the color value and brightness of each pixel in different color channels in the text image sample.

[0045] Among them, image index calculation can be a calculation of text image samples based on image attribute information under at least one feature, such as calculating text image samples based on color features, contour features, shape features, or texture features.

[0046] The reference sample index can be a reference sample index determined by the calculation result based on the image index. The reference sample index can be equivalent to a label for the text image sample, and is used to compare with the prediction sample index obtained by prediction based on the attention feature information extracted by the feature extraction model.

[0047] For example, specifically, for at least one feature, image attribute information of the text image sample can be obtained and calculated to obtain image index information (image index information is the calculation result obtained by performing image index calculation on image attribute information). For example, for color features, the color histogram (e.g., RGB color histogram, HSV color histogram, grayscale histogram) of the text image sample can be calculated from the color values ​​of different channels of each pixel in the text image sample, and / or the color set and color moments, etc.; for contour features, the boundary features of the text image sample can be obtained through Hough transform, the edge orientation histogram of the text image sample can be calculated, and / or the Fourier shape description of the text image sample can be performed to obtain the Fourier descriptor of the text image sample, etc.; for shape features, the edges, corners, regions, and / or ridges of the text image sample can be calculated. Regarding the distribution of text image samples, for texture features, homogeneity, contrast, dissimilarity, entropy, second moment of angle, and / or correlation can be calculated based on the brightness and color value of each pixel. For the overall features of text image samples, attribute information such as the color value of each pixel in different channels can be calculated to obtain image index information about the text image samples (image index information can be a three-dimensional tensor of size Channels*H*W, where H*W is the size of the text image sample, with a length of H pixels and a width of W pixels, and Channels is the number of channels, which can be flexibly set according to different image attribute information).

[0048] Image index information calculated based on the image attribute information of the text image sample is used as the reference sample index of the text image sample. For example, the one-dimensional tensor corresponding to the RGB color histogram calculated based on the color features can be used as the reference sample feature of the text image sample.

[0049] The color histogram can be obtained by dividing a color space (e.g., RGB, HSV, and grayscale) into several small color intervals and calculating the number of pixels that fall within each small interval. Based on the number of pixels of each color, an image index information about the color histogram can be obtained, which can be represented as a one-dimensional tensor. Similarly, the edge orientation histogram can be obtained by calculating the number of pixels in the text image whose edge orientation falls within each small interval.

[0050] A color set can be created by converting the RGB color space into a visually balanced color space (such as HSV space) and quantizing the color space into several color intervals. An image is divided into several regions using automatic color segmentation technology, with each region indexed by a specific color component of the quantized color space. This represents the image as a binary color index set, which can be represented as a one-dimensional tensor. This one-dimensional tensor is an image index calculated based on the image attribute information of the text image sample.

[0051] Color moments can be extracted from text image samples, such as first-order moments, second-order moments, and third-order moments. Based on the extracted color moments, a one-dimensional tensor can be obtained, which can be an image index information.

[0052] The Hough transform can identify geometric shapes in text image samples. Based on the Hough transform, a calculation result can be obtained, which is an image extracted from the geometric shapes of the text image sample. Specifically, performing the Hough transform on a text image sample yields an H*W two-dimensional tensor, where H*W is the size of the text image sample (H pixels long and W pixels wide). This two-dimensional tensor can serve as an image index. Similarly, by performing Fourier shape description on the text image sample, and detecting the distribution of edges, corners, regions, and ridges within the text image sample, corresponding two-dimensional tensors can be obtained. Each two-dimensional tensor represents an image index calculated based on the image attribute information of the text image sample.

[0053] Homogeneity, contrast, dissimilarity, entropy, second moment of angle, and correlation can be measures of the texture features of text image samples. Specifically, the corresponding values ​​can be calculated based on the grayscale and brightness of the text image samples, and each value is an image index information.

[0054] If image index calculations are performed for multiple features, a large number of reference sample indices will be obtained. This necessitates predicting the same number of predicted sample indices, resulting in a large data processing volume. Furthermore, some reference sample indices are in numerical form (e.g., the contrast and entropy of the calculated text image samples) and have a wide value range, which is detrimental to model training. In one embodiment, the calculated image index information can be merged to obtain reference sample indices. Merging multiple calculation results can reduce the number of reference sample indices and lower the training difficulty of the model. Specifically, the step "calculate image indices based on the image attribute information of the text image samples, and determine the reference sample indices of the text image samples based on the calculation results" can include:

[0055] Image indices are calculated based on the image attribute information of the text image samples to obtain at least one image index information.

[0056] At least one image indicator is merged to obtain a reference sample indicator for the text image sample.

[0057] For example, image metrics can be calculated based on the attribute information of text image samples to obtain at least one image metric. All the obtained image metric information can be concatenated into a multidimensional array or tensor, and this multidimensional array or tensor can be used as a reference sample metric.

[0058] Optionally, the image index information in numerical form can be merged, for example, by concatenating them into an array. The resulting array, along with other non-numerical forms, such as one-dimensional tensors and multi-dimensional tensors, can be used as reference index information.

[0059] Optionally, the image indicator information can be merged according to the indicator value expression type of the image indicator to obtain a reference sample indicator, that is, the step "merging at least one image indicator information to obtain a reference sample indicator for the text image sample" can be specifically:

[0060] Obtain the value representation type of at least one image metric;

[0061] Based on the index value expression type, at least one image index information is merged to obtain the reference sample index of the text image sample.

[0062] The index value expression type can represent the expression type of image index information, such as numerical value or tensor (one-dimensional or multi-dimensional).

[0063] For example, each image indicator corresponds to an indicator value expression type. Based on the image indicator, the corresponding indicator value expression type is obtained. The image indicator information with the indicator value expression type of numerical value and one-dimensional tensor is concatenated to obtain a one-dimensional tensor containing more image indicator information about the text image sample. The image indicator information with the indicator value expression type of two-dimensional tensor is concatenated to obtain a three-dimensional tensor about the text image sample. The obtained one-dimensional tensor, three-dimensional tensor, and other indicator value expression types (e.g., three-dimensional tensor, etc.) image indicator information are used as reference sample indicators.

[0064] 103. By using a feature extraction model to perform image feature extraction processing on text image samples, the image feature information of the text image samples is obtained.

[0065] The feature extraction model can be a neural model used to extract image features from text image samples.

[0066] Image feature extraction can be a process of analyzing and transforming text image samples to extract characteristic information from the text image samples. The image feature information can be the information obtained through image feature extraction.

[0067] For example, the feature extraction model could include a Convolutional Recurrent Neural Network (CRNN). The CRNN network contains a Convolutional Neural Network (CNN) and a Recurrent Neural Network (RNN). The CNN network performs convolution processing on the text image samples to obtain feature maps of the text image samples. The RNN network then extracts features from the feature maps to obtain the image feature information of the text image samples.

[0068] Optionally, the feature extraction model can also extract image features using other neural networks. For example, the DenseNet network can be used to encode text image samples, mapping the text image samples to image feature information that can represent the text image samples.

[0069] 104. Based on image feature information, the feature extraction model is used to extract attention features from text image samples to obtain attention feature information of the attention context of the text image samples.

[0070] The text image sample may include multiple image regions, each of which may have corresponding regional feature information. The image feature information of the text image sample may include the regional feature information of each image region. The attention feature information may be information with contextual information obtained by fusing the regional feature information of the image region and the regional feature information of the associated image region for each image region.

[0071] For example, specifically for each image region, the similarity between the regional feature information of the image region and the regional feature information of the associated image region can be used as the weight of the associated image region. Based on the weight, the regional feature information of the image region and the associated image region can be weighted and summed to obtain the attention feature information of the attention context for each image region.

[0072] Attention feature information of the text image to be recognized can be obtained by using the attention feature information of the attention context information corresponding to each image region.

[0073] Among them, the associated image region can be an image region associated with the image region, such as an adjacent image region of the image region. The associated image region can also be all image regions of the text image sample, or other image regions in the text image sample.

[0074] In one embodiment, the image feature information may include image feature vectors, which may contain regional feature vectors for each image region. The similarity between the regional feature vectors of the image regions and those of associated image regions can be obtained based on the distance between them. Furthermore, the image feature vectors can be mapped into an attention space, and the similarity between each image region and its associated image region can be determined based on the distance between vectors in the attention space. Attention feature information with contextual information is then obtained based on this similarity. Specifically, the step "using a feature extraction model to extract attention features from text image samples based on image feature information to obtain attention feature information of the text image samples with attentional contextual information" may include:

[0075] The image feature information is processed by attention space mapping to obtain the spatial vector of each image region in the text image sample in the attention space. The spatial vector includes query vector, content vector and key vector.

[0076] For each image region, the similarity between the image region and the associated image region is calculated based on the distance between the query vector of the image region and the key vector of the associated image region.

[0077] For each image region, the content vectors of the image region and the associated image region are fused based on the similarity between the key vector of the image region and the associated image region to obtain attention feature information that focuses on contextual information.

[0078] The query vector, key vector, and content vector can be spatial vectors obtained by linearly transforming the image feature vectors according to different attention network parameters. The attention network parameters can be the network parameters in the feature extraction model.

[0079] For example, the attention network parameters can specifically include the first attention network parameter W. Q Second attention network parameters W K And the third attention network parameter W V The image feature vector is mapped based on the parameters of the first attention network to obtain the query vector of the text image sample, denoted as Query, or Q for short, Q = Γ·W Q The query vector of the i-th image region in the text image sample is denoted as Q. iBased on the parameters of the second attention network, the image feature vector is mapped to obtain the key vector of the text image sample, denoted as Key, or K for short, K ​​= Γ·W K The query vector of the i-th image region in the text image sample is denoted as K. i The image feature vector is mapped based on the parameters of the third attention network to obtain the content vector of the text image sample, denoted as Value, or V for short, V = Γ·W. V Let Vi be the query vector of the i-th image region in the text image sample, and let Vi be the spatial vector of the text image sample in the attention space. 。

[0080] Calculate the distance between the query vector corresponding to the i-th image region and the key vector corresponding to the associated image region j. For example, the query vector and the key vector can be multiplied by a dot product, such as Q. i ·K j The similarity between image region i and image region j is obtained.

[0081] The same processing is performed on each image region in the text image sample to obtain the similarity between each image region and its corresponding associated region.

[0082] The content vectors of the image region and each associated image region are weighted and summed based on the similarity between the image region and each associated image region to obtain the regional attention feature information of the image region's attention context. Based on the regional attention feature information of each image region, the attention feature information of the text image sample is obtained.

[0083] In one embodiment, an initial similarity matrix for the text image samples can be obtained based on the initial similarity between each image region and its associated image region. This matrix can be denoted as SCORE0, where the element 'score' in the i-th row and i-th column of the initial similarity matrix is ​​the first element. ij The element `score` located in the j-th row and i-th column represents the initial similarity between image region i and image region j (where image region j is a related image region of image region i). ji This can represent the initial similarity between image region j and image region i.

[0084] Typically, for each image region, the text content between distant image regions is almost unrelated and has no impact on the predicted text content of the image region. Therefore, for each image region, a corresponding window matrix can be set for the obtained initial similarity matrix to retain the similarity of the image regions within the area indicated by the window matrix and mask the similarity at other locations.

[0085] The window matrix can be set to 0 for the window position and -∞ or other very large negative numbers, such as 10, for other positions. -16 The initial similarity matrix is ​​then added to the window matrix. This causes the similarity scores of non-window locations in the initial similarity matrix to be set to a large negative number due to the addition of a large negative number. Normalization maps the similarity scores of non-window locations to 0. This normalization process can be summarized by the formula: The first similarity matrix is ​​obtained by normalizing each image region. The similarity between the image region and the associated image region can be determined based on the first similarity matrix. It can be understood that after adding the window matrix, since the similarity of the image regions at non-window positions is 0, the associated image region corresponding to each image region is actually the image region at the window position of the window matrix.

[0086] The window matrix can be a matrix of the same type as the initial similarity matrix. It can be used to retain the similarity within the region indicated by the window matrix (which can be called the window position) and to mask the similarity at other positions, such as setting the similarity at other positions to -∞. The window position can be set according to each image region. For example, for image region i, the window position can be score. ii score ii+1 and score ii+1 This means preserving the similarity between three adjacent image regions while masking the similarity with other image regions.

[0087] Ideally, if image region i has a high similarity to image region j, then image region j has a high similarity to image region i. In other words, ideally, the similarity matrix is ​​a symmetric matrix, and the transpose matrix SCORE can be obtained by interchanging the rows and columns of the first similarity matrix. T Add the transpose matrix and the first similarity matrix to obtain the similarity matrix SCORE = SCORE1 + SCORE T The similarity matrix is ​​a symmetric matrix, and the elements of the similarity matrix are scores. ij =score ji .

[0088] The similarity matrix can be used to determine the similarity between the feature vector of each image region and the feature vector of the associated region. For example, the similarity between image region i and image region j is the score of the element SCORE in the similarity matrix. ij .

[0089] Multiplying the content vector V of the text image sample by the similarity matrix yields the attention feature information C of the text image sample with contextual information, C = V·SCORE, where V is the content vector of the text image sample and SCORE is the similarity matrix.

[0090] In one embodiment, interference information of the text image samples can be added to improve the training effect of the feature extraction model. Specifically, the step "based on image feature information, the feature extraction model extracts attention features from the text image samples to obtain attention feature information of the attention context of the text image samples" can include:

[0091] The image feature information is masked by a feature extraction model to obtain the masked image feature information of the text image sample.

[0092] Attention feature extraction is performed on the masked image feature information to obtain the attention feature information of the attention context of the text image sample.

[0093] Masking can be a method of masking or selecting some features in the image features of text image samples to increase the noise of the training samples of the feature extraction model, so that the trained feature extraction model has a more generalized feature extraction capability.

[0094] For example, one could mask certain features in the image feature information of a text image sample to increase the interference information of the sample, obtain masked image feature information, and then extract attention features from the masked image feature information to obtain attention feature information of the attention context of the text image sample.

[0095] Optionally, step 103 can involve image feature extraction using an attention mechanism included in the feature extraction model. The attention mechanism is a special structure embedded in a machine learning model used to automatically learn and calculate the contribution of input data to output data. To improve the accuracy of feature extraction, the attention mechanism can be a multi-layer attention mechanism. Specifically, step "based on image feature information, the feature extraction model extracts attention features from the text image samples to obtain attention feature information about the context of the text image samples" can include:

[0096] Image feature information is used as input feature information for a multi-layer feature extraction mechanism;

[0097] Attention feature information of text image samples is obtained by sequentially extracting attention features from the input feature information through a multi-layer feature extraction mechanism.

[0098] For example, the multi-layer attention mechanism can be ordered in a certain way. Image feature information is used as the input feature information of the first layer of the multi-layer attention mechanism. The first layer of attention mechanism extracts attention features from the text image sample to obtain the first attention feature information. The first attention feature information is used as the input feature information of the second layer of attention mechanism. The second layer of attention mechanism extracts attention features from the input feature information to obtain the second attention feature information. And so on, until the feature information output by the last layer of attention mechanism is used as the attention feature information.

[0099] Optionally, each attention mechanism layer can include a multi-head attention mechanism. The attention feature information of that layer is obtained based on the sub-attention feature information obtained from the multi-head attention mechanism. Specifically, the step "to extract attention features from the input feature information sequentially through a multi-layer feature extraction mechanism to obtain the attention feature information of the text image sample's attention context" can include:

[0100] By using a multi-layer feature extraction mechanism based on a multi-head attention mechanism, attention features are extracted sequentially from the input feature information to obtain sub-attention feature information under each head attention mechanism.

[0101] By fusing the sub-attention feature information of each attention mechanism through a multi-layer feature extraction mechanism, attention feature information of the attention context of the text image sample is obtained.

[0102] For example, specifically, attention features can be extracted from image feature information through a multi-head attention mechanism at each layer, resulting in sub-attention feature information under each attention feature mechanism. Each sub-attention feature information is then concatenated to obtain concatenated attention feature information. The concatenated attention feature information is then dimensionally transformed to be the same as the dimension of the input image feature information. The attention feature information obtained from this layer of attention mechanism is then output, and the attention feature information is used as the input feature information for the next layer of attention feature mechanism.

[0103] 105. Based on the attention feature information of text image samples, predict the predicted sample index of text image samples.

[0104] Among them, the predicted sample index can be an index that is predicted based on attention feature information.

[0105] For example, it could be a predicted sample index that matches the reference sample index based on attention feature information.

[0106] In one embodiment, attention feature information can be processed in different ways according to the indicator type of the reference sample indicator to obtain a predicted sample indicator that matches the reference sample indicator. That is, the step "predicting the predicted sample indicator of the text image sample based on the attention feature information of the text image sample" can specifically be:

[0107] Determine the feature processing method corresponding to each indicator type;

[0108] For each indicator type, the attention feature information is processed using the corresponding feature processing method to obtain the prediction sample corresponding to each indicator type.

[0109] The indicator type can be the same as the indicator type of the reference sample, such as a one-dimensional indicator, a two-dimensional indicator, or a restoration indicator. A one-dimensional indicator means that the image indicator information it contains is a one-dimensional tensor (the value is a special one-dimensional tensor). A restoration indicator means that the image indicator information it contains is an image tensor.

[0110] For example, it could involve determining the feature processing method corresponding to each indicator type, such as dimensional transformation or image restoration. Based on each indicator type, the attention features could be processed using the corresponding feature processing method to obtain the predicted sample indicator that matches the reference sample indicator.

[0111] If the indicator type is a one-dimensional indicator or a two-dimensional indicator, then the attention feature mechanism is subjected to dimensionality transformation to obtain a tensor with the same size as the reference sample indicator, and the obtained tensor is used as the reference sample indicator.

[0112] Feature extraction models involve downsampling during the extraction of attention feature information from text image samples. During the learning process, the model discards information deemed useless. Increasing the number of channels cannot solve this problem. Feature extraction models cannot effectively store two-dimensional image information in channels. This is due to the network optimization objective. Affected by the task and training data distribution, some effective information is considered redundant by the network. Image restoration tasks can retain image information to the maximum extent, forcing the network to retain more effective information.

[0113] If the indicator type is a three-dimensional tensor corresponding to a text image sample, then the corresponding processing method is image reconstruction processing. The step "processing the attention feature information using the corresponding feature processing method to obtain the prediction sample corresponding to each indicator type" can specifically include:

[0114] The attention feature information is transposed and convolutional to obtain the processed attention feature information.

[0115] The attention feature information is normalized to obtain the predicted sample index.

[0116] Among them, transposed convolution processing can be used to deconvolve the attention feature information to perform image restoration processing based on the attention feature information, and obtain a tensor with the same size as the text image sample.

[0117] For example, one could perform transposed convolution on the attention feature information to obtain processed attention feature information, and then normalize the attention feature information to obtain the predicted sample index.

[0118] Optionally, the attention feature information can be subjected to multiple transposed convolutions, and the processed attention feature information obtained from the transposed convolutions can be batch-normalized and processed using activation functions. The specific network structure for image restoration processing of the attention feature information is as follows: Figure 3 As shown, attention feature information is processed alternately through transposed convolutional layers and activation layers, and finally the processed attention feature information is output. The attention feature information is then normalized to obtain the predicted sample label.

[0119] 106. Based on the predicted sample indicators and the corresponding reference sample indicators, the feature extraction model is trained so as to extract the attention feature information of the text image to be recognized and perform image text recognition.

[0120] For example, the feature extraction model can be trained by the error between the predicted sample index and the corresponding reference sample index, and the network parameters of the feature extraction model can be adjusted to obtain the trained feature extraction model.

[0121] After obtaining the trained feature extraction model, image features can be extracted from the text sample to be recognized using the trained feature extraction model to obtain image feature information. Attention feature extraction is then performed on the image feature information to obtain the attention feature information of the text image to be recognized. Based on the attention feature information, the text content in the text image to be recognized can be predicted.

[0122] Optionally, after training the feature extraction model using predicted sample metrics and reference sample metrics to obtain a pre-trained feature extraction model, the model can be fine-tuned using labeled target text image samples to obtain a trained feature extraction model. This improves the feature extraction capability of the trained model. Specifically, the step "training the feature extraction model based on predicted sample metrics and corresponding reference sample metrics" can be followed by:

[0123] Obtain target text image samples, each carrying a sample label;

[0124] Image feature extraction is performed on the target text image sample using a pre-trained feature extraction model to obtain the image feature information of the target text image sample;

[0125] Based on the image feature information of the target text image sample, the pre-trained feature extraction model extracts attention features of the target text image sample to obtain the attention feature information of the attention context of the target text image sample.

[0126] The prediction result of the target text image sample is obtained by using a text recognition model to predict based on the attention feature information of the target text image sample;

[0127] The text recognition model and the pre-trained feature extraction model are trained based on the sample labels and prediction results to obtain the trained text recognition model and the trained feature extraction model, so as to recognize the text content of the text image to be recognized through the text recognition model.

[0128] For example, specifically, attention feature information of target text image samples can be extracted by a pre-trained feature extraction model, and the target text image samples can be predicted by a text recognition model based on the attention feature information to obtain the prediction result. The pre-trained feature extraction model and the text recognition model can be trained according to the prediction result and the sample label to obtain the trained feature extraction model and the trained text recognition model. The trained feature extraction model can be used to extract features from the text image to be recognized to obtain attention feature information, and the prediction result can be obtained by the trained text recognition model based on the attention feature information.

[0129] As described above, this embodiment of the application obtains text image samples; calculates image metrics based on the image attribute information of the text image samples, and determines reference sample metrics based on the calculation results; performs image feature extraction processing on the text image samples using a feature extraction model to obtain image feature information of the text image samples; extracts attention features from the text image samples based on the image feature information using the feature extraction model to obtain attention feature information of the attention context of the text image samples; predicts predicted sample metrics of the text image samples based on the attention feature information of the text image samples; and trains the feature extraction model based on the predicted sample metrics and the corresponding reference sample metrics, so as to extract the attention feature information of the text image to be recognized for image text recognition through the trained feature extraction model. This scheme trains the feature extraction model using reference sample metrics and predicted sample metrics, and can utilize a large number of unlabeled text image samples to train the feature extraction model, thereby enhancing the training effect of the feature extraction model.

[0130] Based on the above embodiments, the following examples will provide further detailed explanations.

[0131] This embodiment will be described from the perspective of a text recognition device, which can be integrated into a computer device, such as a server or a terminal.

[0132] This application provides a text recognition method, such as... Figure 4 As shown, the text recognition method can be divided into feature extraction tasks, upstream tasks, and downstream tasks. The specific process is as follows:

[0133] 1. Feature Extraction Task: The feature extraction model extracts features from the input image to obtain attention feature information, which is then used for upstream and downstream tasks.

[0134] 201. The server obtains text image samples, performs image feature extraction processing on the text image samples using a feature extraction model, and obtains the image feature information of the text image samples.

[0135] For example, the server can obtain text image samples from the database, encode the text image samples through the DenseNet network of the feature extraction model, and map the text image samples into image feature information Γ that can represent the text image samples. The image feature information can be a feature embedding sequence.

[0136] 202. The server extracts attention features from image features using a feature extraction model to obtain attention feature information of the attention context of the text image sample.

[0137] For example, the server can use a feature extraction model to map the image feature vector based on the parameters of the first attention network with an attention mechanism to obtain the query vector of the text image sample, denoted as Query, or Q for short, Q = Γ·W Q The query vector of the i-th image region in the text image sample is denoted as Q. i Based on the parameters of the second attention network, the image feature vector is mapped to obtain the key vector of the text image sample, denoted as Key, or K for short, K ​​= Γ·W K The query vector of the i-th image region in the text image sample is denoted as K. i The image feature vector is mapped based on the parameters of the third attention network to obtain the content vector of the text image sample, denoted as Value, or V for short, V = Γ·W. V The query vector of the i-th image region in the text image sample is denoted as V. i .

[0138] Calculate the distance between the query vector corresponding to the i-th image region and the key vector corresponding to the associated image region j. For example, the query vector and the key vector can be multiplied by a dot product, such as Q.i ·K j The similarity between image region i and associated image region j is obtained.

[0139] The same processing is performed on each image region in the text image sample to obtain the similarity between each image region and its corresponding associated region.

[0140] The content vectors of the image region and each associated image region are weighted and summed based on the similarity between the image region and each associated image region to obtain the region attention feature information of the attention context of the image region. Based on the region attention feature information of each image region, the attention feature information of the text image sample is obtained.

[0141] Optionally, to improve feature extraction capabilities, the feature extraction model may include a feature extraction network, the network structure of which may be as follows: Figure 1 As shown, after obtaining the attention feature information through the attention mechanism, the data can be normalized through the first BN layer to obtain... Then, the data is fed into a feedforward network. This network extracts attention features and reduces feature redundancy. The feedforward network typically consists of two fully connected layers and one activation layer. After passing through the feedforward network, C* is normalized by a second batch normalization layer. Finally, the attention features obtained from this feature extraction network are output, resulting in the final data. ,in , , ,and Here are the network parameters for the fully connected layer of the feedforward network, where the Batch Normalization (BN) layer contains the following BN algorithm:

[0142] For any input sequence Calculate the mean and variance.

[0143]

[0144]

[0145] Normalize the input X, where It is a very small constant to prevent invalid calculations caused by a variance of 0.

[0146]

[0147] Through learnable weights With bias Get the final output .

[0148]

[0149] Wherein, input sequence It can be image feature information containing a set of text image samples. The text image samples captured by the feature extraction model in one training session constitute a set of text image samples. This refers to relevant information about one of the text image samples in a set of text image samples, such as attention feature information of the text image sample obtained through an attention feature mechanism.

[0150] Understandably, once the feature extraction model is trained, the mean and variance can be calculated based on the mean and variance obtained during training, such as by averaging or moving average.

[0151] Optionally, to improve the ability of the feature extraction model to extract attention features, the feature extraction model can include multiple layers of attention mechanisms, which can be distributed in different feature extraction networks, for example, such as... Figure 5 The structure shown uses the output of the previous feature extraction network as the input feature information for the next feature extraction network, and uses the output of the last feature extraction network as the attention feature information.

[0152] II. Upstream Task: Calculate the reference sample index of the text image sample using the image attribute information of the text image sample, and predict the predicted sample index based on the attention feature information. Train the feature extraction model based on the reference sample index and the predicted sample index.

[0153] 203. The server calculates image metrics based on the image attribute information of the text image sample, including at least one image metric.

[0154] For example, the server can calculate image index information (the image index information is the calculation result obtained by calculating the image index based on the image attribute information) by obtaining image attribute information for at least one feature. For example, for color features, the server can calculate the color histogram (e.g., RGB color histogram, HSV color histogram, grayscale histogram) and / or color set and color moments of the text image sample by calculating the color values ​​of each pixel in different channels of the text image sample; and for the overall features of the text image sample, the server can calculate the color values ​​and other attribute information of each pixel in different channels of the text image sample to obtain the image index information of the text image sample.

[0155] 204. The server merges the image indicator information according to the indicator value expression type of the image indicator to obtain the reference sample indicator of the text image sample. The reference sample indicator includes one-dimensional indicator, two-dimensional indicator and restoration indicator.

[0156] For example, each image indicator could correspond to an indicator value representation type. The server obtains the corresponding indicator value representation type based on the image indicator, and concatenates the image indicator information with numerical values ​​and one-dimensional tensors to obtain a one-dimensional tensor containing more image indicator information for the text image sample. This one-dimensional tensor is then used as a one-dimensional indicator. Similarly, concatenating the image indicator information with two-dimensional indicator value representation types yields a three-dimensional tensor for the text image sample, which is then used as a two-dimensional indicator.

[0157] And the three-dimensional tensor obtained by merging the image index information calculated from the color values ​​of text image samples in different channels is used as the restoration index.

[0158] One-dimensional indicators, two-dimensional indicators, and restoration indicators are used as reference sample indicators for text image samples.

[0159] 205. The server processes the attention feature information according to the corresponding processing method of the indicator type of the reference sample indicator to obtain the prediction indicator information, which includes one-dimensional prediction indicator, two-dimensional prediction indicator and prediction restoration indicator.

[0160] If the indicator type is one-dimensional or two-dimensional, the attention feature information can be dimensionally transformed using different network structures to obtain a tensor of the same size as the reference sample indicator. This tensor is then used as the reference sample indicator. For example, for a one-dimensional indicator, the attention feature can be dimensionally transformed using two fully connected layers: To obtain the predicted sample index For one-dimensional metrics, the attention features are transformed in dimension through three fully connected layers: ,in, , These are network parameters.

[0161] If the indicator type is a restoration indicator, the corresponding processing method is image restoration processing. The specific network structure for image restoration processing of attention feature information is as follows: Figure 3 As shown, attention feature information is processed alternately through transposed convolutional layers and activation layers, and finally the processed attention feature information is output. The processed attention feature information is then normalized to obtain the predicted sample label.

[0162] After the first transposed convolutional layer, we can obtain: Where C represents the attention feature information, which is the final processed attention feature information output. The processed attention feature information is then normalized. ,in, For predicting sample indicators.

[0163] Optionally, the activation function in the last activation layer can be the Tanh function, and the activation function in the BN& activation layer can be the ReLU function.

[0164] Optionally, the number of transposed convolutional layers and the number of BN& activation layers in the network structure for image restoration processing can be flexibly set as needed, and are not limited here.

[0165] 206. The server trains the feature extraction model using reference sample metrics and predicted sample metrics to obtain a pre-trained feature extraction model.

[0166] For example, the MSE loss function can be used to calculate the error between the predicted sample index and the corresponding reference sample index to train the feature extraction model, and the network parameters of the feature extraction model can be adjusted to obtain a pre-trained feature extraction model.

[0167] 3. Downstream task: Load the pre-trained feature extraction model, connect it to the text recognition model, the text recognition model predicts the attention feature information extracted from the target text image sample based on the pre-trained features, obtains the prediction result, and trains the initial feature extraction model based on the prediction result and sample labels to obtain the trained feature extraction model.

[0168] 207. The server performs image feature extraction and attention feature extraction on the target text image sample through a pre-trained feature extraction model to obtain the attention feature information of the target text image sample's attention context.

[0169] For example, a pre-trained feature extraction mechanism can be used to extract image features from the target text image sample to obtain the image feature information of the target text image sample. Then, attention feature extraction can be performed on the image feature information to obtain the attention feature information of the attention context of the target text image sample.

[0170] 208. The server uses the text recognition model to make predictions based on the attention feature information of the target text image samples, obtains the prediction results, and trains the pre-trained feature extraction model and text recognition model based on the prediction results and sample labels, to obtain the trained feature extraction model and trained text recognition model.

[0171] For example, a text recognition model can predict target text image samples based on attention feature information to obtain prediction results. Based on the prediction results and sample labels, a pre-trained feature extraction model and a text recognition model can be trained to obtain a trained feature extraction model and a trained text recognition model. The trained feature extraction model can then extract features from the text image to be recognized to obtain attention feature information. Finally, the trained text recognition model can use the attention feature information to obtain prediction results.

[0172] The text recognition model can be a classifier model, for example, such as... Figure 6 As shown, the text recognition model based on the CRNN structure inputs attention feature information into the fully connected classifier of the text recognition model to predict the probability that each image region in the target text image is each character in the dictionary, and obtains the prediction result. Based on the prediction result and sample labels, the loss is calculated based on CTCLOss, and the pre-trained feature extraction model and text recognition model are trained.

[0173] Or like Figure 6 As shown, the text recognition model based on the Attention mechanism inputs the attention feature information into an LSTM decoder containing the Attention mechanism to predict the probability that each image region in the target text image is each character in the dictionary, and obtains the prediction result. Based on the prediction result and sample labels, the loss is calculated based on cross-entropy (CE Loss) to train the pre-trained feature extraction model and the text recognition model.

[0174] As can be seen from the above, in this embodiment, the server acquires text image samples, performs image feature extraction processing on the text image samples using a feature extraction model, and obtains image feature information of the text image samples; it then performs attention feature extraction on the image feature information using the feature extraction model to obtain attention feature information of the attention context of the text image samples; it calculates image indicators based on the image attribute information of the text image samples, obtaining at least one image indicator; the server merges the image indicator information according to the indicator value expression type of the image indicators to obtain reference sample indicators of the text image samples, the reference sample indicators including one-dimensional indicators, two-dimensional indicators, and restoration indicators; and it processes the attention feature information according to the processing method corresponding to the indicator type of the reference sample indicators to obtain predicted indicator information, the predicted indicator information including one-dimensional predicted indicators, two-dimensional predicted indicators, and restoration indicators. The scheme employs a pre-trained feature extraction model, which is trained using reference sample metrics and predicted sample metrics. This model is then used to extract image features and attention features from target text image samples, resulting in attention feature information about the target text image samples' attention context. A text recognition model then predicts these attention feature information based on the target text image samples, yielding prediction results. Finally, the pre-trained feature extraction model and the text recognition model are trained using the prediction results and sample labels, resulting in a trained feature extraction model and a trained text recognition model. This approach, by using reference sample metrics and predicted sample metrics to train the feature extraction model, can utilize a large number of unlabeled text image samples to enhance the training effect of the feature extraction model.

[0175] To facilitate better implementation of the text recognition method provided in the embodiments of this application, a text recognition device is also provided in one embodiment. The meanings of the terms are the same as in the text recognition method described above, and specific implementation details can be found in the description of the method embodiments.

[0176] This text recognition device can be integrated into a computer device, such as... Figure 7 As shown, the text recognition device may include: an acquisition unit 301, a calculation unit 302, a first feature extraction unit 303, a second feature extraction unit 304, a prediction unit 305, and a training unit 306, as detailed below:

[0177] (1) Acquisition Unit 301

[0178] Acquisition unit 301: Used to acquire text image samples.

[0179] (2) Calculation unit 302

[0180] Calculation unit 302: used to calculate image indicators based on the image attribute information of text image samples, and determine reference sample indicators of text image samples based on the calculation results.

[0181] Optionally, the calculation unit 302 may include a calculation subunit and a merging subunit, specifically:

[0182] Calculation subunit: used to calculate image indicators based on the image attribute information of text image samples, and obtain at least one image indicator information;

[0183] Merging subunit: Used to merge at least one image indicator information to obtain a reference sample indicator for the text image sample.

[0184] Optionally, the merged subunit may include an acquisition module and a merge module, specifically:

[0185] Acquisition module: Used to acquire the value representation type of at least one image indicator;

[0186] Merging module: Used to merge at least one image indicator information according to the indicator value expression type to obtain the reference sample indicator of the text image sample.

[0187] (3) First feature extraction unit 303

[0188] First feature extraction unit 303: used to perform image feature extraction processing on text image samples through a feature extraction model to obtain image feature information of text image samples.

[0189] (4) Second feature extraction unit 304

[0190] The second feature extraction unit 304 is used to extract attention features from text image samples based on image feature information through a feature extraction model, thereby obtaining attention feature information of the attention context of the text image samples.

[0191] Optionally, the second feature extraction unit 304 may include a mapping subunit, a similarity calculation subunit, and a first fusion subunit, specifically:

[0192] Mapping subunit: Used to perform attention space mapping on image feature information to obtain the spatial vector corresponding to each image region in the attention space in the text image sample. The spatial vector includes query vector, content vector and key vector.

[0193] Similarity calculation subunit: For each image region, it calculates the similarity between the image region and the associated image region based on the distance between the query vector of the image region and the key vector of the associated image region;

[0194] The first fusion subunit is used to fuse the image feature information of each image region and the associated image region based on the similarity between the key vector of the image region and the associated image region, so as to obtain attention feature information that focuses on context information.

[0195] Optionally, the second feature extraction unit 304 may include a masking subunit and an extraction subunit, specifically:

[0196] Masking subunit: Used to mask image feature information through a feature extraction model to obtain masked image feature information of text image samples;

[0197] Extraction sub-unit: used to extract attention features from the masked image feature information to obtain attention feature information of the attention context of the text image sample.

[0198] Optionally, the second feature extraction unit 304 may include both a sub-unit and a first feature extraction sub-unit, specifically:

[0199] As a sub-unit: used to use image feature information as input feature information for a multi-layer feature extraction mechanism;

[0200] The first feature extraction subunit is used to sequentially extract attention features from the input feature information through a multi-layer feature extraction mechanism, thereby obtaining attention feature information of the attention context of the text image sample.

[0201] Optionally, the second feature extraction unit 304 may include a second feature extraction subunit and a second fusion subunit, specifically:

[0202] The second feature extraction subunit is used to sequentially extract attention features from the input feature information through a multi-layer feature extraction mechanism based on a multi-head attention mechanism, so as to obtain the sub-attention feature information under each head attention mechanism.

[0203] The second fusion subunit is used to fuse the sub-attention feature information of each attention mechanism through a multi-layer feature extraction mechanism to obtain the attention feature information of the attention context of the text image sample.

[0204] (5) Prediction unit 305:

[0205] Prediction unit 305: Used to predict the prediction sample index of text image samples based on the attention feature information of text image samples.

[0206] Optionally, the prediction unit 305 may include a determination subunit and a processing subunit, specifically:

[0207] Determine the sub-unit: used to determine the feature processing method corresponding to each indicator type;

[0208] Processing subunit: For each indicator type, the attention feature information is processed using the corresponding feature processing method to obtain the predicted sample indicator for each indicator type.

[0209] Optionally, the processing subunit may include a transposed convolution module and a normalization module, specifically:

[0210] Transposed convolution module: Used to perform transposed convolution processing based on attention feature information to obtain processed attention feature information;

[0211] Normalization module: Used to normalize attention feature information to obtain predicted sample indicators.

[0212] (6) Training Unit 306:

[0213] Training unit 306: Used to train the feature extraction model based on the predicted sample indicators and the corresponding reference sample indicators, so as to extract the attention feature information of the text image to be recognized through the trained feature extraction model for image text recognition.

[0214] Optionally, the text recognition device may include a sample acquisition unit, a third feature extraction unit, a fourth feature extraction unit, a result prediction unit, and a supervised training unit, specifically:

[0215] Sample acquisition unit: used to acquire target text image samples, which carry sample labels;

[0216] The third feature extraction unit is used to perform image feature extraction processing on the target text image sample through a pre-trained feature extraction model to obtain the image feature information of the target text image sample.

[0217] The fourth feature extraction unit is used to extract attention features from the target text image sample based on the image feature information of the target text image sample through a pre-trained feature extraction model, so as to obtain the attention feature information of the attention context of the target text image sample.

[0218] Result prediction unit: used to predict the target text image sample based on the attention feature information of the target text image sample through the text recognition model, and obtain the prediction result of the target text image sample;

[0219] Supervised training unit: Used to train the text recognition model and the pre-trained feature extraction model based on sample labels and prediction results, so as to obtain the trained text recognition model and the trained feature extraction model, so as to recognize the text content of the text image to be recognized through the text recognition model.

[0220] This application also provides a computer device, which can be a terminal or a server, such as... Figure 8 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0221] The computer device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art will understand that... Figure 8 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0222] The processor 1001 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1002, and by calling data stored in the memory 1002, it performs various functions of the computer device and processes data, thereby performing overall detection of the computer device. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and computer programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1001.

[0223] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1002 may also include a memory controller to provide the processor 1001 with access to the memory 1002.

[0224] The computer equipment also includes a power supply 1003 that supplies power to the various components. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 1003 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0225] The computer device may also include an input unit 1004, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0226] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1001 in the computer device loads the executable files corresponding to the processes of one or more computer programs into the memory 1002 according to the following instructions, and the processor 1001 runs the computer programs stored in the memory 1002 to realize various functions, as follows:

[0227] Obtain text image samples;

[0228] Image indices are calculated based on the image attribute information of the text image samples, and reference sample indices are determined based on the calculation results.

[0229] Image feature extraction is performed on text image samples using a feature extraction model to obtain the image feature information of the text image samples;

[0230] Based on image feature information, attention feature extraction is performed on text image samples using a feature extraction model to obtain attention feature information of the attention context of the text image samples.

[0231] Based on the attention feature information of text image samples, predict the predicted sample index of text image samples.

[0232] Based on the predicted sample indicators and the corresponding reference sample indicators, the feature extraction model is trained so that the attention feature information of the text image to be identified can be extracted through the trained feature extraction model to perform image text recognition.

[0233] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0234] As can be seen from the above, the computer device in this embodiment can acquire text image samples; calculate image indicators based on the image attribute information of the text image samples, and determine reference sample indicators for the text image samples based on the calculation results; perform image feature extraction processing on the text image samples using a feature extraction model to obtain image feature information of the text image samples; extract attention features from the text image samples based on the image feature information using the feature extraction model to obtain attention feature information of the attention context of the text image samples; predict the predicted sample indicators of the text image samples based on the attention feature information of the text image samples; and train the feature extraction model based on the predicted sample indicators and the corresponding reference sample indicators, so as to extract the attention feature information of the text image to be recognized for image text recognition through the trained feature extraction model. This scheme trains the feature extraction model using reference sample indicators and predicted sample indicators, and can utilize a large number of unlabeled text image samples to train the feature extraction model, thereby enhancing the training effect of the feature extraction model.

[0235] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0236] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0237] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute any of the text recognition methods provided in embodiments of this application.

[0238] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0239] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0240] Since the computer program stored in the computer-readable storage medium can execute any of the text recognition methods provided in the embodiments of this application, it can achieve the beneficial effects that any of the text recognition methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0241] The foregoing has provided a detailed description of a text recognition method, apparatus, computer device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A text recognition method, characterized in that, include: Obtain text image samples; Image metrics are calculated based on the image attribute information of the text image sample, and reference sample metrics are determined based on the calculation results; wherein, the reference sample metrics include reference sample metrics of at least two metric types; The text image sample is processed by a feature extraction model to extract image features, thereby obtaining the image feature information of the text image sample. Based on the image feature information, the feature extraction model extracts attention features from the text image sample to obtain attention feature information of the attention context of the text image sample. Based on the attention feature information of the text image samples, predict the predicted sample index of the text image samples, including: determining the feature processing method corresponding to each index type; for each index type, processing the attention feature information using the corresponding feature processing method to obtain the predicted sample index matching the reference sample index corresponding to each index type. The feature extraction model is trained based on the predicted sample index and the corresponding reference sample index, so as to extract the attention feature information of the text image to be recognized and perform image text recognition through the trained feature extraction model.

2. The method according to claim 1, characterized in that, The step of calculating image indices based on the image attribute information of the text image sample, and determining reference sample indices for the text image sample based on the calculation results, includes: Based on the image attribute information of the text image sample, image indexes are calculated to obtain at least one image index information; The at least one indicator information is merged to obtain the reference sample indicator of the text image sample.

3. The method according to claim 2, characterized in that, The step of merging the at least one indicator information to obtain the reference sample indicator of the text image sample includes: Obtain the value representation type of at least one image metric; The at least one image indicator information is merged according to the indicator value expression type to obtain the reference sample indicator of the text image sample.

4. The method according to claim 1, characterized in that, The feature processing method includes image restoration processing. The step of processing attention feature information using the corresponding feature processing method to obtain the predicted sample index corresponding to each index type includes: Based on the attention feature information, a transposed convolution process is performed to obtain the processed attention feature information; The processed attention feature information is normalized to obtain the predicted sample index.

5. The method according to claim 1, characterized in that, The image feature information includes an image feature vector. The attention feature extraction model extracts attention features from the text image sample based on the image feature information to obtain attention feature information of the text image sample's attention context, including: The image feature vector is subjected to attention space mapping processing to obtain the spatial vector corresponding to each image region in the text image sample in the attention space. The spatial vector includes a query vector, a content vector, and a key vector. For each image region, the similarity between the image region and the associated image region is calculated based on the distance between the query vector of the image region and the key vector of the associated image region. For each image region, the content vectors of the image region and the associated image region are fused based on the similarity between the key vector of the image region and the associated image region to obtain the attention feature information of the attention context information.

6. The method according to claim 1, characterized in that, The feature extraction model includes a multi-layer feature extraction mechanism. The feature extraction model extracts attention features from the text image sample based on the image feature information to obtain attention feature information of the text image sample's attention context, including: The image feature information is used as the input feature information for the multi-layer feature extraction mechanism; The attention feature information of the text image sample is obtained by sequentially extracting attention features from the input feature information through the multi-layer feature extraction mechanism.

7. The method according to claim 6, characterized in that, Each feature extraction mechanism includes a multi-head feature extraction mechanism. The attention feature extraction process, performed sequentially on the input feature information through these multi-layer feature extraction mechanisms, yields attention feature information of the text image sample's attention context, including: The multi-layer feature extraction mechanism extracts attention features from the sequentially input feature information based on the multi-head attention mechanism to obtain sub-attention feature information under each head attention mechanism. The multi-layer feature extraction mechanism fuses the sub-attention feature information of each attention mechanism to obtain the attention feature information of the attention context of the text image sample.

8. The method according to claim 1, characterized in that, The step of extracting attention features from the text image sample based on the image feature information using the feature extraction model to obtain attention feature information of the attention context of the text image sample includes: The image feature information is masked by the feature extraction model to obtain the masked image feature information of the text image sample. Attention feature extraction is performed on the masked image feature information to obtain the attention feature information of the attention context information of the text image sample.

9. The method according to claim 1, characterized in that, After training the feature extraction model based on the predicted sample indicators and the corresponding reference sample indicators, the method further includes: Obtain a target text image sample, wherein the text image sample carries a sample label; The target text image sample is processed by a pre-trained feature extraction model to obtain the image feature information of the target text image sample. The pre-trained feature extraction model is trained by the predicted sample index and the corresponding reference sample index. Based on the image feature information of the target text image sample, the pre-trained feature extraction model extracts attention features from the target text image sample to obtain attention feature information of the attention context of the target text image sample. The prediction result of the target text image sample is obtained by making a prediction based on the attention feature information of the target text image sample through a text recognition model; The text recognition model and the pre-trained feature extraction model are trained based on the sample labels and the prediction results to obtain the trained text recognition model and the trained feature extraction model, which are used to perform image text recognition on the text image to be recognized.

10. A text recognition device, characterized in that, include: The acquisition unit is used to acquire text image samples; The calculation unit is used to calculate image indicators based on the image attribute information of the text image sample, and determine reference sample indicators of the text image sample based on the calculation results; wherein, the reference sample indicators include reference sample indicators of at least two indicator types; The first feature extraction unit is used to perform image feature extraction processing on the text image sample through a feature extraction model to obtain the image feature information of the text image sample. The second feature extraction unit is used to extract attention features from the text image sample based on the image feature information using the feature extraction model, so as to obtain attention feature information of the attention context information of the text image sample; The prediction unit is used to predict the predicted sample index of the text image sample based on the attention feature information of the text image sample. The training unit is used to train the feature extraction model based on the predicted sample index and the corresponding reference sample index, so as to extract the attention feature information of the text image to be identified through the trained feature extraction model for image text recognition. The prediction unit includes a determination subunit and a processing subunit; Determine the sub-unit: used to determine the feature processing method corresponding to each indicator type; Processing subunit: Used to process the attention feature information using the corresponding feature processing method for each indicator type, so as to obtain the predicted sample indicator matching the reference sample indicator for each indicator type.

11. The apparatus according to claim 10, characterized in that, The calculation unit includes a calculation subunit and a merging subunit; Calculation subunit: used to calculate image indicators based on the image attribute information of text image samples, and obtain at least one image indicator information; Merging subunit: Used to merge at least one image indicator information to obtain a reference sample indicator for the text image sample.

12. The apparatus according to claim 11, characterized in that, The merging subunit includes an acquisition module and a merging module; Acquisition module: Used to acquire the value representation type of at least one image indicator; Merging module: Used to merge at least one image indicator information according to the indicator value expression type to obtain the reference sample indicator of the text image sample.

13. The apparatus according to claim 10, characterized in that, The second feature extraction unit includes a mapping subunit, a similarity calculation subunit, and a first fusion subunit; Mapping subunit: Used to perform attention space mapping on image feature information to obtain the spatial vector corresponding to each image region in the attention space in the text image sample. The spatial vector includes query vector, content vector and key vector. Similarity calculation subunit: For each image region, it calculates the similarity between the image region and the associated image region based on the distance between the query vector of the image region and the key vector of the associated image region; The first fusion subunit is used to fuse the image feature information of each image region and the associated image region based on the similarity between the key vector of the image region and the associated image region, so as to obtain attention feature information that focuses on context information.

14. The apparatus according to claim 10, characterized in that, The second feature extraction unit includes a mask subunit and an extraction subunit; Masking subunit: Used to mask image feature information through a feature extraction model to obtain masked image feature information of text image samples; Extraction sub-unit: used to extract attention features from the masked image feature information to obtain attention feature information of the attention context of the text image sample.

15. The apparatus according to claim 10, characterized in that, The second feature extraction unit includes a sub-unit and a first feature extraction sub-unit; As a sub-unit: used to use image feature information as input feature information for a multi-layer feature extraction mechanism; The first feature extraction subunit is used to sequentially extract attention features from the input feature information through a multi-layer feature extraction mechanism, thereby obtaining attention feature information of the attention context of the text image sample.

16. The apparatus according to claim 10, characterized in that, The second feature extraction unit includes a second feature extraction subunit and a second fusion subunit; The second feature extraction subunit is used to sequentially extract attention features from the input feature information through a multi-layer feature extraction mechanism based on a multi-head attention mechanism, so as to obtain the sub-attention feature information under each head attention mechanism. The second fusion subunit is used to fuse the sub-attention feature information of each attention mechanism through a multi-layer feature extraction mechanism to obtain the attention feature information of the attention context of the text image sample.

17. The apparatus according to claim 10, characterized in that, The processing subunit includes a transposed convolution module and a normalization module; Transposed convolution module: Used to perform transposed convolution processing based on attention feature information to obtain processed attention feature information; Normalization module: Used to normalize attention feature information to obtain predicted sample indicators.

18. The apparatus according to claim 10, characterized in that, The text recognition device includes a sample acquisition unit, a third feature extraction unit, a fourth feature extraction unit, a result prediction unit, and a supervised training unit; Sample acquisition unit: used to acquire target text image samples, which carry sample labels; The third feature extraction unit is used to perform image feature extraction processing on the target text image sample through a pre-trained feature extraction model to obtain the image feature information of the target text image sample. The fourth feature extraction unit is used to extract attention features from the target text image sample based on the image feature information of the target text image sample through a pre-trained feature extraction model, so as to obtain the attention feature information of the attention context of the target text image sample. Result prediction unit: used to predict the target text image sample based on the attention feature information of the target text image sample through the text recognition model, and obtain the prediction result of the target text image sample; Supervised training unit: Used to train the text recognition model and the pre-trained feature extraction model based on sample labels and prediction results, so as to obtain the trained text recognition model and the trained feature extraction model, so as to recognize the text content of the text image to be recognized through the text recognition model.

19. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform the text recognition method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which is loaded by a processor to perform the text recognition method according to any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text recognition method according to any one of claims 1 to 9.