Text recognition method and device, electronic equipment and readable storage medium

By detecting and correcting text region features in images, and combining global and local features to generate contextual features, the problem of conveniently identifying text in communication device nameplate images is solved, thus improving the accuracy of text recognition.

CN115171116BActive Publication Date: 2026-04-17CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA TELECOM CORP LTD
Filing Date
2022-06-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately extract and recognize text information from images, particularly nameplate images of communication devices.

Method used

By acquiring the text region features and reference text features of the image to be recognized, the regional information of the text region is detected, and after correction, the global and local features of the target text region are generated. Text recognition is then performed by combining the context features.

Benefits of technology

It enables convenient identification of text contained in images and improves the accuracy of text recognition, especially in communication equipment nameplate images, ensuring the integrity and accuracy of text information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171116B_ABST
    Figure CN115171116B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text recognition method and device, electronic equipment and readable storage medium. In the method, the text region features of the to-be-recognized image and the reference text features are obtained. According to the text region features of the to-be-recognized image and the reference text features, the region information of the text region of the to-be-recognized image is detected. The text region is corrected based on the region information to obtain the corrected target text region. According to the global features and the local features of the target text region, the context features of the target text region are generated, and the text recognition of the target text region is performed based on the context features. In this way, the text contained in the picture can be determined conveniently to some extent, and the accuracy of text recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network technology, and in particular relates to a text recognition method, device, electronic device, and readable storage medium. Background Technology

[0002] Currently, it is often necessary to process the text contained in images. For example, during the installation and maintenance of communication equipment, the text contained in the nameplate images of the equipment plays a crucial guiding role in on-site operations and information management within the communication management system. Specifically, the installation location, installation method, maintenance method, and so on, of the communication equipment need to be determined based on the text contained in the nameplate image, i.e., the nameplate information.

[0003] Therefore, how to easily identify the text contained in an image has become a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This invention provides a text recognition method, apparatus, electronic device, and readable storage medium to conveniently determine the text contained in an image.

[0005] In a first aspect, the present invention provides a text recognition method, the method comprising:

[0006] Obtain the text region features and reference text features of the image to be identified;

[0007] Based on the text region features of the image to be identified and the reference text features, the region information of the text region in the image to be identified is detected;

[0008] The text region is corrected based on the region information to obtain the corrected target text region;

[0009] Based on the global and local features of the target text region, the contextual features of the target text region are generated;

[0010] Text recognition is performed on the target text region based on the contextual features.

[0011] In a second aspect, the present invention provides a text recognition device, the device comprising:

[0012] The acquisition module is used to acquire the text region features and reference text features of the image to be recognized;

[0013] The detection module is used to detect the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features;

[0014] A correction module is used to correct the text region based on the region information to obtain the corrected target text region;

[0015] The generation module is used to generate contextual features of the target text region based on the global and local features of the target text region.

[0016] The recognition module is used to perform text recognition on the target text region based on the context features.

[0017] Thirdly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the above-described text recognition method when executing the program.

[0018] Fourthly, the present invention provides a readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the above-described text recognition method.

[0019] In this embodiment of the invention, text region features and reference text features of the image to be recognized are obtained. Based on the text region features and reference text features of the image to be recognized, region information of the text region in the image to be recognized is detected. The text region is corrected based on the region information to obtain the corrected target text region. Context features of the target text region are generated based on the global and local features of the target text region, and text recognition is performed on the target text region based on the context features. Thus, by automatically performing text recognition based on the target text region, the text contained in an image can be conveniently determined. Simultaneously, by performing text region correction and using context features determined based on the corrected global and local features of the target text region for text recognition, the accuracy of text recognition can be improved to a certain extent while conveniently determining the text contained in an image. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of the steps of a text recognition method provided in an embodiment of the present invention;

[0022] Figure 2 This is a schematic diagram of a scenario provided by an embodiment of the present invention;

[0023] Figure 3This is a schematic diagram of a data test provided in an embodiment of the present invention;

[0024] Figure 4 This is a structural diagram of a text recognition device provided in an embodiment of the present invention;

[0025] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Figure 1 This is a flowchart illustrating the steps of a text recognition method provided in an embodiment of the present invention, as follows: Figure 1 As shown, the method may include:

[0028] Step 101: Obtain the text region features and reference text features of the image to be recognized.

[0029] In this embodiment of the invention, the image to be identified can be an image in which text needs to be determined. The image to be identified can be collected according to actual needs; for example, the image to be identified can be a nameplate image, which can be obtained by taking a picture of the nameplate of a communication device using an image acquisition device. The image recognition device performs text recognition on the acquired image to determine the text information contained in the image. The image acquisition device can be a drone with a shooting function, a handheld shooting device, etc. The image processing device can be a device with processing capabilities, such as a computer, a server, etc. The image to be identified collected by the image acquisition device can be sent to the image processing device for recognition. Alternatively, the image to be identified can also be an image downloaded from the network by the image processing device; this embodiment of the invention does not limit this.

[0030] Text region features and reference text features can be considered as text-related features of the image to be recognized. Text-related features can be features used to provide information for text region detection. For example, text-related features can include semantic information representing text regions in the image to be recognized. Text-related features can be features extracted from the image to be recognized that can provide feature representation for the detector in the text region detection process. The specific types of text-related features can be selected according to actual needs, as long as the extracted features can provide information for text region detection, thereby ensuring that text regions in the image to be recognized can be detected based on these features.

[0031] Step 102: Detect the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features.

[0032] In this embodiment of the invention, region information can be used to characterize information about the text region. For example, region information may include the coordinates of the region boundary constituting the text region. The region boundary of the text region can be understood as a text bounding box, which selects the text portion of the image to be recognized. It should be noted that the text region in this embodiment of the invention can also be considered a character region; the text included in the text region can be composed of characters, and text recognition can be character recognition.

[0033] Step 103: Correct the text region based on the region information to obtain the corrected target text region.

[0034] Since the determined text region has a significant impact on the accuracy of subsequent text recognition—for example, if the text region includes too many non-text areas, or if some text is omitted, it will lead to inaccurate text recognition, such as recognition errors, incomplete text recognition, or missing text—this embodiment of the invention can first correct the text region based on the region information before recognition, obtaining a corrected target text region.

[0035] Step 104: Generate contextual features of the target text region based on its global and local features.

[0036] Step 105: Perform text recognition on the target text region based on the contextual features.

[0037] In text recognition, the recognition is based on the corrected target text region, thus ensuring accuracy to a certain extent. Furthermore, when recognizing the target text region, contextual features of the target text region are generated by combining global and local features, and text recognition is performed based on these contextual features. Since global-local features can more fully represent the target text region, using contextual features generated from global and local features for recognition can further improve the accuracy of text recognition to some extent.

[0038] In summary, the text recognition method provided in this embodiment of the invention acquires text region features and reference text features of an image to be recognized. Based on these features, the region information of the text region in the image is detected. The text region is corrected based on the region information to obtain a corrected target text region. Context features of the target text region are generated based on its global and local features, and text recognition is performed on the target text region based on these context features. Thus, by automatically performing text recognition based on the target text region, the text contained in an image can be conveniently determined. Furthermore, by performing text region correction and using context features determined based on the corrected global and local features of the target text region for text recognition, the accuracy of text recognition can be improved to some extent while conveniently determining the text contained in an image.

[0039] Optionally, in one implementation, the acquisition of text region features and reference text features of the image to be recognized may specifically include:

[0040] Step 1011: Extract features from the image to be identified based on the first coding network to obtain the text region features of the image to be identified.

[0041] The text region features can also be referred to as text region feature spectra. For example, a pre-trained, mature encoding network can be directly adopted as the first encoding network. Specifically, the first encoding network can be a pre-trained model selected as needed, such as a self-supervised learning framework model. Alternatively, the first encoding network can be trained manually; this embodiment of the invention does not impose any limitations on this. Specifically, the image to be recognized can be used as the input to the first encoding network, and the output of the first encoding network can be obtained to obtain the text region features of the image to be recognized.

[0042] Step 1012: Use the image to be recognized as the input of the second encoding network, obtain the output of the second encoding network, and obtain the reference text features of the image to be recognized; the second encoder is pre-trained using multiple sample reference texts, which are obtained from images of the same type as the image to be recognized.

[0043] The reference text features can also be called reference text feature spectra. The sample reference text can be text from a sample image, which is obtained from the scene to which the image to be identified belongs. Therefore, the sample image can be considered as an image of the same type as the image to be identified. The scene to which the image to be identified belongs refers to the scene in which the image to be identified exists. For example, if the image to be identified is a nameplate image obtained by photographing the nameplate of a communication device, then the sample image can also be a nameplate image of a communication device. For example, images from a nameplate image library of communication devices can be obtained as sample images. The text in these sample images is used as the sample reference text. The sample reference text can be obtained directly from a sample text library, which can include text included in images from the nameplate image library. The sample text library can include language materials that have actually appeared in actual use, that is, it can include text from nameplate images that have actually appeared. The sample text library can be a large-scale electronic text library that has been pre-sampled and processed to ensure the quality of the sample reference text. Alternatively, other corpora can be used to train the second coding network, for example, using the Natural Language Processing Toolkit (NLTK) corpus in the field of natural language processing. This embodiment of the invention does not limit this.

[0044] Furthermore, the type of the second encoding network can be selected according to actual needs. For example, the second encoding network can be a text encoder, which can be constructed using two deep learning networks based on self-attention mechanisms (e.g., transformer networks) to fully capture the reference text and the relationships between it and the text; that is, to fully capture the reference text features that are related to the text in the image to be recognized. This reference text can belong to the sample reference text.

[0045] During the training phase, the sample reference texts can be transformed into a dictionary composed of sample reference text vectors. Each sample reference text vector is then represented as a fixed-dimensional language feature vector using a parameterized lookup table, yielding the reference text features for each sample reference text. That is, the reference text is encoded as word vectors. Correspondingly, sample images and pre-annotated reference text features related to the sample images can be used as training data. The sample images are input to the second encoding network to be trained, and the reference text features output by the second encoding network are obtained. The network loss value is calculated based on the output reference text features and the pre-annotated reference text features, and the parameters of the second encoding network are optimized based on the network loss value. For example, stochastic gradient descent can be used for parameter tuning. After multiple rounds of training, training stops when a stopping condition is met, such as when the network loss value is less than a preset loss value threshold or the number of training rounds exceeds a preset number of rounds threshold, resulting in the second encoding network. In this way, through training, the second encoding network is equipped to output reference text features related to the input image.

[0046] In this step, the image to be recognized is used as the input to the second encoding network. The second encoding network can output reference text features related to the image to be recognized, thereby obtaining the reference text features H of the image to be recognized. ref In this process, the image to be recognized can be considered as the visual region where the text to be recognized is located. The reference text features output by the second encoding network can be one or more, and these multiple reference text features can be concatenated to form the reference text features of the image to be recognized. Based on the output reference text features, an association representation between the reference text and the text in the image to be recognized can be established.

[0047] In this embodiment of the invention, by additionally using reference text features, two dimensions of features are extracted from the image to be identified: text region features and reference text features. Text detection is then performed based on these two dimensions of features. This can improve the richness of text-related features to a certain extent, allowing them to provide more sufficient information for text region detection, thereby improving the accuracy of text region detection and further enhancing the subsequent text detection performance.

[0048] Optionally, the step of extracting features from the image to be recognized based on the first coding network to obtain the text region features of the image to be recognized may specifically include:

[0049] Step 1011a: Encode the image to be identified into a first feature vector based on the preset encoding network in the first encoding network.

[0050] Step 1011b: Extract the second feature vector of the image to be identified from the image to be identified based on the preset context network in the first encoding network.

[0051] The first encoding network can be considered as the basic encoder in the text processing flow. The first encoding network may include: a Convolutional Neural Network (CNN)-based encoding network and a transformer-based context network. The CNN-based encoding network is the preset encoding network, and the transformer-based context network is the preset context network. The preset encoding network can be considered as a spatial encoder, and the preset context network can be considered as a context encoder.

[0052] The CNN-based encoding network is responsible for encoding the raw text samples (i.e., the image to be recognized) into implicit features. Specifically, the image to be recognized can be input into a pre-defined encoding network. Based on multiple normalization layers and non-linear activation function layers, such as ReLU activation layers, the image to be recognized is encoded into a compressed feature vector H. S Among them, H S That is, the first eigenvector, H S It can be used for subsequent text region location estimation. Based on a pre-defined context network, the image to be recognized is taken as input, and a self-attention mechanism is used to obtain the relationships between semantic content in the image to be recognized, that is, to capture the semantic context relationships in the image to be recognized. Based on the obtained information, a feature vector H is output for identifying the text. T H T This is the second feature vector of the image to be recognized. The second feature vector can be a feature vector representing the semantic context of the image to be recognized.

[0053] Step 1011c: Concatenate the first feature vector and the second feature vector to obtain the text region features.

[0054] In this step, the first feature vector and the second feature vector can be concatenated. Specifically, by... S With H T Combined, they form the final text region features:

[0055] H text =concatenate(H S H T )

[0056] In this embodiment of the invention, by spatially encoding the image to be recognized based on a first encoding network in the text region encoding part, the image to be recognized is encoded into a first feature vector, and a preset context network is introduced. A second feature vector is extracted based on the preset context network, and text region features are generated based on the first feature vector and the second feature vector. This allows the text region features to better represent the spatial features of the text region in the image to be recognized, thereby facilitating subsequent text region detection by the detector and ensuring the detection effect of text region detection.

[0057] Optionally, the step of detecting the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features may specifically include:

[0058] Step 10211: Concatenate the text region features and the reference text features to obtain a joint feature representation.

[0059] Step 10212: Detect the boundary point coordinate information of the text region based on the joint feature representation; the boundary point coordinate information is used to characterize the region boundary of the text region, and the region enclosed by the region boundary is the text region.

[0060] The area enclosed by the regional boundary refers to the area surrounded by the regional boundary.

[0061] Specifically, text region features can be concatenated with reference text features to obtain a joint feature representation. This joint feature representation is then used as input to a preset detector, and its output is obtained to acquire the boundary point coordinates of the text region. In this way, by simply concatenating the text region features and reference text features as input, the boundary point coordinates of the text region can be obtained, thus improving detection efficiency to some extent. The boundary point coordinates represent the spatial coordinates of the text bounding box, i.e., the coordinates of points on the region boundary. These coordinates indicate the location of the text instance (i.e., the text region in the image to be recognized). The preset detector can be selected according to actual needs; for example, it can be a target detection model: the Yolo-v5 model, which only requires one look. Of course, other detection models can also be used.

[0062] In this embodiment of the invention, the text region features and reference text features are connected together as input, which facilitates feature processing during the text region detection process and ensures the convenience of the text region detection operation to a certain extent.

[0063] Optionally, the operation of generating contextual features of the target text region based on its global and local features may specifically include:

[0064] Step 1041: Obtain the global and local features of the target text region.

[0065] Global features can be extracted based on the entire target text region, while local features can be extracted based on a portion of the target text region. The number of global features can be 1, and the number of local features can be no less than 2.

[0066] Optionally, in one implementation, obtaining the global and local features of the target text region may specifically include: treating each part of the target text region as a word sequence and encoding the word sequence to obtain the local features; and treating the target text region as a text sequence and encoding the text sequence to obtain the global features.

[0067] Specifically, for any target text region, it can be divided into multiple parts. For any given part, this part is used as input to a traditional fully connected long short-term memory (FC-LSTM) network, and the network's output is obtained to acquire the corresponding local features. Further, all parts are stacked to form the entire text sequence, which is then used as input to the FC-LSTM network to acquire its output, yielding the corresponding global features. In other words, in this implementation, global and local features are extracted for each target text region. Subsequently, based on these global and local features, the contextual features of each target text region are determined, and text recognition is performed on each region to obtain the characters contained within it.

[0068] In this embodiment of the invention, the target text region is divided into multiple parts. By extracting global features and local features of each part respectively, the extracted features can fully characterize the target text region, thereby improving the effect of subsequent text recognition of the target text region to a certain extent.

[0069] It should be noted that there are often multiple target text regions. In another implementation, each target text region can be treated as a word sequence, and the word sequence can be encoded to obtain a local feature. All target text regions can be treated as a text sequence, and the text sequence can be encoded to obtain the global feature. Specifically, each target text region can be used as an FC-LSTM, and the features output by the FC-LSTM can be obtained to obtain multiple local features. Then, all target text regions are stacked to form a complete text sequence, and this complete text sequence is used as the input to an FC-LSTM to obtain the network's output, thus obtaining the global feature. That is, in this implementation, multiple target text regions correspond to one global feature, and one target text region corresponds to one local feature. Subsequently, based on the global and local features of the overall target text region, the context features of the overall target text region are determined. Text recognition is then performed based on these context features to obtain the characters included in all target text regions, thereby improving recognition efficiency to some extent.

[0070] Step 1042: For any of the local features, calculate the weight corresponding to the local feature based on the local feature and the global feature.

[0071] With H i H represents the i-th local feature. g This represents the global feature. The value of i can be 1, 2, ..., X, where X is the total number of local features. For local features H... i In terms of local features H i The corresponding weight can be represented as ai.

[0072] Specifically, ai = softmax(ei). Both local and global features can be in matrix form, with local features H... i The corresponding weight ei can be specifically defined as: ei = H i T ·H g Among them, H i T Representing local features H i The transpose of ei can be used to represent global features and local features H. i The correlation between global features and local features H is normalized using the softmax function in this embodiment of the invention. i The correlation between the local features H is then obtained. i The corresponding weights.

[0073] In this embodiment of the invention, global features and local features H are calculated. iThe correlations are determined, and weights are calculated based on these correlations. That is, the relationship between these correlations and global features is calculated, and then the local features are aligned to form the context features of the final representation.

[0074] Step 1043: Calculate the context features of the target text region based on each of the local features and the weights corresponding to each local feature.

[0075] Specifically, a weighted average can be performed based on local features and their corresponding weights to obtain a global-local context representation, i.e., context features. The context feature Hc can be specifically defined as follows:

[0076] In this embodiment of the invention, since the weights corresponding to local features are calculated by combining the local features and the global features, using local features and their corresponding weights is equivalent to calculating context features based on local and global features. Furthermore, since each local feature corresponds to a portion of the target text region, and there are contextual relationships between these portions, the context features of the target text region can be calculated based on the local and global features. Further, in this embodiment of the invention, by unifying the global and local features of the target text region into context features, processing can be facilitated during text recognition operations.

[0077] It should be noted that in this embodiment of the invention, after obtaining the context features, a Softmax classifier can be used for text recognition. Specifically, the context features can be used as output to obtain a probability list of words corresponding to the target text region output by the Softmax classifier. Specifically, this probability list can include a probability list of words corresponding to each region in the target text region. The words involved in the probability list can be words from a preset word library, and the probability list can represent the probability that a character in the region is that word. Further, for any region, the word with the highest probability in the probability list corresponding to that region is taken as the character included in that region. Finally, the characters included in each region are concatenated sequentially to obtain the text included in the image to be recognized. For example, the characters included in the regions can be combined according to the distribution of the regions to obtain the text included in the image to be recognized.

[0078] For example, the probability list output by the Softmax classifier can be represented as: y = softmax(WcHc + bc). Here, Wc and bc represent the trainable parameters in the Softmax classifier, and Wc and bc can be matrices. Wc and bc can be continuously optimized through parameter tuning during the training phase of the Softmax classifier, and the final trainable parameters can be obtained at the end of training. The Softmax classifier can use cross-entropy loss as the objective function (i.e., the loss function) to train and obtain these trainable parameters. Specifically, the loss value can be calculated based on the cross-entropy loss function, and parameter tuning can be performed based on the loss value.

[0079] It should be noted that other methods can also be used to achieve text recognition based on target text regions. For example, convolutional features of the target text region can be extracted first in a convolutional layer. The extracted convolutional features are then input into a recurrent network layer, such as a deep bidirectional LSTM network, to further extract text sequence features based on the convolutional features. The output of the recurrent network layer is used as the output of a softmax classifier, which outputs the characters. In this recurrent network layer, each feature vector in the output feature sequence can be generated in a left-to-right reading order, with each feature sequence representing a specific part of the target text region.

[0080] Furthermore, in one implementation, each part can be identified separately, and then the characters corresponding to each part can be connected to obtain the final identified text. This embodiment of the invention does not limit this approach.

[0081] In this embodiment of the invention, text region detection is first used to locate the text within the image to be recognized and to define the range of the text. Next, text region correction is performed to improve the accuracy of the text region recognition. Finally, text recognition based on the text region context is performed to identify the located text region and determine the specific text within it, thus converting the text region in the image to be recognized into character information.

[0082] Optionally, the step of correcting the text region based on the region information to obtain the corrected target text region may specifically include:

[0083] Step 1031: Sample the boundary point coordinate information of the text region according to the joint feature representation to obtain M control points.

[0084] The quantity M can be set according to actual needs. M can be an integer not less than 2, for example, M can be 20, 30, etc. In one implementation, M can be set according to the computing power of the image processing device. The specific value of M can be positively correlated with the computing power of the image processing device.

[0085] Furthermore, the boundary point coordinate information can include the coordinates of multiple boundary points. Specifically, a control point selection model can be pre-trained, which can combine the aforementioned joint feature representation to select M coordinates from the coordinates of multiple boundary points. For example, the sample joint feature representation and sample boundary point coordinate information of a sample image can be obtained, and M sample control points can be selected from the sample boundary point coordinate information to obtain training data. The sample control points can be selected through manual annotation. Then, the sample joint feature representation and sample boundary point coordinate information of the sample image are used as input to the control point selection model to be trained, and the output control points of the model are obtained. Based on the output control points and the actual sample control points, the loss value of the model is calculated, and then the parameters of the model are adjusted based on the loss value, for example, using stochastic gradient descent. After multiple rounds of training, training stops when a stopping condition is met, such as when the network loss value is less than a preset loss value threshold or the number of training rounds exceeds a preset number of rounds threshold, resulting in a trained control point selection model. Accordingly, in this step, the joint feature representation of the image to be recognized and the coordinate information of the boundary points of the text region can be used as input to obtain the output of the trained control point selection model, thus conveniently completing the sampling and obtaining M control points. These M control points can be used to represent the text bounding box. Alternatively, when selecting control points, a preset sampling algorithm can be used to select M coordinates from the coordinates of multiple boundary points. These M coordinates can then represent the control points. Specifically, the boundary points corresponding to these M coordinates are the determined control points. For example, the control points can be determined based on the thin-plate spline (TPS) algorithm.

[0086] Step 1032: Perform boundary fitting based on the M control points to determine the corrected region boundary; wherein the region enclosed by the corrected region boundary is the target text region.

[0087] In this step, the original image coordinates of M control points can be used for fitting using a two-dimensional polynomial fitting method. The two-dimensional polynomial fitting is used to fit a polygonal region that optimally represents the spatial position of these control points. For example, the coordinates of these M control points can be used as input to the two-dimensional polynomial fitting algorithm to obtain the region boundary information output by the algorithm. This region boundary information indicates the corrected region boundary. The area enclosed by the corrected region boundary is the target text region, thus providing more accurate text location information for subsequent processing. The region boundary information can include multiple coordinate information, which can be the pixel coordinates of points in the image to be recognized. Based on the coordinates of the control points in the fitted region boundary, the coordinates of other points in the fitted region boundary can be determined to obtain the final region boundary information. The boundary formed by the points indicated by the coordinates in the region boundary information is the corrected region boundary.

[0088] In this embodiment of the invention, the coordinate information of the boundary points of the text region is sampled according to the joint feature representation to obtain M control points. Boundary fitting is then performed based on these M control points to obtain the corrected region boundary, where the region whose corrected boundary matches is the target text region. Thus, by performing boundary fitting based on control points selected according to the joint feature representation, the corrected region boundary can be made more accurate to some extent.

[0089] Optionally, the text region features include multiple first sub-features, and the reference text features include multiple second sub-features; the joint feature representation can be obtained through the following steps:

[0090] Step A1: For any first sub-feature, calculate the first weight factor of the first sub-feature based on the first sub-feature and the second sub-feature corresponding to the first sub-feature.

[0091] The number of first sub-features and the number of second sub-features can be the same. For example, a vector of a predetermined number of dimensions in the text region features can be divided into one first sub-feature, and a vector of a predetermined number of dimensions in the reference text features can be divided into one second sub-feature, thus obtaining multiple first sub-features and multiple second sub-features. Alternatively, one of the second sub-features in the reference text features of the image to be recognized can be a reference text feature output by the second encoding network for the image to be recognized. Accordingly, the text region features can be divided into a corresponding number of first sub-features based on the number of second sub-features.

[0092] The first sub-feature can be represented as The second sub-feature can be represented as in, H textH represents the text region features of the image to be recognized. ref The reference text features represent the image to be recognized, where n takes values ​​of 1, 2, ..., N, and t takes values ​​of 1, 2, ..., T. The second sub-feature corresponding to the first sub-feature can be a second sub-feature with the same index as the first sub-feature, for example, in... When n is 1, the second sub-feature corresponding to the first sub-feature can be t = 1. Alternatively, the second sub-feature corresponding to the first sub-feature can also be any of the second sub-features, for example, in When n is 1, the second sub-feature corresponding to the first sub-feature can be t for 1, 2, ..., T.

[0093] The first weighting factor can be represented as a t,n Specifically, a t,n It can be calculated in the following way:

[0094]

[0095]

[0096] in, It can be in matrix form. express The transpose of .

[0097] Furthermore, the first weighting factor can be the cumulative factor of the first sub-feature, and the degree of sharing of the first sub-feature in the text sequence can be controlled based on this first weighting factor.

[0098] Step A2: Based on each of the first sub-features and the first weight factor of the first sub-features, perform a weighted summation on the multiple first sub-features to obtain the target feature representation of the text region feature.

[0099] Step A3: Calculate the filtering weights based on the target feature representation and the second sub-feature;

[0100] Step A4: Calculate the fusion feature component corresponding to the second sub-feature based on the target feature representation, the second sub-feature, and the filtering weight.

[0101] Among them, the target feature representation can be in the form of c t Specifically, the target feature representation can be:

[0102] Furthermore, the fused feature components can be represented as y tWhen calculating the fused feature components, the filtering weight g of the second sub-feature in the reference text features can be calculated first based on the target feature representation and the second sub-feature. The filtering weight g can be positively correlated with the correlation of the second sub-feature; that is, the higher the correlation, the higher the filtering weight g, thereby increasing the contribution of reference text features with high correlation and reducing the influence of reference text features with low correlation.

[0103] Specifically, the filter weight g can be:

[0104]

[0105] Where W, U, and b represent preset parameters.

[0106] Furthermore, the fused feature components can be specifically defined as follows:

[0107]

[0108] Step A5: Concatenate the fused feature components corresponding to each of the second sub-features to obtain the joint feature representation.

[0109] For example, the fused feature components y1, y2, ..., y2 corresponding to each second sub-feature can be... T The features are then concatenated to obtain the final optimized and filtered joint feature representation.

[0110] In this embodiment of the invention, the fused feature components can represent the filtered feature representation, and the above processing can be implemented using a text filtering gate based on attention fusion. Specifically, the text filtering gate operation can be implemented using linear projection, summation, and sigmoid activation. Linear projection corresponds to the step of calculating the first weight factor, and summation corresponds to the calculation of c. t In the process of calculating the filter weights g, the sigmoid activation corresponds to the process of calculating the filter weights g mentioned above.

[0111] Furthermore, directly concatenating the text region features of the image to be recognized with the reference text features may result in redundant information and imbalanced distribution in the joint feature representation. In the above processing, calculating a first weight factor and then calculating yt based on the first weight factor and filtering weights is equivalent to filtering the first and second sub-features. This reduces redundant information and alleviates the imbalanced distribution to some extent, thus avoiding compromise of feature robustness and improving the quality of the final joint feature representation. Correspondingly, subsequent text region detection and correction using this joint feature can improve both detection and correction performance to some extent.

[0112] It should be noted that, in this embodiment of the invention, the text region features can also be directly concatenated with the reference text features to obtain a joint feature representation. The boundary point coordinates of the text region are detected based on this joint feature representation. Then, the boundary point coordinates of the text region are sampled based on this joint feature representation to perform text correction and obtain the corrected region boundary. Next, a new joint feature representation is obtained using steps A1-A5 described above. Then, the boundary point coordinates of the corrected region boundary are sampled again using this new joint feature representation to obtain M new control points. Boundary fitting is performed based on the new M control points to perform correction again, thereby further improving the correction effect. The implementation methods of each step can be referred to the aforementioned descriptions, and will not be repeated here.

[0113] The following describes an application scenario related to an embodiment of the present invention. During the installation and maintenance of communication equipment, obtaining the model information provided by the nameplate is a crucial task for on-site equipment deployment and network management systems. For example, it is frequently necessary to obtain the information provided by the nameplates of the baseband processing unit (BBU) and the radio frequency remote unit (RRU). Figure 2 This is a schematic diagram of a scenario provided by an embodiment of the present invention. In this application scenario, the image to be identified can be a nameplate image. First, text region detection can be performed on the nameplate image based on the backbone network. Then, text region correction is performed based on the text region features obtained from the text region detection, reference text features, and the region information of the text region. Finally, text recognition based on global-local context is performed to obtain the text included in the nameplate image. Here, character local context can refer to acquiring multiple local features, and global-local association can refer to calculating the context features of the target text region based on each local feature and its corresponding weight. Character region recognition can refer to text recognition based on context features.

[0114] Manually examining nameplates to obtain information is error-prone and time-consuming. Traditional Optical Character Recognition (OCR) technology for detecting and recognizing text in visual images often leads to text recognition errors due to the inability to accurately determine text boundaries. This invention addresses this issue by using visual information processing to automatically recognize text based on the context of the text region. Through correction, more accurate boundaries are obtained, and the global-local context mechanism of the corrected text region is used to more precisely recognize the nameplate text. This rapid and accurate recognition of text in nameplate images reduces error rates and time consumption, effectively assisting in the installation and maintenance of communication equipment, thereby improving the efficiency and quality of actual work.

[0115] Furthermore, compared to directly using text detection boxes obtained from visual object detection technology as input for subsequent text recognition, in this embodiment of the invention, the initial text region is corrected based on a joint feature representation constructed from the text region feature spectrum and the reference text feature spectrum to obtain a corrected target text region, thus providing a more accurate representation of the text's location. Moreover, obtaining the contextual features of the text region can, to some extent, extract the feature parts in the text region that are beneficial to the correct recognition result, thereby obtaining a more accurate text recognition result.

[0116] It should be noted that for scenarios requiring high recognition accuracy, the text recognition method described in this embodiment can be tested before being put into use. Only when the text recognition accuracy meets a preset accuracy threshold can the text recognition method be used to recognize the nameplate image. For example, Figure 3 This is a schematic diagram of a data test provided in an embodiment of the present invention, such as... Figure 3 As shown, the text identified from each text region in the image (i.e., the region enclosed by the dashed box) can be displayed above each text region. Accordingly, the actual text in a text region can be compared with the displayed identified text to determine whether the text region was accurately identified. The ratio of the number of accurately identified text regions to the total number of text regions is calculated to obtain the text recognition accuracy.

[0117] Figure 4 This is a structural diagram of a text recognition device provided in an embodiment of the present invention. The device 20 may include:

[0118] The acquisition module 201 is used to acquire the text region features and reference text features of the image to be recognized;

[0119] Detection module 202 is used to detect the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features;

[0120] The correction module 203 is used to correct the text region based on the region information to obtain the corrected target text region;

[0121] The generation module 204 is used to generate contextual features of the target text region based on the global and local features of the target text region.

[0122] The recognition module 205 is used to perform text recognition on the target text region based on the context features.

[0123] Optionally, the generation module 204 is specifically used for:

[0124] Obtain the global and local features of the target text region;

[0125] For any of the local features, calculate the weight corresponding to the local feature based on the local feature and the global feature;

[0126] The context features of the target text region are calculated based on each of the local features and the weights corresponding to each local feature.

[0127] Optionally, the generation module 204 is further specifically used for:

[0128] Each part of the target text region is taken as a word sequence, and the word sequence is encoded to obtain the local features;

[0129] The target text region is treated as a text sequence, and the text sequence is encoded to obtain the global feature.

[0130] Optionally, the acquisition module 201 is specifically used for:

[0131] Based on the first coding network, feature extraction is performed on the image to be identified to obtain the text region features of the image to be identified;

[0132] The image to be identified is used as the input of the second encoding network, and the output of the second encoding network is obtained to obtain the reference text features of the image to be identified; the second encoder is pre-trained using multiple sample reference texts, which are obtained from images of the same type as the image to be identified.

[0133] Optionally, the acquisition module 201 is further specifically used for:

[0134] The image to be identified is encoded into a first feature vector based on the preset encoding network in the first encoding network;

[0135] Based on the preset context network in the first coding network, the second feature vector of the image to be identified is extracted from the image to be identified;

[0136] The first feature vector and the second feature vector are concatenated to obtain the text region features.

[0137] Optionally, the detection module 202 is specifically used for:

[0138] The text region features of the image to be identified and the reference text features of the image to be identified are concatenated to obtain a joint feature representation;

[0139] Based on the joint feature representation, the boundary point coordinate information of the text region is detected; the boundary point coordinate information is used to characterize the region boundary of the text region, and the region enclosed by the region boundary is the text region.

[0140] Optionally, the correction module 203 is specifically used for:

[0141] Based on the joint feature representation, the boundary point coordinate information of the text region is sampled to obtain M control points;

[0142] Boundary fitting is performed based on the M control points to determine the corrected region boundary; wherein the region enclosed by the corrected region boundary is the target text region.

[0143] Optionally, the text region features include multiple first sub-features, and the reference text features include multiple second sub-features; the detection module 202 is further specifically used for:

[0144] For any first sub-feature, a first weighting factor of the first sub-feature is calculated based on the first sub-feature and the second sub-feature corresponding to the first sub-feature.

[0145] Based on each first sub-feature and its first weight factor, the multiple first sub-features are weighted and summed to obtain the target feature representation of the text region feature.

[0146] Calculate the filtering weights based on the target feature representation and the second sub-feature;

[0147] Based on the target feature representation, the second sub-feature, and the filtering weight, calculate the fusion feature component corresponding to the second sub-feature;

[0148] The joint feature representation is obtained by concatenating the fused feature components corresponding to each of the second sub-features.

[0149] In summary, the text recognition device provided in this embodiment of the invention acquires text region features and reference text features of an image to be recognized. Based on the text region features and reference text features of the image to be recognized, it detects the region information of the text region in the image to be recognized. Based on the region information, it corrects the text region to obtain the corrected target text region. Based on the global and local features of the target text region, it generates context features of the target text region and performs text recognition on the target text region based on the context features. Thus, by automatically performing text recognition based on the target text region, the text contained in an image can be conveniently determined. Furthermore, by performing text region correction and using context features determined based on the global and local features of the corrected target text region for text recognition, the accuracy of text recognition can be improved to a certain extent while conveniently determining the text contained in an image.

[0150] The present invention also provides an electronic device, see [link to relevant documentation]. Figure 5 The system includes: a processor 901, a memory 902, and a computer program 9021 stored in the memory and executable on the processor. When the processor executes the program, it implements the text recognition method of the foregoing embodiments.

[0151] The present invention also provides a readable storage medium, wherein when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device is able to perform the text recognition method of the foregoing embodiments.

[0152] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0153] It should be noted that all information and data obtained in the embodiments of the present invention were obtained with the authorization of the information / data holder.

[0154] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0155] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0156] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.

[0157] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0158] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present invention. The present invention can also be implemented as a device or apparatus program for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0159] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0160] The user information (including but not limited to user device information, user personal information, etc.) and related data involved in this invention are all information authorized by the user or by the parties.

[0161] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0163] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A text recognition method, characterized by, The method includes: Obtain the text region features and reference text features of the image to be identified; Based on the text region features of the image to be identified and the reference text features, the region information of the text region in the image to be identified is detected; The text region is corrected based on the region information to obtain the corrected target text region; Based on the global and local features of the target text region, the contextual features of the target text region are generated; Text recognition is performed on the target text region based on the aforementioned contextual features; The acquisition of text region features and reference text features of the image to be identified includes: Based on the first coding network, feature extraction is performed on the image to be identified to obtain the text region features of the image to be identified; The image to be identified is used as the input of the second encoding network, and the output of the second encoding network is obtained to obtain the reference text features of the image to be identified. The second encoding network is pre-trained using multiple sample reference texts, which are obtained from images of the same type as the image to be identified.

2. The method of claim 1, wherein, The step of generating contextual features of the target text region based on its global and local features includes: Obtain the global and local features of the target text region; For any of the local features, calculate the weight corresponding to the local feature based on the local feature and the global feature; The context features of the target text region are calculated based on each of the local features and the weights corresponding to each local feature.

3. The method of claim 2, wherein, The acquisition of global and local features of the target text region includes: Each part of the target text region is taken as a word sequence, and the word sequence is encoded to obtain the local features; The target text region is treated as a text sequence, and the text sequence is encoded to obtain the global feature.

4. The method of claim 1, wherein, The step of extracting features from the image to be identified based on the first coding network to obtain the text region features of the image to be identified includes: The image to be identified is encoded into a first feature vector based on the preset encoding network in the first encoding network; Based on the preset context network in the first coding network, the second feature vector of the image to be identified is extracted from the image to be identified; The first feature vector and the second feature vector are concatenated to obtain the text region features.

5. The method according to any of claims 1 to 4, characterized in that, The step of detecting the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features includes: The text region features and the reference text features are concatenated to obtain a joint feature representation; Based on the joint feature representation, the boundary point coordinate information of the text region is detected; the boundary point coordinate information is used to characterize the region boundary of the text region, and the region enclosed by the region boundary is the text region.

6. The method of claim 5, wherein, The step of correcting the text region based on the region information to obtain the corrected target text region includes: Based on the joint feature representation, the boundary point coordinate information of the text region is sampled to obtain M control points; Boundary fitting is performed based on the M control points to determine the corrected region boundary; wherein the region enclosed by the corrected region boundary is the target text region.

7. The method of claim 5, wherein, The text region features include multiple first sub-features, and the reference text features include multiple second sub-features; The step of concatenating the text region features and the reference text features to obtain a joint feature representation includes: For any first sub-feature, a first weighting factor of the first sub-feature is calculated based on the first sub-feature and the second sub-feature corresponding to the first sub-feature. Based on each first sub-feature and its first weight factor, the multiple first sub-features are weighted and summed to obtain the target feature representation of the text region feature. Calculate the filtering weights based on the target feature representation and the second sub-feature; Based on the target feature representation, the second sub-feature, and the filtering weight, calculate the fusion feature component corresponding to the second sub-feature; The joint feature representation is obtained by concatenating the fused feature components corresponding to each of the second sub-features.

8. A text recognition device, characterized in that, The device includes: The acquisition module is used to acquire the text region features and reference text features of the image to be recognized; The detection module is used to detect the region information of the text region in the image to be identified based on the text region features of the image to be identified and the reference text features; A correction module is used to correct the text region based on the region information to obtain the corrected target text region; The generation module is used to generate contextual features of the target text region based on the global and local features of the target text region. The recognition module is used to perform text recognition on the target text region based on the context features; The acquisition module is specifically used for: Based on the first coding network, feature extraction is performed on the image to be identified to obtain the text region features of the image to be identified; The image to be identified is used as the input of the second encoding network, and the output of the second encoding network is obtained to obtain the reference text features of the image to be identified. The second encoding network is pre-trained using multiple sample reference texts, which are obtained from images of the same type as the image to be identified.

9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1-7.

10. A readable storage medium, characterized by, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Text area detection method and device, computer equipment and storage medium

    CN112613402A

  • Text recognition method and device

    CN112749695A