Text recognition method based on sequence recognition and related device
By combining the sequence recognition method of deep convolutional neural network and recurrent neural network model, the problem of inefficient recognition in existing text recognition technologies is solved, and more efficient and accurate text recognition is achieved.
Patent Information
- Application Number
- CN202510286214.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
Existing text recognition technology is affected by factors such as image quality, font style and language complexity, resulting in bias or error in recognition results. It consumes a lot of time and computing resources to process large-scale text data, making the recognition efficiency in inefficient.
The text recognition method based on sequence recognition is adopted, combined with the deep convolutional neural network model and the recurrent neural network model, and the target text recognition efficiency is determined through feature sequence extraction, labeling and transcription.
Improves the accuracy and efficiency of text recognition, reduces dependence on image quality and font styles, and can process large-scale text data faster.
Smart Images

Figure CN120220161A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text recognition, and in particular, to a text recognition method and related device based on sequence recognition. Background Art
[0002] In today's digital and information age, efficient and accurate text recognition technology plays a crucial role in many fields, such as automated document processing, intelligent image understanding, information retrieval, etc. Traditional text recognition methods have many limitations. With the rise of deep learning technology, deep convolutional neural network models perform well in image feature extraction, but their results need to be further processed. Current text recognition technology is often affected by various factors such as image quality, font style, language complexity, etc., resulting in deviation or error in the recognition result, and it is difficult to achieve a satisfactory level of recognition accuracy. At the same time, when dealing with large-scale text data, existing technologies often require a lot of time and computing resources, and the recognition efficiency is very low. Therefore, how to improve text recognition efficiency is an urgent problem to be solved. Summary of the Invention
[0003] Embodiments of the present application provide a text recognition method and related device based on sequence recognition, which improve text recognition efficiency.
[0004] In a first aspect, embodiments of the present application provide a text recognition method based on sequence recognition, which is applied to a text recognition system. The text recognition system includes a deep convolutional neural network model and a recurrent neural network model. The method includes:
[0005] Obtain a target image;
[0006] Extract a feature sequence from the target image through the deep convolutional neural network model to obtain n feature sequences; n is an integer greater than 1;
[0007] Label each of the n feature sequences to obtain n sequence labels;
[0008] Transcribe the n sequence labels through the recurrent neural network model to obtain a target label sequence;
[0009] Determine a target text based on the target label sequence.
[0010] In a second aspect, embodiments of the present application provide a text recognition device, which is applied to a text recognition system. The text recognition system includes a deep convolutional neural network model and a recurrent neural network model. The device includes: an obtaining unit and a processing unit;
[0011] The obtaining unit is used to obtain a target image;
[0012] The processing unit is configured to extract a feature sequence from the target image through the deep convolutional neural network model, obtaining n feature sequences; n is an integer greater than 1;
[0013] Annotate each of the n feature sequences to obtain n sequence annotations;
[0014] Transcribe the n sequence annotations through the recurrent neural network model to obtain a target label sequence;
[0015] Determine a target text based on the target label sequence.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, including: a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the processor so that the electronic device executes the method according to the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method according to the first aspect.
[0018] In a fifth aspect, an embodiment of the present invention provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, such that the computer executes the method according to the first aspect.
[0019] Implementing the embodiments of the present invention has the following beneficial effects:
[0020] It can be seen that the text recognition method based on sequence recognition described in the embodiments of the present invention is applied to a text recognition system, which includes a deep convolutional neural network model and a recurrent neural network model. The method includes: obtaining a target image, extracting a feature sequence from the target image through the deep convolutional neural network model to obtain n feature sequences, where n is an integer greater than 1, annotating each of the n feature sequences to obtain n sequence annotations, transcribing the n sequence annotations through the recurrent neural network model to obtain a target label sequence, and determining a target text based on the target label sequence, thereby improving the text recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.
[0022] Figure 1It is a schematic structural diagram of a text recognition system provided by an embodiment of the present application;
[0023] Figure 2 It is a flowchart of a text recognition method based on sequence recognition provided by an embodiment of the present application;
[0024] Figure 3 It is a flowchart of determining n feature sequences provided by an embodiment of the present application;
[0025] Figure 4 It is a flowchart of determining k quality evaluation values provided by an embodiment of the present application;
[0026] Figure 5 It is a flowchart of determining a target label sequence provided by an embodiment of the present application;
[0027] Figure 6 It is an example diagram of a target text display interface provided by an embodiment of the present application;
[0028] Figure 7 It is a schematic structural diagram of a text recognition device based on sequence recognition provided by an embodiment of the present application;
[0029] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments
[0030] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0031] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0032] References to "embodiments" in this disclosure mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment each time, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0033] First, the relevant terms involved in the present application are explained as follows:
[0034] DCNN: Deep Convolutional Neural Network (DCNN), which is a deep learning architecture that contains multiple convolutional layers to automatically extract features from input data (usually images). It captures local patterns and features through convolutional operations and uses a multi-layer structure to achieve high-level abstract representation and learning of complex data, thereby being able to effectively handle tasks such as image classification, object detection, and image generation.
[0035] RNN: Recurrent Neural Network (RNN), which is a type of neural network used to process sequential data. Its characteristic is that there are directed cyclic connections between neurons, enabling the network to utilize historical information to process the current input and affect subsequent inputs. This gives it unique advantages in processing tasks such as text in natural language processing and time series prediction.
[0036] Please refer to Figure 1 , Figure 1It is a schematic structural diagram of a text recognition system provided by an embodiment of the present application. The text recognition system 10 includes a deep convolutional neural network model 101 and a recurrent neural network model 102. The text recognition system 10 is an overall framework or architecture for converting an input target image into an understandable and processable target text. It integrates multiple technologies and models to achieve effective recognition and conversion from an image to text. The main purpose of this system is to accurately extract meaningful text information from complex images and present it in a structured and understandable manner. The deep convolutional neural network model 101 is an important component in the text recognition system 10. It is mainly responsible for extracting feature sequences from the input target image. With its powerful ability in image processing, it can automatically discover and extract various valuable features from the image. Due to its deep structure, it can learn features at different levels and abstraction degrees, thereby capturing subtle and complex patterns in the image. The recurrent neural network model 102 is also a key component of the text recognition system 10. It receives n sequence annotations extracted and labeled by the deep convolutional neural network model 101 as input. Its function is to transcribe these sequence annotations, that is, by learning the internal relationships and patterns between the sequence annotations, convert them into a target label sequence. This transcription process can utilize the order and interdependence of the sequence annotations to generate more semantically and logically coherent results. Finally, based on the target label sequence generated by the recurrent neural network model 102, the text recognition system 10 can determine the final target text.
[0037] In this embodiment, by combining the deep convolutional neural network model 101 and the recurrent neural network model 102, the text recognition efficiency is improved. Specifically, the combination of the deep convolutional neural network model 101 and the recurrent neural network model 102 can be implemented by the following method: The convolutional layer is based on the standard recurrent neural network architecture, removes the fully connected layer, extracts the image feature map through convolutional and pooling operations, and converts it into a sequence of feature vectors by columns to provide input for subsequent sequence modeling. The recurrent layer adopts a bidirectional long short-term memory network structure, which can utilize forward and backward context information to capture long-distance dependencies and perform frame-by-frame prediction on the feature sequence generated by the convolutional layer. The transcription layer converts the frame-by-frame prediction of the recurrent layer into the final label sequence through the connectionist temporal classification algorithm, supporting both dictionary-free and dictionary-based modes, simplifying the training process and improving the recognition efficiency. It should be noted that the implementation of end-to-end training relies on the joint training mechanism. By organically combining the convolutional layer, the recurrent layer, and the transcription layer, the entire network can be jointly trained through a unified loss function. During the training process, the feature sequence extracted by the convolutional layer is directly passed to the recurrent layer, and the prediction result of the recurrent layer is converted into the final label sequence through the transcription layer. The loss function is defined based on the negative log-likelihood of the true label sequence, and the learning rate of each dimension is automatically calculated through stochastic gradient descent and algorithms, thereby accelerating the convergence of the network. Since the connectionist temporal classification algorithm is adopted, the model can directly learn from the sequence labels (such as words) without detailed annotation of the position of each character. This design greatly reduces the annotation workload, improves the training efficiency and generalization ability of the model. Through the end-to-end training mechanism, each component of the network can be optimized collaboratively to achieve the globally optimal parameter configuration. Compared with the traditional phased training method, end-to-end training can better capture the mapping relationship between image features and sequence labels and improve the overall performance of the model. The embodiment of the present invention realizes the organic combination of the deep convolutional neural network and the recurrent neural network, constructs an end-to-end trainable neural network architecture, and thus provides an efficient, flexible and easy-to-implement solution for the image sequence recognition task.
[0038] In this embodiment, first, a target image is obtained, then the deep convolutional neural network model is used to extract a feature sequence from the target image to obtain n feature sequences. Next, each of the n feature sequences is annotated to obtain n sequence annotations. Then, the recurrent neural network model is used to transcribe the n sequence annotations to obtain a target label sequence. Finally, the target text is determined based on the target label sequence, improving the text recognition efficiency.
[0039] Please refer to Figure 2 , Figure 2It is a flowchart of a text recognition method based on sequence recognition provided by an embodiment of the present application, including but not limited to the following steps:
[0040] S201: Obtain a target image.
[0041] In this embodiment, the target image represents a specific object that needs to be processed for text recognition. It may be various scene images containing text, such as book pages, billboards, handwritten notes, product labels, etc. Its role is to provide the original data input for the subsequent text recognition process. By processing and analyzing the target image, valuable text information can be extracted from it. The target image is the starting point of the entire text recognition system, and factors such as its quality, clarity, and content complexity will directly affect the accuracy and effect of subsequent feature extraction, annotation, transcription, and final text determination. A high-quality, clear, and content-defined target image helps to improve the success rate and accuracy of the entire text recognition process, thereby achieving the accurate conversion of text information in the image into computer-readable text that can be edited and processed.
[0042] In this embodiment, the ways to obtain the target image can be: an image acquisition device, which can directly capture scenes or objects in the real world through a camera (such as a digital camera, smartphone camera) to obtain the target image; a scanning device, which scans paper documents, pictures, etc. and converts them into digital images as the target image; reading from a database or storage medium, where images previously stored in a computer hard disk, server, cloud storage, etc. are read out as the target image; network acquisition, downloading images from the Internet or receiving images through a network interface, etc.
[0043] S202: Extract feature sequences from the target image through the deep convolutional neural network model to obtain n feature sequences.
[0044] In this embodiment, n is an integer greater than 1. Please refer to Figure 3 , Figure 3 It is a flowchart of a method for determining n feature sequences provided by an embodiment of the present application, including but not limited to the following steps:
[0045] S301: Extract feature sequences from the target image through the deep convolutional neural network model to obtain k feature sequences.
[0046] In this embodiment, the feature sequence is a series of feature representations with specific patterns and meanings extracted from the target image through a deep convolutional neural network model. Specifically, when the deep convolutional neural network model processes the target image, it automatically extracts various features from different regions and levels of the image. These features do not exist in isolation but are combined together in a certain order and relationship to form a sequence. Each feature has a specific position and meaning in the sequence. These features may include information such as the edges, textures, color distributions, and shapes of the image. The role of the feature sequence is to provide a basis and foundation for subsequent processing steps (such as annotation, transcription, etc.). By further analyzing and processing these feature sequences, the content of the image can be understood more deeply and ultimately transformed into the target text.
[0047] In this embodiment, k is an integer greater than or equal to n. The following method can be used to extract the feature sequence from the target image through the deep convolutional neural network model: The deep convolutional neural network model consists of multiple convolutional layers, pooling layers, fully connected layers, etc. When the target image is input into the deep convolutional neural network model, first, in the convolutional layer, the target image is convolved using multiple different convolutional kernels (also called filters). Each convolutional kernel slides on the target image and performs an inner product operation with the local region of the target image, thereby extracting the local features of the target image. Then, the feature map obtained by convolution is downsampled through the pooling layer to reduce the dimension of the features, while retaining the main features and reducing the computational amount and the risk of overfitting. After alternating processing through multiple convolutional layers and pooling layers, the deep convolutional neural network model gradually extracts more abstract and higher-level features. Finally, these features are flattened and processed through the fully connected layer to obtain the final feature sequence. Due to the complexity and multi-layer structure of the deep convolutional neural network model, as well as the roles of different convolutional kernels and parameters, multiple different feature sequences are extracted from the target image, and thus k feature sequences are obtained. Each feature sequence represents the feature representation of the target image in a specific dimension or perspective. Combining these feature sequences can more comprehensively describe the information of the target image. For example, the role of the feature sequence is to provide a basis and foundation for subsequent processing steps (such as annotation, transcription, etc.). By further analyzing and processing these feature sequences, the content of the image can be understood more deeply and ultimately transformed into the target text.
[0048] S302: Determine the quality evaluation value corresponding to each feature sequence among the k feature sequences to obtain k quality evaluation values.
[0049] In this embodiment, the quality assessment value is a quantitative index used to measure the quality of each feature sequence. By comparing it with a set quality assessment threshold, feature sequences with higher quality can be screened out. Among the numerous extracted feature sequences, only those feature sequences with quality assessment values greater than the threshold are selected for subsequent processing, thereby removing feature sequences that may contain noise, lack representativeness, or have poor quality, improving the accuracy and efficiency of subsequent processing. The quality assessment value can compare and rank different feature sequences, helping to determine which feature sequences are more likely to contain information useful for text recognition. During the training and optimization of the model, the quality assessment value can serve as a feedback signal to help adjust the parameters of the deep convolutional neural network model to improve the quality of future-extracted feature sequences. According to the quality assessment value of the feature sequence, computing resources or attention can be reasonably allocated, with more resources used to process feature sequences with higher quality, improving the performance of the entire text recognition system. The quality assessment value is a quantitative measure of the quality of the feature sequence, helping to make selections and decisions among numerous feature sequences to optimize the text recognition effect.
[0050] Please refer to Figure 4 , Figure 4 which is a flowchart for determining k quality assessment values provided by an embodiment of this application, including but not limited to the following steps:
[0051] S401: Determine at least one feature value in the first feature sequence.
[0052] In this embodiment, the first feature sequence is any one of the k feature sequences, and the feature value is a numerical value existing in the first feature sequence. It should be noted that a feature sequence is an ordered arrangement of a series of features obtained by performing a certain processing or analysis on a target image, and the feature value is the specific quantitative representation of these features. For example, if the feature sequence describes the color features of a certain region in an image, the feature value may be the specific numerical values representing the red, green, and blue color components of that region; if the feature sequence describes the texture features of an image, the feature value may be the numerical values representing characteristics such as texture roughness and directionality.
[0053] In this embodiment, statistics such as the mean, median, mode, and standard deviation of the first feature sequence can be calculated, and these statistics can be used as feature values. Also, the maximum and minimum values in the first feature sequence can be found as feature values, or the numerical values at the beginning, middle, or end positions of the first feature sequence can be selected as feature values, or the numerical values with higher occurrence frequencies can be determined as feature values. Additionally, the value range of the first feature sequence can be divided into several intervals (bins), and the representative value (such as the midpoint of the interval) of each interval can be used as a feature value.
[0054] Since the first feature sequence is any one of the k feature sequences, the method for determining the quality evaluation value corresponding to each feature sequence among the k feature sequences is the same as that for determining the quality evaluation value corresponding to the first feature sequence. In this embodiment, taking the first feature sequence as an example, the quality evaluation value corresponding to the first feature sequence is determined first.
[0055] S402: Determine the information entropy and variance corresponding to the at least one eigenvalue.
[0056] In this embodiment, the information entropy can reflect the uncertainty or randomness of the eigenvalues. By calculating the information entropy, we can understand the degree of uniformity of the eigenvalue distribution and the amount of information contained. If the information entropy is high, it indicates that the eigenvalue distribution is relatively dispersed and the uncertainty is large, which may mean a higher degree of diversity and complexity of the features. Conversely, if the information entropy is low, it indicates that the eigenvalue distribution is relatively concentrated and the certainty is strong. The variance reflects the degree of dispersion of the eigenvalues relative to the average value. A larger variance indicates that the eigenvalue distribution is relatively dispersed, which may imply a larger range of feature changes and poorer stability. A smaller variance indicates that the eigenvalues are relatively close to the average value and are relatively stable. Considering the information entropy and variance comprehensively can more comprehensively evaluate the characteristics and quality of the feature sequence, which helps to subsequently determine the quality evaluation value of the feature sequence based on these evaluation values, so as to screen out feature sequences with higher quality, greater representativeness and usefulness, and improve the accuracy and reliability of the entire text recognition process.
[0057] In this embodiment, first determine the information entropy corresponding to the at least one eigenvalue. Specifically, for the selected at least one eigenvalue, assume that these eigenvalues form a data set. Then, calculate the probability of occurrence of each different eigenvalue, and then calculate the information entropy corresponding to the at least one eigenvalue through the following formula:
[0058] H = -∑p i log2(p i )
[0059] where H is the information entropy corresponding to the at least one eigenvalue, and p i is the probability of occurrence of the i-th eigenvalue.
[0060] Then determine the variance corresponding to the at least one eigenvalue. Specifically, first, calculate the mean value of the selected at least one eigenvalue, that is, the sum of all eigenvalues divided by the number of eigenvalues. Then, for each eigenvalue, calculate the difference between it and the mean value, and square the difference. Finally, divide the sum of all squared differences by the number of eigenvalues, and the result is the variance, which reflects the degree of dispersion of these eigenvalues relative to the mean value.
[0061] S403: Determine a first reference quality evaluation value corresponding to the information entropy and a second reference quality evaluation value corresponding to the variance.
[0062] In this embodiment, it may be a mapping relationship between a preset information entropy and a reference quality evaluation value. Based on this mapping relationship, the first reference quality evaluation value corresponding to the information entropy can be determined. It may be a mapping relationship between the variance and the reference quality evaluation value. Based on this mapping relationship, the second reference quality evaluation value corresponding to the variance can be determined.
[0063] S404: Determine a quality evaluation value corresponding to the first feature sequence based on the first reference quality evaluation value and the second reference quality evaluation value.
[0064] In this embodiment, exemplarily, determine a first reference weight value corresponding to the first reference quality evaluation value and a second reference weight value corresponding to the second reference quality evaluation value. The sum of the first reference weight value and the second reference weight value is 1. Specifically, it may be a mapping relationship between a preset reference quality evaluation value and a reference weight value. Based on this mapping relationship, the first reference weight value corresponding to the first reference quality evaluation value and the second reference weight value corresponding to the second reference quality evaluation value can be determined.
[0065] Exemplarily, determine the number of features in the first feature sequence. Specifically, the number of features may be related to the importance or representativeness of the feature sequence. A larger number of features may mean that the feature sequence contains richer information, but there may also be redundancy. A smaller number of features may be more concise, but the information may not be comprehensive enough. By determining the number of features, the weights can be optimized accordingly to more reasonably balance the roles of different feature sequences in quality evaluation. The number of features reflects the complexity of the feature sequence to a certain extent. By considering the number of features, different complex feature sequences can be better adapted to ensure that the quality evaluation can accurately reflect its contribution to text recognition. In actual calculation and processing, the number of features will affect the calculation cost and efficiency. Understanding the number of features helps to reasonably control the allocation of computing resources on the premise of ensuring the evaluation accuracy. Therefore, it is necessary to determine the number of features in the first feature sequence.
[0066] Exemplarily, determine a target optimization factor corresponding to the number of features. Exemplarily, it may be a mapping relationship between a preset number of features and an optimization factor. Based on this mapping relationship, the target optimization factor corresponding to the number of features can be determined.
[0067] Exemplarily, optimize the first reference weight value based on the target optimization factor to obtain a first target weight value. Specifically, the first target weight value is calculated according to the following formula:
[0068] The first target weight = the first reference weight × (1 + the target optimization factor);
[0069] According to the above formula, the first reference weight can be optimized based on the target optimization factor to obtain the first target weight.
[0070] Exemplarily, the second reference weight is adjusted based on the first target weight to obtain the second target weight. The sum of the first target weight and the second target weight is 1. Since the sum of the first reference weight and the second reference weight is 1, and the sum of the first target weight and the second target weight is also 1, after obtaining the first target weight, the second reference weight can be adjusted based on the first target weight to obtain the second target weight.
[0071] Exemplarily, a weighted calculation is performed based on the first reference quality evaluation value, the second reference quality evaluation value, the first target weight, and the second target weight to obtain the quality evaluation value corresponding to the first feature sequence. Specifically, the quality evaluation value corresponding to the first feature sequence is calculated according to the following formula:
[0072] The quality evaluation value corresponding to the first feature sequence = the first reference quality evaluation value × the first target weight + the second reference quality evaluation value × the second target weight;
[0073] According to the above formula, a weighted calculation can be performed based on the first reference quality evaluation value, the second reference quality evaluation value, the first target weight, and the second target weight to obtain the quality evaluation value corresponding to the first feature sequence.
[0074] It can be seen that by introducing a dynamic adjustment mechanism for weights, weights can be flexibly allocated according to different feature situations, better adapting to various complex data features, improving the adaptability and accuracy of quality evaluation. Considering the number of features and their corresponding optimization factors to adjust the weights can more finely balance the importance of different features, thereby improving the accuracy of quality evaluation and more accurately reflecting the actual contribution of the feature sequence to text recognition. Considering the number of features and their corresponding optimization factors to adjust the weights can more finely balance the importance of different features, thereby improving the accuracy of quality evaluation and more accurately reflecting the actual contribution of the feature sequence to text recognition. Controlling the allocation of computing resources according to the number of features can optimize the computing cost and efficiency on the premise of ensuring evaluation accuracy, avoiding unnecessary resource waste. A more reasonable weight allocation helps reduce evaluation fluctuations caused by feature differences, making the performance of the model more stable on different data, providing a more targeted basis for model optimization, and being able to more effectively adjust feature selection and model structure according to the quality evaluation results.
[0075] Since the first feature sequence is any one of the k feature sequences, the quality evaluation value corresponding to each of the k feature sequences can be determined in the same way as the quality evaluation value corresponding to the first feature sequence, obtaining k quality evaluation values.
[0076] It can be seen that by combining information entropy and variance to determine the quality evaluation value, the quality of the feature sequence can be comprehensively evaluated from different perspectives. Information entropy reflects the uncertainty and randomness of feature values, and variance reflects the dispersion degree of feature values. Considering both can more accurately measure the advantages and disadvantages of features. A single evaluation index may have limitations, while considering information entropy and variance simultaneously can make up for their respective deficiencies, thereby improving the accuracy and reliability of the quality evaluation of the feature sequence, helping to screen out feature sequences with higher quality, and thus improving the performance and prediction ability of the model in subsequent model training and applications. Accurately evaluating feature quality can avoid the interference of low-quality features on the model, make the model more stable, and reduce model fluctuations and errors caused by poor feature quality, providing a clear basis for selecting the most representative and effective feature sequences, helping to reduce the feature dimension, lower the computational complexity, and improve the data processing efficiency.
[0077] S303: Determine n quality evaluation values among the k quality evaluation values that are greater than the quality evaluation threshold.
[0078] In this embodiment, after obtaining the k quality evaluation values corresponding to the k feature sequences respectively, these k quality evaluation values are compared with the quality evaluation threshold respectively, and the quality evaluation values that are greater than the quality evaluation threshold are screened out from these k values to obtain n quality evaluation values.
[0079] It can be seen that by screening out the part with quality evaluation values higher than the threshold, the feature sequences that are considered more valuable and more likely to have a positive impact on the final result can be concentrated for processing and analysis, reducing unnecessary calculations and processing, thereby improving the operating efficiency of the entire system. At the same time, it also helps to improve the accuracy of the final result. The quality evaluation threshold can help exclude those feature sequences with poor quality, which may contain more noise or are not representative, so as to avoid the interference of these low-quality data on subsequent analysis and decision-making, and make subsequent processing based on more reliable and meaningful data. Only focusing on the feature sequences corresponding to the n quality evaluation values with higher quality can more effectively allocate computational resources, storage resources, and time resources, concentrate limited resources on the more potential part, improve resource utilization efficiency, and excluding low-quality feature sequences can reduce the negative impact of outliers or bad data on model training and prediction, make the model more robust and stable, and improve its generalization ability in different scenarios.
[0080] S304: Determine the n feature sequences corresponding to the n quality evaluation values.
[0081] In this embodiment, after determining the n quality evaluation values greater than the quality evaluation threshold among the k quality evaluation values, the feature sequence corresponding to each quality evaluation value among the n quality evaluation values can be determined, so as to determine the n feature sequences with better quality.
[0082] It can be seen that through quality evaluation and threshold screening, it is possible to accurately select n truly valuable and high-quality feature sequences from numerous extracted feature sequences, avoid the interference of low-quality or irrelevant features to subsequent processing, reduce the number of features participating in subsequent processing, reduce the computational complexity, save computational resources and time costs, especially when dealing with large-scale data or complex models, the effect is remarkable. Using high-quality feature sequences for subsequent analysis and modeling can improve the accuracy, robustness and generalization ability of the model, enable the model to better learn the key information and patterns in the data, avoid excessive features causing the model to overfit the training data, thereby improving the performance and adaptability of the model on new data, and help adjust the architecture and parameters of the model according to the feature quality evaluation results, making the model more suitable for effective features and improving the efficiency and effect of the model.
[0083] S203: Label each of the n feature sequences to obtain n sequence labels.
[0084] In this embodiment, sequence labeling is to assign specific marks or labels to each extracted feature sequence. These marks or labels are descriptions of the meaning or feature type represented by the feature sequence. For example, if the feature sequence represents a certain morphological feature of the text in an image, then the sequence label may be "clear text form", "blurred text form", "curved text stroke", etc. It should be noted that the purpose of sequence labeling is to convert the abstract feature sequence into a representation form with clear semantic or classification information, so that subsequent models (such as recurrent neural network models) can process and analyze these labels, so as to realize the conversion from features to the final target text. Sequence labeling is usually determined based on specific rules, task requirements and prior knowledge, and needs to be consistent and accurate to ensure the reliability and effectiveness of subsequent processing.
[0085] In this embodiment, first, the purpose and rules of annotation need to be clarified, which are usually formulated based on specific application scenarios and task requirements. For example, in the case of text recognition tasks in images, the annotation rules may be based on the categories, attributes, etc. of characters, words, or sentences. When starting the annotation, for each feature sequence, first, the features are understood, and the information and feature patterns contained in the feature sequence are carefully analyzed and understood. This may require in-depth research on aspects such as the numerical values, distributions, and structures in the feature sequence. Then, referring to the annotation rules, according to the pre-determined annotation rules and standards, it is determined to which category the feature sequence should belong or what label should be assigned. Finally, the annotation content is selected, and a suitable annotation item is selected from the pre-defined annotation set to label this feature sequence.
[0086] Exemplarily, each of the n feature sequences is annotated to obtain n sequence annotations. Specifically, first, we need to clarify the annotation criteria and rules, which may be formulated based on previous research, the specific requirements of the task, or an in-depth understanding of the data features. For example, if the feature sequence is about the color features of objects in an image, the annotation rule may be to label the feature sequence of a red object as "red", the feature sequence of a green object as "green", etc. Then, start annotating the first feature sequence, carefully analyze the specific feature information contained in this feature sequence, and compare and match it with the pre-determined annotation rules. Suppose this feature sequence exhibits a certain specific pattern and conforms to "category A" defined in the annotation rules, then it is labeled as "category A". Next, the second feature sequence is processed in the same way, and its features are also analyzed in depth to determine the appropriate annotation according to the annotation rules. This process is repeated continuously until the nth feature sequence is completed. During the entire annotation process, there may be some cases where the features are not obvious or it is difficult to determine the annotation. At this time, further analysis, reference to more relevant information, or discussion among multiple people may be required to reach a consistent annotation result. Finally, after completing the annotation of the n feature sequences, the corresponding n sequence annotations are obtained, and each sequence annotation accurately describes the features or categories of its corresponding feature sequence.
[0087] It can be seen that the labeled feature sequences can enable the recurrent neural network model to more clearly understand the meaning or category represented by each feature sequence, so as to transcribe and perform subsequent processing more accurately. With the labels, the recurrent neural network model can learn the correspondence between the feature sequences and the final target text faster, reduce the guessing and trial-and-error of the recurrent neural network model, and improve the accuracy of training and prediction. Rich and accurate labels can help the recurrent neural network model better understand different types of feature sequences, so that when facing new and unseen images, it can also more effectively extract and process relevant features, improving the generalization ability of the recurrent neural network model. By analyzing the labeling results, it can be found that the recurrent neural network model has deficiencies or errors in processing certain feature sequences, so as to optimize and adjust the recurrent neural network model targeted. The labels enable the feature sequences extracted by the deep convolutional neural network to better match and fuse with the processing method of the recurrent neural network model, ensuring the smooth operation of the entire system.
[0088] S204: Transcribe the n sequence labels through the recurrent neural network model to obtain a target label sequence.
[0089] In this embodiment, the label sequence refers to a sequence composed of a series of labels with specific orders and meanings. These labels are symbols or identifiers with specific meanings obtained by further summarizing, transforming, or encoding the n sequence labels. The content included in the label sequence depends on the specific task and data characteristics. For example, in an image recognition task, the labels may be the category names representing different objects, scenes, and attributes in the image. In a natural language processing task, the labels may be specific symbols representing word parts of speech, sentence components, semantic roles, etc. Generally speaking, the label sequence is a higher-level, more general and semantic representation form of the sequence labels, facilitating subsequent processing, analysis, or generation of the final target output (such as the target text).
[0090] In this embodiment, the recurrent neural network model has a memory function and can process sequential data. When transcribing n sequential annotations, the recurrent neural network model processes these annotations one by one in order. First, the recurrent neural network model receives the first sequential annotation as input and calculates an intermediate result based on its internal parameters and weights. This intermediate result contains the processing information of the input annotation and a certain initial memory. Then, when processing the second sequential annotation, the recurrent neural network model not only considers the current input annotation but also takes the previously calculated intermediate result (i.e., memory) as additional information to participate in the calculation. This memory mechanism enables the recurrent neural network model to capture the dependencies and context information between sequential annotations. As it processes subsequent sequential annotations one by one, the recurrent neural network model continuously updates its internal state (memory) and comprehensively considers the previous processing results and the current input to generate a new intermediate result. After processing all n sequential annotations, the recurrent neural network model generates an output sequence based on the final internal state. This output sequence is the target label sequence. The way to generate the target label sequence may be through decoding the final state of the recurrent neural network model or by generating partial labels in each processing step and finally combining them into a complete target label sequence. Throughout the process, the recurrent neural network model utilizes its ability to process sequential data, learns the patterns and rules in sequential annotations, and thus realizes the conversion from the input sequential annotations to the target label sequence.
[0091] It can be seen that by transcribing the n sequential annotations through the recurrent neural network model, a target label sequence is obtained. The recurrent neural network is good at processing sequential data, can fully capture the temporal or logical order relationship between sequential annotations, thus better understanding and integrating this information, and can effectively take into account the context information of each sequential annotation. This means not only focusing on a single annotation itself but also combining the annotations before and after it to make more accurate transcription decisions. For n sequential annotations of different lengths, the recurrent neural network can adaptively process them without a fixed limit on the length of the input sequence. It can learn the potential patterns, rules, and semantic relationships from a large amount of sequential annotation data, thereby improving the accuracy and rationality of transcription, helping to generate a target label sequence with internal logic and coherence rather than an isolated and irrelevant label combination. The trained recurrent neural network model can be applied to new similar sequential annotation data, has a certain adaptability and generalization ability, and can transcribe more efficiently compared to manual processing or simple models, reducing human errors and uncertainties and improving the accuracy of the final result.
[0092] It should be noted that, in this embodiment, when an abnormal sequence annotation appears in the n sequence annotations, the abnormal sequence annotation needs to be deleted first and then transcribed to obtain the target tag sequence. Please refer to Figure 5 , Figure 5 which is a flowchart for determining a target tag sequence provided by an embodiment of the present application, including but not limited to the following steps:
[0093] S501: Determine p sequence annotations with abnormalities among the n sequence annotations.
[0094] In this embodiment, p is an integer less than n. Exemplarily, obtain the historical images in the text recognition system, and determine at least one image in the historical images whose similarity to the target image is greater than a preset similarity. Specifically, first, search for all previously processed and saved images in the database or storage area of the text recognition system. These historical images may be stored and managed according to a certain classification, time sequence, or other identifiers. Through the preset interface or data access method of the system, read the data of these historical images and extract them for subsequent analysis and processing. Then, determine at least one image in the historical images whose similarity to the target image is greater than the preset similarity. This generally involves using methods for calculating image similarity. Common methods include pixel comparison, feature extraction and matching, etc. Representative features such as color, shape, texture, etc. can be extracted from the target image first. Then, the same feature extraction operation is performed on each historical image. Next, a specific similarity metric algorithm, such as Euclidean distance, cosine similarity, etc., is used to calculate the similarity value between the features of the target image and the features of each historical image. Preset a similarity threshold. If the calculated similarity value between a certain historical image and the target image is greater than this threshold, it is considered that this historical image meets the requirements. Repeat this process to find all historical images that meet the similarity greater than the preset value, and the number is at least one. Thus, at least one image in the historical images whose similarity to the target image is greater than the preset similarity can be determined.
[0095] Exemplarily, obtain the historical image recognition data corresponding to the at least one image. Specifically, these historical image recognition data contain the past processing and analysis results of similar images, which can serve as an important reference for the current processing, helping to discover possible rules and patterns. To determine the occurrence frequency of the current n sequence annotations in similar situations, sufficient relevant data is required. The historical image recognition data provides such a rich data source, making the frequency statistics more reliable and representative. By comparing the occurrence frequency of the current sequence annotations in the historical data, those annotations with extremely low frequencies can be detected, thereby discovering possible problems or abnormal situations. Therefore, it is necessary to obtain the historical image recognition data corresponding to the at least one image.
[0096] Exemplarily, determine the occurrence frequency of the n sequence annotations in the historical image recognition data to obtain n frequency values. Specifically, first, parse and organize the historical image recognition data, which may be stored in various formats, such as records in a database, information in a text file, etc. Then, for each sequence annotation, search and count in the organized historical data. For example, for the first sequence annotation, traverse all the historical image recognition data. Whenever this sequence annotation is encountered, increment its occurrence count by 1. After completing the search of all historical data, obtain the occurrence count of this sequence annotation. Then, divide this occurrence count by the total number of historical image recognition data to obtain the occurrence frequency of this sequence annotation. In the same way, calculate the occurrence frequencies of the n sequence annotations respectively to obtain n frequency values. In this way, the occurrence frequency of each sequence annotation in the historical image recognition data can be accurately determined.
[0097] Exemplarily, determine p frequency values among the n frequency values that are less than a preset frequency value. Specifically, the frequency values of the n sequence annotations in the historical image recognition data have been obtained. Then, set a preset frequency value as the judgment criterion. Compare these n frequency values with the preset frequency value respectively. Find out those frequency values that are less than the preset frequency value, and the number is p, so as to obtain p frequency values among the n frequency values that are less than the preset frequency value.
[0098] Exemplarily, determine the sequence annotations corresponding to the p frequency values to obtain the p sequence annotations. Specifically, because each frequency value corresponds to a specific sequence annotation. After determining p frequency values that are less than the preset frequency value, according to the previously established correspondence between the frequency values and the sequence annotations, find out the sequence annotations corresponding to these p frequency values. In this way, the p sequence annotations that are considered abnormal or uncommon are obtained.
[0099] It can be seen that by comparing and analyzing with historical images, abnormal or uncommon sequence annotations that may exist in the target image being currently processed can be discovered, thereby correcting and optimizing the recognition results, improving the accuracy of the final text recognition. Using historical data as a reference can reduce errors caused by factors such as the particularity or noise of the target image, making the performance of the text recognition system more stable and reliable in different situations. It can reveal problems that may be overlooked in the current processing. For example, an abnormally low occurrence frequency of certain sequence annotations may imply deficiencies of the model in specific scenarios, providing clues for further improving the system. According to the analysis of historical data and current sequence annotations, the parameters of the model can be adjusted and the algorithm can be improved targeted to enhance the performance and adaptability of the system, avoiding processing all sequence annotations equally, but focusing on those parts that may be abnormal, thereby saving computing resources and improving processing efficiency. This method of analysis and comparison based on historical data provides a certain explanation and basis for the results of text recognition, making the entire recognition process more transparent and understandable.
[0100] S502: Delete the p sequence annotations among the n sequence annotations to obtain n - p sequence annotations.
[0101] In this embodiment, an index list can be established for the n sequence annotations. After determining the p sequence annotations to be deleted, these annotations are removed from the original sequence annotation set according to their corresponding indexes. A flag bit can also be set for each sequence annotation, initially valid. After determining the p sequence annotations to be deleted, the corresponding flag bits are set to invalid. In subsequent processing, only the sequence annotations marked as valid are selected, thereby achieving the effect of deleting the p sequence annotations. A new set or array can also be created, and the n - p sequence annotations other than the p sequence annotations to be deleted among the n sequence annotations are copied or added to this new set. Any of the above methods can be used to delete the p sequence annotations among the n sequence annotations to obtain n - p sequence annotations.
[0102] It can be seen that removing abnormal sequence annotations can reduce the interference of noise and incorrect data on subsequent processing, thereby improving the overall quality and reliability of the data, inputting cleaner and more accurate sequence annotations to the recurrent neural network model, which helps the model learn more representative and regular patterns, improving the accuracy of the transcription results. Abnormal sequence annotations may lead to incorrect learning and prediction by the model. Deleting them can prevent the model from being misled, making the training and prediction of the model more targeted and effective, reducing the amount of data input to the model, lowering the computational complexity, and improving the computational efficiency. Especially when dealing with large-scale data, it can save time and computational costs. By removing abnormal data that may affect the generalization ability of the model, the model can better adapt to new and unseen data, improving its performance and generality in different scenarios.
[0103] S503: Transcribe the n-p sequence annotations through the recurrent neural network model to obtain the target label sequence.
[0104] In this embodiment, first, the n-p sequence annotations obtained after screening and processing are input into the recurrent neural network model. The neurons in the recurrent neural network model receive the first sequence annotation as input and perform calculations based on its internal weights and biases to generate an output. This output depends not only on the currently input sequence annotation but also on the information stored in the internal memory unit of the model from previous inputs. Then, when the second sequence annotation is input, the model combines the previous calculation results and the newly input annotation to update its internal state and generate a new output. This process is repeated in sequence. The model continuously receives new sequence annotations, updates its internal state, and generates corresponding outputs. After processing all the n-p sequence annotations, the model generates a continuous output sequence based on the final internal state. This output sequence is the target label sequence. Throughout the transcription process, the recurrent neural network model utilizes its learning and memory capabilities for sequence data to capture the long-term dependence relationships and patterns between sequence annotations, thereby generating meaningful and accurate target label sequences.
[0105] It can be seen that removing the abnormal p sequence annotations can reduce the interference of error or noise data on the final result, thereby improving the accuracy and reliability of the target label sequence, reducing the amount of data input to the recurrent neural network model, reducing the processing burden of the model, helping to improve the training and inference speed of the model, optimizing the performance and efficiency of the model. Abnormal annotations may lead to unstable model training or large fluctuations. Deleting them can make the training process of the model smoother, improve the stability and robustness of the model. The remaining n-p sequence annotations are more likely to contain key information valuable for generating the target label sequence, enabling the model to focus more on processing and learning this valid information, thereby generating a more accurate and meaningful target label sequence. By removing abnormal data, the model can better learn the patterns and regularities in normal data, and thus has better generalization ability when facing new and unseen data.
[0106] S205: Determine the target text based on the target label sequence.
[0107] In this embodiment, a comprehensive and detailed mapping table can be created first, which clearly stipulates the correspondence between each possible label and the corresponding text segment or element. When processing the target label sequence, each label is read in turn, and then it is converted into the corresponding text content according to the mapping table. It is also possible to perform decoding based on probability. If each label in the target label sequence has a certain probability value indicating its likelihood of occurrence, then the algorithm will comprehensively consider these probabilities and select the most likely text combination. It is also possible to design a series of templates, each of which stipulates the pattern of the label sequence and the corresponding text structure. The target label sequence is matched with these templates to find the most suitable template. Then, the target text is generated according to the provisions of the template. It is also possible to utilize a trained language model. This model has learned a large number of language rules and patterns. When generating text based on the target label sequence, the language model will predict the next most likely word or character according to the information of the generated partial text and the label sequence, and repeat this process continuously to gradually generate the complete target text. It is also possible to perform grammar and spelling checks on the initially generated text, correct possible errors, check the semantic consistency of the text, and ensure that it is logically reasonable, smooth and easy to understand. It is also possible to further optimize and adjust the text according to specific domain knowledge or context information.
[0108] For example, assume that our target image is a picture containing various fruits. After extracting the feature sequences through a deep convolutional neural network model, five feature sequences (n = 5) are obtained. These five feature sequences are respectively: the feature sequence representing the shape and color of an apple, the feature sequence representing the curved shape of a banana, the feature sequence representing the texture of an orange, the feature sequence representing the distribution of seeds on the surface of a strawberry, and the feature sequence representing the shape of a grape cluster. Then, these five feature sequences are labeled to obtain five sequence labels: such as "apple feature", "banana feature", "orange feature", "strawberry feature", "grape feature". Through a recurrent neural network model, these five sequence labels are transcribed to obtain the target label sequence: "fruit type label". The target text determined based on this target label sequence may be: "The picture contains several fruits including apples, bananas, oranges, strawberries and grapes." Please refer to Figure 6 , Figure 6 FIG. Figure 6 is an example diagram of a target text display interface provided by an embodiment of the present application. In the blank area of the target text display interface 60, the target text: "This beautiful picture contains several attractive fruits including brightly colored apples, curved bananas, fragrant oranges, delicate strawberries and clustered grapes" is directly displayed.
[0109] For example, an image containing a license plate is obtained from a street view camera. The height of the image is scaled to a fixed value (e.g., 32 pixels), and the width is adjusted according to the original aspect ratio to keep the image ratio unchanged. For example, if the original image is 100×200 pixels, after normalization, it may become 32×64 pixels. The normalized image is used as the input and fed into the convolutional layer. The convolutional layer is based on the standard recurrent neural network architecture and extracts the depth feature map of the image through convolutional and pooling operations. Assume that the size of the feature map output by the convolutional layer is 32×16×512 (height×width×number of channels). The feature map is unfolded by columns to generate a sequence of feature vectors, with each column having a width of 1 pixel. Therefore, the length of the generated feature sequence is 16, and the dimension of each feature vector is 512. These sequences of feature vectors will be used as the input to the recurrent layer. The recurrent layer adopts a bidirectional long short-term memory network structure to model the feature sequence from both the forward and backward directions. The bidirectional long short-term memory network structure can capture the context information between characters in the sequence. For example, the character 'B' in 'ABC' depends not only on the previous 'A' but also on the subsequent 'C'. The bidirectional long short-term memory network structure predicts each feature vector and outputs the character probability distribution corresponding to each frame. For example, for a feature vector with an input sequence length of 16, the recurrent layer outputs 16 character probability distributions. The transcription layer uses the connectionist temporal classification algorithm to convert the frame-by-frame predictions of the recurrent layer into the final label sequence. The connectionist temporal classification algorithm decodes the most likely character sequence directly from the sequence by calculating the conditional probability and ignoring the label position information. After being decoded by the connectionist temporal classification algorithm, the model outputs the final recognition result 'ABC123'. In the dictionary-free mode, the model directly outputs the sequence with the highest probability; in the dictionary-based mode, the model will further optimize the result in combination with the dictionary to improve the accuracy and recognition efficiency.
[0110] It can be seen that the complex image processing and analysis process is finally transformed into clear and understandable target text and directly presented in the display interface, enabling users to intuitively obtain key information without having to laboriously interpret complex data or charts. Through detailed and vivid descriptions, such as words like 'bright color', 'curved shape', and 'fragrant smell', the characteristics of the fruits in the picture are conveyed more vividly, enabling users to understand the picture content more vividly. The simple and clear display interface and attractive text description enable users to quickly and easily obtain the required information, reducing the cognitive burden of users and enhancing the satisfaction of users when using related systems or applications. If there are multiple similar target images and their corresponding texts, this unified display method helps users make horizontal comparisons and analyses to discover the differences and characteristics between different pictures.
[0111] In summary, implementing the embodiments of the present invention has the following beneficial effects:
[0112] It can be seen that the text recognition method based on sequence recognition described in the embodiments of the present invention is applied to a text recognition system, and the text recognition system includes a deep convolutional neural network model and a recurrent neural network model. The method includes: obtaining a target image, extracting a feature sequence from the target image through the deep convolutional neural network model to obtain n feature sequences, where n is an integer greater than 1, labeling each of the n feature sequences to obtain n sequence labels, transcribing the n sequence labels through the recurrent neural network model to obtain a target label sequence, and determining a target text based on the target label sequence, thereby improving the text recognition efficiency.
[0113] Please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of a text recognition device based on sequence recognition provided by an embodiment of the present application, which is applied to a text recognition system. The text recognition system includes a deep convolutional neural network model and a recurrent neural network model. The device includes: an obtaining unit 701 and a processing unit 702;
[0114] The obtaining unit 701 is configured to obtain a target image;
[0115] The processing unit 702 is configured to extract a feature sequence from the target image through the deep convolutional neural network model to obtain n feature sequences; n is an integer greater than 1;
[0116] Label each of the n feature sequences to obtain n sequence labels;
[0117] Transcribe the n sequence labels through the recurrent neural network model to obtain a target label sequence;
[0118] Determine a target text based on the target label sequence.
[0119] In some possible embodiments, in terms of extracting a feature sequence from the target image through the deep convolutional neural network model to obtain n feature sequences, the processing unit 702 is specifically configured to:
[0120] Extract a feature sequence from the target image through the deep convolutional neural network model to obtain k feature sequences; k is an integer greater than or equal to n;
[0121] Determine a quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values;
[0122] Determine n quality evaluation values among the k quality evaluation values that are greater than a quality evaluation threshold;
[0123] Determine the n feature sequences corresponding to the n quality evaluation values.
[0124] In some possible implementation manners, in terms of determining the quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values, the processing unit 702 is specifically configured to:
[0125] Determine at least one feature value in the first feature sequence; the first feature sequence is any one of the k feature sequences; the feature value is a numerical value existing in the first feature sequence;
[0126] Determine the information entropy and variance corresponding to the at least one feature value;
[0127] Determine a first reference quality evaluation value corresponding to the information entropy and a second reference quality evaluation value corresponding to the variance;
[0128] Determine the quality evaluation value corresponding to the first feature sequence based on the first reference quality evaluation value and the second reference quality evaluation value.
[0129] In some possible implementation manners, in terms of determining the quality evaluation value corresponding to the first feature sequence based on the first reference quality evaluation value and the second reference quality evaluation value, the processing unit 702 is specifically configured to:
[0130] Determine a first reference weight corresponding to the first reference quality evaluation value and a second reference weight corresponding to the second reference quality evaluation value; the sum of the first reference weight and the second reference weight is 1;
[0131] Determine the number of features in the first feature sequence;
[0132] Determine a target optimization factor corresponding to the number of features;
[0133] Optimize the first reference weight based on the target optimization factor to obtain a first target weight;
[0134] Adjust the second reference weight based on the first target weight to obtain a second target weight; the sum of the first target weight and the second target weight is 1;
[0135] Perform weighted calculation based on the first reference quality evaluation value, the second reference quality evaluation value, the first target weight, and the second target weight to obtain the quality evaluation value corresponding to the first feature sequence.
[0136] In some possible embodiments, when calculating the weighted sum based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain the quality assessment value corresponding to the first feature sequence, the processing unit 702 is specifically configured to:
[0137] Calculate a weighted sum based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain a third reference quality assessment value;
[0138] Obtain the sharpness and contrast of the target image;
[0139] Determine target fine-tuning parameters corresponding to the sharpness and the contrast;
[0140] Adjust the third reference quality assessment value based on the fine-tuning parameters to obtain the quality assessment value corresponding to the first feature sequence.
[0141] In some possible embodiments, the processing unit 702 is further specifically configured to:
[0142] Determine p sequence annotations that are abnormal among the n sequence annotations; p is an integer less than n;
[0143] Delete the p sequence annotations from the n sequence annotations to obtain n - p sequence annotations;
[0144] Transcribe the n - p sequence annotations through the recurrent neural network model to obtain the target label sequence.
[0145] In some possible embodiments, when determining p sequence annotations that are abnormal among the n sequence annotations, the processing unit 702 is specifically configured to:
[0146] Obtain historical images in the text recognition system;
[0147] Determine at least one image in the historical images whose similarity to the target image is greater than a preset similarity;
[0148] Obtain historical image recognition data corresponding to the at least one image;
[0149] Determine the frequencies of occurrence of the n sequence annotations in the historical image recognition data to obtain n frequency values;
[0150] Determine p frequency values that are less than a preset frequency value among the n frequency values;
[0151] Determine the sequence annotations corresponding to the p frequency values to obtain the p sequence annotations.
[0152] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in Figure 8 , the electronic device 800 includes a transceiver 801, a processor 802, and a memory 803. They are connected through a bus 804. The memory 803 is used to store computer programs and data, and the transceiver 801 can transmit the data stored in the memory 803 to the processor 802. The above program includes instructions for performing the following steps:
[0153] Obtain a target image;
[0154] Extract feature sequences from the target image through the deep convolutional neural network model to obtain n feature sequences; n is an integer greater than 1;
[0155] Annotate each of the n feature sequences to obtain n sequence annotations;
[0156] Transcribe the n sequence annotations through the recurrent neural network model to obtain a target label sequence;
[0157] Determine a target text based on the target label sequence.
[0158] In some possible embodiments, in terms of extracting n feature sequences from the target image through the deep convolutional neural network model, the above program includes instructions for performing the following steps:
[0159] Extract k feature sequences from the target image through the deep convolutional neural network model; k is an integer greater than or equal to n;
[0160] Determine a quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values;
[0161] Determine n quality evaluation values among the k quality evaluation values that are greater than the quality evaluation threshold;
[0162] Determine the n feature sequences corresponding to the n quality evaluation values.
[0163] In some possible embodiments, in terms of determining a quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values, the above program includes instructions for performing the following steps:
[0164] Determine at least one eigenvalue in a first feature sequence; the first feature sequence is any one of the k feature sequences; the eigenvalue is a value existing in the first feature sequence;
[0165] Determine the information entropy and variance corresponding to the at least one eigenvalue;
[0166] Determine a first reference quality assessment value corresponding to the information entropy and a second reference quality assessment value corresponding to the variance;
[0167] Determine the quality assessment value corresponding to the first feature sequence based on the first reference quality assessment value and the second reference quality assessment value.
[0168] In some possible implementation manners, in terms of determining the quality assessment value corresponding to the first feature sequence based on the first reference quality assessment value and the second reference quality assessment value, the above program includes instructions for performing the following steps:
[0169] Determine a first reference weight corresponding to the first reference quality assessment value and a second reference weight corresponding to the second reference quality assessment value; the sum of the first reference weight and the second reference weight is 1;
[0170] Determine the number of features in the first feature sequence;
[0171] Determine a target optimization factor corresponding to the number of features;
[0172] Optimize the first reference weight based on the target optimization factor to obtain a first target weight;
[0173] Adjust the second reference weight based on the first target weight to obtain a second target weight; the sum of the first target weight and the second target weight is 1;
[0174] Perform a weighted calculation based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain the quality assessment value corresponding to the first feature sequence.
[0175] In some possible implementation manners, in terms of performing a weighted calculation based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain the quality assessment value corresponding to the first feature sequence, the above program includes instructions for performing the following steps:
[0176] Perform a weighted calculation based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain a third reference quality assessment value;
[0177] Obtain the sharpness and contrast of the target image;
[0178] Determine a target fine-tuning parameter corresponding to the sharpness and the contrast;
[0179] Adjust the third reference quality assessment value based on the fine-tuning parameter to obtain the quality assessment value corresponding to the first feature sequence.
[0180] In some possible implementation manners, the above program includes instructions for performing the following steps:
[0181] Determine p sequence annotations that are abnormal among the n sequence annotations; p is an integer less than n;
[0182] Delete the p sequence annotations from the n sequence annotations to obtain n - p sequence annotations;
[0183] Transcribe the n - p sequence annotations through the recurrent neural network model to obtain the target label sequence.
[0184] In some possible implementation manners, in terms of determining p sequence annotations that are abnormal among the n sequence annotations, the above program includes instructions for performing the following steps:
[0185] Obtain historical images in the text recognition system;
[0186] Determine at least one image in the historical images whose similarity to the target image is greater than a preset similarity;
[0187] Obtain historical image recognition data corresponding to the at least one image;
[0188] Determine the frequencies at which the n sequence annotations appear in the historical image recognition data to obtain n frequency values;
[0189] Determine p frequency values among the n frequency values that are less than a preset frequency value;
[0190] Determine the sequence annotations corresponding to the p frequency values to obtain the p sequence annotations.
[0191] It should be understood that the electronic devices in the present application may include text recognition devices based on sequence recognition, smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, palmtop computers, laptop computers, mobile Internet devices MID (Mobile Internet Devices, abbreviated as MID) or wearable devices, or servers, edge computing nodes, etc. The above electronic devices are only examples and not exhaustive, including but not limited to the above electronic devices.
[0192] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, which is executed by a processor to implement part or all of the steps of any one of the text recognition methods based on sequence recognition described in the above method embodiments.
[0193] Embodiments of the present application also provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the text recognition methods based on sequence recognition described in the above method embodiments.
[0194] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0195] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0196] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0197] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] In addition, the functional units in various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software program modules.
[0199] When the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. And the aforementioned memory includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other media that can store program codes.
[0200] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories (abbreviation: ROM, English: Read-Only Memory), random access memories (abbreviation: RAM, English: Random Access Memory), magnetic disks, or optical discs, etc.
[0201] The above has introduced the embodiments of this application in detail. Specific examples are used in this article to elaborate on the principle and embodiments of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific embodiments and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A text recognition method based on sequence recognition, characterized in that: Applied to a text recognition system, the text recognition system includes a deep convolutional neural network model and a recurrent neural network model, and the method includes: Get the target image; Extracting feature sequences from the target image using the deep convolutional neural network model to obtain n feature sequences; n is an integer greater than 1; Annotate each of the n feature sequences to obtain n sequence annotations; The n sequence tags are transcribed by the recurrent neural network model to obtain a target label sequence; A target text is determined based on the target tag sequence.
2. The method according to claim 1, characterized in that The extracting feature sequences of the target image by the deep convolutional neural network model to obtain n feature sequences includes: Extracting feature sequences from the target image using the deep convolutional neural network model to obtain k feature sequences, where k is an integer greater than or equal to n; Determine a quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values; Determining n quality assessment values that are greater than a quality assessment threshold among the k quality assessment values; The n feature sequences corresponding to the n quality evaluation values are determined.
3. The method according to claim 2, characterized in that The step of determining a quality evaluation value corresponding to each of the k feature sequences to obtain k quality evaluation values comprises: Determine at least one characteristic value in a first characteristic sequence; the first characteristic sequence is any one of the k characteristic sequences; the characteristic value is a numerical value existing in the first characteristic sequence; Determining information entropy and variance corresponding to the at least one eigenvalue; Determine a first reference quality assessment value corresponding to the information entropy and a second reference quality assessment value corresponding to the variance; A quality assessment value corresponding to the first feature sequence is determined based on the first reference quality assessment value and the second reference quality assessment value.
4. The method according to claim 3, characterized in that The determining, based on the first reference quality evaluation value and the second reference quality evaluation value, a quality evaluation value corresponding to the first feature sequence includes: Determine a first reference weight corresponding to the first reference quality assessment value and a second reference weight corresponding to the second reference quality assessment value; the sum of the first reference weight and the second reference weight is 1; determining the number of features in the first feature sequence; Determining a target optimization factor corresponding to the number of features; Optimizing the first reference weight based on the target optimization factor to obtain a first target weight; The second reference weight is adjusted based on the first target weight to obtain a second target weight; the sum of the first target weight and the second target weight is 1; A weighted calculation is performed based on the first reference quality evaluation value, the second reference quality evaluation value, the first target weight, and the second target weight to obtain a quality evaluation value corresponding to the first feature sequence.
5. The method according to claim 4, characterized in that The performing weighted calculation based on the first reference quality evaluation value, the second reference quality evaluation value, the first target weight, and the second target weight to obtain a quality evaluation value corresponding to the first feature sequence includes: Perform weighted calculation based on the first reference quality assessment value, the second reference quality assessment value, the first target weight, and the second target weight to obtain a third reference quality assessment value; Acquiring the clarity and contrast of the target image; determining a target fine-tuning parameter corresponding to the clarity and the contrast; The third reference quality assessment value is adjusted based on the fine-tuning parameter to obtain a quality assessment value corresponding to the first feature sequence.
6. The method according to any one of claims 2 to 5, characterized in that: The method further comprises: Determine p sequence annotations that are abnormal among the n sequence annotations; p is an integer less than n; Deleting the p sequence labels from the n sequence labels to obtain np sequence labels; The np sequence annotations are transcribed through the recurrent neural network model to obtain the target label sequence.
7. The method according to claim 6, characterized in that The determining of p sequence labels having abnormalities among the n sequence labels includes: Acquire historical images in the text recognition system; Determine at least one image in the historical images whose similarity to the target image is greater than a preset similarity; Acquiring historical image recognition data corresponding to the at least one image; Determine the frequency at which the n sequence annotations appear in the historical image recognition data to obtain n frequency values; Determine p frequency values among the n frequency values that are less than a preset frequency value; Determine the sequence labels corresponding to the p frequency values to obtain the p sequence labels.
8. A text recognition device based on sequence recognition, characterized in that: Applied to a text recognition system, the text recognition system includes a deep convolutional neural network model and a recurrent neural network model, and the device includes: an acquisition unit and a processing unit; The acquisition unit is used to acquire a target image; The processing unit is used to extract feature sequences from the target image through the deep convolutional neural network model to obtain n feature sequences; n is an integer greater than 1; Annotate each of the n feature sequences to obtain n sequence annotations; The n sequence tags are transcribed by the recurrent neural network model to obtain a target label sequence; A target text is determined based on the target tag sequence.
9. An electronic device, characterized in that: The method comprises a processor, a memory, a communication interface and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the one or more programs include instructions for executing the steps in the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 7.