Multi-scale text recognition method, electronic equipment, storage medium and program product

By using deep learning models and knowledge graph technology, the problem of low accuracy in multi-scale text recognition of traditional OCR has been solved, achieving efficient and accurate recognition of multi-scale text, adapting to different font styles and reducing human intervention.

CN120913232APending Publication Date: 2025-11-07AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511014224.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional OCR technology has low accuracy when recognizing multi-scale text, especially in unstructured documents where there are problems such as tilting, blurring, or shadow interference, resulting in low efficiency and easy errors in manual data entry.

Method used

By employing deep learning models combined with Transformer and knowledge graph technologies, image distortion is corrected and sharpness is adjusted through image optimization, text recognition models, and semantic recognition. Furthermore, the model is trained using multi-style samples to identify and correct multi-scale text.

Benefits of technology

It achieves accurate recognition of multi-scale text, improves recognition efficiency, reduces manual intervention, and enhances adaptability and recognition accuracy to different font styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913232A_ABST
    Figure CN120913232A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-scale text recognition method, electronic equipment, a storage medium and a program product. The method comprises the steps of obtaining image information carrying a to-be-recognized multi-scale text, and performing optimization processing on the image information to obtain an optimized image; an image range in which a to-be-recognized text exists in the optimized image is determined, the to-be-recognized text in the image range is recognized through a preset character recognition model, a first recognition result is obtained, the character recognition model is obtained through machine learning of multiple sets of sample data, and the first recognition result is obtained; each group of data in the multiple groups of sample data comprises sample words of different styles and entity words corresponding to the sample words, and the first recognition result at least comprises recognition results of partial texts in the to-be-recognized multi-scale text; and performing semantic recognition on the first recognition result, and determining a recognition result of semantic recognition as a recognition result of the to-be-recognized text. The method is used for achieving the effect of accurately recognizing the multi-scale text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition, and in particular to a multi-scale text recognition method, an electronic device, a storage medium and a program product. BACKGROUND

[0002] With the deepening of the digital transformation of the banking industry, a large number of manually filled forms need to be processed every day for counter services, loan approval, and public account scene. In the traditional manual input mode, the operator needs to manually locate the key information (such as account name, ID number, etc.) from the paper form or scanned copy and then input into the system, which not only consumes a lot of time and is low in efficiency, but also is easily affected by human errors. In addition, the printed font and the first page font in the related form are mixed, and the multi-scale font size and style vary from person to person, further increasing the difficulty of information extraction. Moreover, unstructured documents generally have problems such as inclination, blur, or shadow interference, which further increase the difficulty of the work of business personnel. In the prior art, the recognition accuracy of traditional OCR for irregular fonts is low.

[0003] At present, there is an urgent need for a technical solution that can accurately recognize multi-scale text. SUMMARY

[0004] The embodiments of the present application provide a multi-scale text recognition method, an electronic device, a storage medium and a program product, to achieve the effect of accurately recognizing multi-scale text.

[0005] In a first aspect, the embodiments of the present application provide a multi-scale text recognition method, comprising: obtaining image information carrying multi-scale text to be recognized, and performing optimization processing on the image information to obtain an optimized image, wherein the optimization processing is used to correct the distortion of the image information and / or adjust the clarity of the image information; determining an image range in which the text to be recognized exists in the optimized image, and identifying the text to be recognized in the image range through a pre-set character recognition model to obtain a first recognition result, wherein the character recognition model is obtained through machine learning of a plurality of sample data, wherein each group of data in the plurality of sample data includes sample characters of different styles and entity characters corresponding to the sample characters, and the first recognition result at least includes the recognition result of part of the multi-scale text to be recognized; and performing semantic recognition on the first recognition result, and determining the recognition result of the semantic recognition as the recognition result of the text to be recognized.

[0006] In a second aspect, an embodiment of the present application provides a multi-scale text recognition device, comprising: an acquisition module, configured to acquire image information carrying a multi-scale text to be recognized, and perform optimization processing on the image information to obtain an optimized image, wherein the optimization processing is used to correct distortion of the image information and / or adjust the definition of the image information; a determination module, configured to determine an image range in which the multi-scale text to be recognized exists in the optimized image, and identify the multi-scale text to be recognized in the image range through a preset character recognition model to obtain a first recognition result, wherein the character recognition model is obtained through machine learning of a plurality of sample data, wherein each group of data in the plurality of sample data comprises: a sample character of different styles and an entity character corresponding to the sample character, and the first recognition result at least comprises a recognition result of a part of the multi-scale text to be recognized; and an identification module, configured to perform semantic recognition on the first recognition result, and determine a recognition result of the semantic recognition as a recognition result of the multi-scale text to be recognized.

[0007] In a third aspect, an embodiment of the present application provides a multi-scale text recognition device, comprising: a memory and a processor.

[0008] The memory stores computer execution instructions.

[0009] The processor executes the computer execution instructions stored in the memory, so that the processor executes the first aspect and / or various possible implementation manners of the first aspect.

[0010] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the first aspect and / or various possible implementation manners of the first aspect.

[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the first aspect and / or various possible implementation manners of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.

[0013] Figure 1 A flowchart of a multi-scale text recognition method provided by the present application Figure 1 ;

[0014] Figure 2 A flowchart of a multi-scale text recognition method provided by the present application Figure 2 ;

[0015] Figure 3 The structural schematic diagram of the multi-scale text recognition device provided in the present application is shown in the following figure:

[0016] Figure 4 The structural schematic diagram of the multi-scale text recognition device provided in the present application is shown in the following figure.

[0017] Through the above-mentioned figures, the specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These figures and the written description are not intended to limit the scope of the present application concept in any way, but to illustrate the present application concept to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0018] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.

[0019] First, the terms involved in the present application are explained:

[0020] Unstructured document: Unstructured document refers to an information carrier without fixed format or predefined structure, such as scanned documents, contracts, bills, reports, emails, etc. The content in such documents often exists in the form of natural language, pictures or handwriting, lacking clear field division and format markers, and the data is presented in a mixed manner of text, images, etc. Unstructured documents have diversity and complexity in practical applications, commonly found in enterprise financial statements, contract archiving, administrative documents, etc. Due to the non-uniformity of their format, layout, font, etc., traditional rule matching or template recognition methods are often difficult to apply directly, thus advanced algorithms and intelligent processing means are needed to extract key information, realizing automated data management and analysis.

[0021] Deep learning: Deep learning is a machine learning method based on artificial neural networks, which realizes automatic feature extraction and complex pattern modeling by constructing multi-layer neural network structures. Deep learning technology has achieved great success in image, speech, natural language processing and other fields, and its core idea lies in training network models with large amounts of data, so that they can automatically learn and abstract high-level features from data. Compared with traditional machine learning methods, deep learning does not need manual feature engineering, and can discover the internal laws from raw data. This makes deep learning have significant advantages in processing unstructured data such as document images and natural language texts. Currently, deep learning has become an important technical support for intelligent information processing and automation systems.

[0022] Knowledge Graph: A knowledge graph is a technique for expressing knowledge and information through a graph structure. It organizes and stores entities (such as people, objects, places, events, etc.) and their relationships in the form of nodes and edges. Knowledge graphs not only integrate knowledge from different data sources, but also enable retrieval, reasoning, and updating of knowledge through graph databases, reasoning algorithms, and rule engines. In the information extraction task, knowledge graphs are often used to provide domain background and contextual association information. Through pre-built rules and entity relationships, the extraction results are verified and completed, thereby improving the accuracy and robustness of information extraction. Knowledge graph technology has wide applications in natural language processing, recommendation systems, and intelligent question answering, and is an important tool for realizing semantic understanding and logical reasoning.

[0023] Transformer: Transformer is a deep learning model based on self-attention mechanism. Its core idea is to use multi-head self-attention mechanism to capture global dependencies in input sequences, thereby greatly improving parallel computing capability. Transformer consists of an encoder and a decoder, where the encoder is responsible for mapping input to high-dimensional representation, and the decoder uses these representations to generate output step by step. Each layer contains multi-head self-attention, feedforward neural network, residual connection, and layer normalization. During training, position encoding is usually used to preserve sequence information, and cross-entropy or other loss functions are used to optimize the model. Due to its efficient computing architecture, Transformer has made breakthroughs in machine translation, text generation, OCR recognition, speech processing, and other tasks, and has given birth to powerful deep learning models such as BERT, GPT, and ViT.

[0024] Figure 1 The flowchart of the multi-scale text recognition method provided in the present application Figure 1 As shown in Figure 1 , the method comprises:

[0025] S101, acquiring image information carrying a multi-scale text to be recognized, and performing optimization processing on the image information to obtain an optimized image, wherein the optimization processing is used to correct distortion of the image information and / or adjust the clarity of the image information;

[0026] Optionally, acquiring image information can be to acquire an image containing a multi-scale text to be recognized from input. This image can come from a camera, scanner, or other image acquisition device. Due to the influence of device performance, environmental light, shooting angle, etc., the quality of scanned images and photographed images is often affected by uneven brightness, low contrast, and serious noise interference, which in turn affects the accuracy of subsequent OCR recognition.

[0027] To improve the accuracy of text recognition, the image can be corrected for distortion and adjusted for clarity to make the text in the image easier to recognize. Alternatively, the image is geometrically corrected, for example, the tilt, perspective distortion, etc., so that the text in the image is as horizontal or regularly arranged as possible. The clarity of the image is improved through image enhancement techniques (such as sharpening, denoising, contrast adjustment, etc.) to make the text easier to be recognized.

[0028] S102, determine the image range of the optimization image where the text to be recognized exists, identify the text to be recognized in the image range through a preset text recognition model to obtain a first recognition result, wherein the text recognition model is obtained by machine learning a plurality of sample data.

[0029] Each of the plurality of sample data includes sample characters of different styles and entity characters corresponding to the sample characters, and the first recognition result includes at least the recognition result of the middle part of the multi-scale text to be recognized.

[0030] In the optimized image, the image range where the text to be recognized exists is determined through image segmentation or object detection technology. For example, find the area where the text is located, and exclude the background or other irrelevant content. The text in the image range is identified using a pre-set text recognition model (such as an OCR model). The model is trained through machine learning, and the training data includes sample characters of different styles (such as handwritten, printed, artistic characters, etc.) and their corresponding entity characters (i.e. correct text content). The model outputs a preliminary recognition result, called "first recognition result". Since the text may have multi-scale problems, the first recognition result may only include part of the text recognition result.

[0031] Optionally, the multi-scale text recognition module fuses text features of different scales and deformations, wherein an improved ResNet-50 is used to replace part of the standard convolution with a variability convolution, so that the recognition ability of the model for irregular and curved text is stronger.

[0032] The adversarial font adaptive module separates the content features and font style features, and introduces an adversarial mechanism to effectively suppress the interference caused by font changes and enhance the adaptability of the model to various fonts. It also builds a font prototype memory bank to store the typical features of common fonts (such as SimSun, bold, KaiTi, handwritten, etc.) as key-value pairs. In the training and inference stages, the attention mechanism is used to correct the current input style features, thereby further enhancing the adaptability to different font styles. The calculation formula of the adversarial loss is as follows:

[0033] L adv =E[log D(style)+E[log(1-D(content))

[0034] where D(style) represents the font discriminator predicts the style features of the separated font, and is expected to accurately determine the font type information, and D(content) represents the prediction of the content features, which ideally does not contain font style.

[0035] The visual-linguistic decoder fuses image visual information and text language information in the decoding stage, and dynamically balances the weights of the two through a gating mechanism. This method not only utilizes the fine-grained information of visual features, but also maintains the context semantics of the language model, significantly improving the recognition accuracy of complex text. The gating weight calculation formula is as follows:

[0036] gate = σ(FFN([visual_feat, lang_feat]))

[0037] where [visual_feat, lang_feat] represents the concatenation of visual features and language features in the channel dimension, the visual features are extracted through a convolutional neural network (CNN), and the language features are extracted from the hidden state of the Transformer decoder; FFN (feedforward neural network) processes the concatenated features to learn the fusion strategy, and σ represents mapping the FFN output to the [0, 1] interval to obtain a gating weight representing the proportion of visual features. The final fusion feature is represented as follows:

[0038] final_feat = gate * visual_feat + (1-gate) * lang

[0039] To prevent the gating value from tending to 0 or 1 too early in the early training stage, leading to unbalanced fusion, a regularization term is introduced in the loss function:

[0040]

[0041] This regularization term keeps close to 0.5 during training, ensuring that visual and language features can be balanced and fused, thereby improving the accuracy of the final character recognition.

[0042] In order to enable the model to gradually adapt to complex situations from simple scenes, a progressive training strategy is adopted. In the first stage, samples of a single font and fixed font size are used, and a higher learning rate is adopted, and the main goal is to enable the model to first master basic feature extraction; in the second stage, three fonts and twice the font size difference are introduced, and a medium learning rate is applied, and the purpose is to enable the model to learn more multi-scale features and improve font robustness; in the third stage, ten fonts, three times the font size difference are used, and the lowest learning rate is adopted, to ensure that the model can also achieve good generalization in complex scenes. Through this progressive training strategy, the input image region can be directly mapped to the character sequence, and the corresponding position information is also output, providing accurate text data for the information extraction stage

[0043] S103, performing semantic recognition on the first recognition result, and determining a recognition result of the semantic recognition as the recognition result of the to-be-recognized text.

[0044] Through semantic recognition, errors in the first recognition result can be corrected, or un-recognized part of the text can be supplemented.

[0045] The multi-scale text recognition method provided by the embodiment of the application corrects image distortion and improves clarity through geometric correction and image enhancement technology, providing standardized input for subsequent recognition. By positioning the text region, the text content is extracted in combination with the character recognition model based on machine learning. During model training, multi-style samples (such as handwriting and printed matter) are used to enable the model to have generalization ability, but due to the influence of multi-scale, the model can only recognize part of the text. Language rule checking is performed on the preliminary recognition result to correct incorrect characters or supplement missing parts, and finally a complete and logical text means is output, achieving the effect of accurately recognizing multi-scale text.

[0046] Figure 2 Flowchart of the multi-scale text recognition method provided by the application Figure 2 As shown in Figure 1 the embodiment is based on Figure 3 the embodiment, the multi-scale text recognition method is described in detail, and the method comprises the following steps:

[0047] S201, adjusting image parameters of the image information according to a preset image parameter range, to obtain a preliminary optimized image, wherein the image parameters at least include one of the following: image brightness, image contrast; performing noise reduction processing on the preliminary optimized image to obtain a noise reduction image; detecting the degree of inclination of the noise reduction image, and performing correction processing on the noise reduction image according to the detection result to obtain a corrected image; and determining the corrected image as an optimized image corresponding to the image information.

[0048] The system pre-defines a set of reasonable image parameter ranges, such as brightness, contrast, etc. These parameter ranges are usually set according to image quality standards or actual needs. According to the preset range, the brightness, contrast, etc. of the input image are adjusted to make the image reach a preliminary optimization state. For example: if the image is too dark, increase the brightness; if the image contrast is insufficient, increase the contrast. After parameter adjustment, a preliminary optimized image is obtained.

[0049] The preliminary optimized image is subjected to noise reduction processing to remove noise (such as salt and pepper noise, Gaussian noise, etc.) in the image.

[0050] Specifically: Specifically:

[0051] Step 1: Adjust the brightness according to the preset rules;

[0052] Step 2: Adjust the contrast according to the preset rules;

[0053] Step 3: Adjust the noise level;

[0054] Step 4: Document skew correction;

[0055] Step 5: Target detection.

[0056] First, the input image is quality evaluated, and three key indicators of brightness, contrast and noise level are extracted.

[0057] The noise level calculation formula is as follows:

[0058] The noise level N calculation formula is as follows:

[0059]

[0060] G x (i) represents the gradient value of the i-th pixel in the horizontal direction calculated by the edge detection method such as Sobel operator; G y (i) represents the gradient value of the i-th pixel in the vertical direction, when the noise index is at a low level, it means that the image noise is less, and a simple Gaussian filter can effectively denoise with less impact on image details; when the noise level is medium and the edge information needs to be preserved well, bilateral filtering is used, which can smooth the noise while preserving the edge sharpness; and when the noise level is high, in order to remove noise while preserving image details and texture information as much as possible, non-local mean filtering (NLM) is used. This method performs better in high noise conditions.

[0061] Further, the degree of inclination of the denoised image is detected, and the denoised image is corrected according to the detection result to obtain a corrected image, specifically:

[0062] Exemplarily, edge detection is performed on the denoised image to obtain a set of edge coordinates of edge pixel points of the denoised image; a dispersion degree and a direction of edge pixel point distribution of the denoised image are calculated according to a coordinate mean value corresponding to the set of edge coordinates; the dispersion degree and the direction of edge pixel point distribution determine an inclination angle and a first inclination direction corresponding to the denoised image; and the denoised image is rotated by an angle identical to the inclination angle value according to a second inclination direction, wherein the second inclination direction is opposite to the first inclination direction.

[0063] That is, edge information is extracted from the denoised image to determine edge pixel points in the image. An edge detection algorithm (such as Canny edge detection, Sobel operator, etc.) is used to process the denoised image to obtain a set of edge coordinates of edge pixel points. These edge pixel points are usually located in the outline of the object in the image or the area where the texture changes obviously. The set of edge coordinates contains the coordinates of all edge pixel points. According to the set of edge coordinates, the coordinate mean value of all edge pixel points is calculated. For example, the x-coordinate mean value and the y-coordinate mean value of all points are calculated to obtain a center point. The dispersion degree of the distribution of edge pixel points is measured by calculating the average distance of edge pixel points to the coordinate mean value. The greater the dispersion degree, the more dispersed the distribution of edge pixel points. By analyzing the direction of the distribution of edge pixel points, the main inclination direction of the image is determined. For example, if the edge pixel points are overall shifted to the right side, the image may be inclined to the right.

[0064] According to the dispersion degree and the direction of the edge pixel points, the inclination angle of the image is calculated. For example, if the edge pixel points are overall shifted to the upper right, the image may be inclined to the upper right, and the inclination angle can be calculated by geometry. According to the direction of the distribution of edge pixel points, the inclination direction of the image is determined. For example, if the edge pixel points are overall shifted to the right, the first inclination direction is “right”. The second inclination direction is obtained by taking the opposite of the first inclination direction. For example, if the first inclination direction is “right”, the second inclination direction is “left”. The image is rotated by an angle identical to the inclination angle value according to the second inclination direction. For example, if the inclination angle is 10 degrees, the image is rotated 10 degrees to the left. After rotation, the inclination problem is corrected.

[0065] Exemplarily, for the processing of the inclined document, an image inclination correction method based on global edge information and principal component analysis (PCA) is used, which does not depend on the pre-segmented text area, but directly utilizes the structural features of the entire image for correction. First, the entire image is processed by Canny edge detection to extract the coordinate set of all edge points, The mean value of edge points And The covariance matrix is constructed:

[0066]

[0067] The PCA decomposition is performed on the covariance matrix to obtain the eigenvector v = [v x , v y ] of the largest eigenvalue, and then the image tilt angle is calculated as Then the rotation matrix is constructed as follows:

[0068]

[0069] And the image is rotated by -θ degrees using affine transformation (using bicubic interpolation to maintain image details) to restore the text region to a horizontal state. This method estimates the tilt angle through PCA analysis of global edge features, does not require additional text region segmentation, reduces computational complexity, and can still achieve high-precision tilt correction in the presence of background noise.

[0070] Exemplarily, a preliminary semantic recognition is performed on the first recognition result to obtain a plurality of entity names contained in the first recognition result and an entity logical relationship chain between the plurality of entity names; in a case where the plurality of entity names pass entity format verification, the entity logical relationship is matched with a plurality of preset logical chains contained in a preset knowledge graph to obtain a matching result; in a case where the matching result indicates that there is a preset logical chain in the preset knowledge graph that has a relationship matching degree greater than a preset threshold with the entity logical chain, the entity logical chain is corrected according to the matched preset logical chain, and a correction result is determined as the text recognition result.

[0071] Through named entity recognition technology, a plurality of entity names (such as names, place names, organization names, etc.) are extracted from the first recognition result. The logical relationship between the extracted entity names (such as names, organization names, amounts, etc.) and the logical relationship between the entities (such as the "opening bank-bank card number" corresponding relationship) is compared with the preset logical chain in the knowledge graph, and the matching degree is calculated. If the matching degree is greater than the preset threshold, it is considered that the matching preset logical chain is found. If the matching preset logical chain is found, the extracted entity logical relationship chain is corrected according to the logical chain. The corrected logical relationship chain is consistent with the preset logical chain in the knowledge graph, ensuring the accuracy of the semantics. The corrected entity logical relationship chain is taken as the final text recognition result.

[0072] Specifically: the BERT model is used for named entity recognition (NER) and relationship extraction of the text recognized by OCR. After pre-training on a large amount of corpus, the model has strong semantic understanding ability and can accurately identify entities and logical relationships between them from complex text.

[0073] To further improve the accuracy of the extraction results, a post-processing scheme based on the reasoning rules of the knowledge graph. The system contains a triple verification mechanism: first, field type verification, through regular expression to verify the number, amount, date and other formats whether to meet the pre-defined rules; second, logical relationship verification, such as the value matching between the bill amount and the capital amount; third, context association verification, using the semantic relationship between entities in the knowledge graph (such as the corresponding relationship of "opening bank-bank card number") to verify and correct the extraction results. In this way, not only can some recognition errors be automatically corrected, but also the fault tolerance of the system to abnormal data can be enhanced by using domain knowledge.

[0074] To optimize the forward propagation of the correction effect of the knowledge graph, the following reward function is defined

[0075] R(s,a)=α*sim(a,s*)-β*err(s,s * )

[0076] Where s represents the original text output by the OCR model, a represents the knowledge graph or the corrected text, S * represents the expected correct text, sim(a,s * ) represents the similarity between the corrected result and the correct text, the higher the data, the closer to the correct, err(s,s * ) represents the error between the original OCR output and the correct text, defined as the normalized edit distance, and a and β are weight coefficients, used to balance the influence of similarity and error penalty.

[0077] Thus the loss function of the OCR model also increases a feedback signal brought by the reward function, and the reward function is converted into an additional loss term L RL by REINFORCE algorithm.

[0078]

[0079] The loss term together with the L adv , L gate_reg defined above and the traditional L CTC constitutes the loss function of the OCR model, so as to make the OCR model more inclined to generate correct text that meets the feedback of the knowledge graph, and finally achieve the purpose of automatic correction.

[0080] Exemplarily, a type of the entity corresponding to the entity name is determined, and a target entity type matching the type of the entity is matched from the entity standard library; the type of the entity is used to indicate an entity format corresponding to the entity; a format of the entity name is compared with an entity format indicated by the target entity type, and in a case where there is a difference between the format of the entity name and the entity format indicated by the target entity type, it is determined that the format corresponding to the entity name is modified according to the entity format indicated by the target entity type.

[0081] S202, identifying a printed text region and a multi-scale text region in the optimized image through a text detection model; and determining the multi-scale text region as an image range in which text to be recognized exists.

[0082] That is, the printed text region and the multi-scale text region are identified from the optimized image through the text detection model. A pre-trained text detection model (such as a target detection model based on deep learning, such as CTPN, EAST, DB, etc.) is used to scan the optimized image, locate the text region therein, and output the coordinate range of the printed text region and the multi-scale text region. The printed text region refers to a text region (such as a title, a main text, etc.) with fixed font and layout rules in the image. The multi-scale text region refers to a text region (such as text with mixed large and small fonts) with irregular font size and layout. The multi-scale text region is determined as an image range in which text to be recognized exists. Since the multi-scale text region contains more complex and more text to be recognized (such as small fonts and irregular layout), it is regarded as the main range of the text to be recognized. The printed text region has been clearly identified or does not need further processing, so the multi-scale text region is focused on. The text region is automatically located through the text detection model, reducing manual intervention. The multi-scale text region is focused on for processing, improving the recognition efficiency and accuracy.

[0083] Exemplarily, the recognition result of the text to be recognized is sent to a target object, and feedback information issued by the target object based on the recognition result is received, wherein the feedback information is used to represent the accuracy of the recognition result of the text to be recognized; in a case where the feedback information carries modification information, the modification information and the recognition result of the text to be recognized are recorded, and the text to be recognized is sent to an operation and maintenance object.

[0084] The recognition result (such as text content recognized by OCR) output by the text recognition model is sent to a target object through a network or an interface. The target object can be an artificial auditor, an automated system or other entities that need to verify the recognition result. The feedback information issued by the target object based on the recognition result is received, which is used to evaluate the accuracy of the recognition result.

[0085] Optionally, the target object verifies the recognition result to determine its accuracy. The feedback information can include:

[0086] Accuracy: For example, "the recognition accuracy is 90%" or "completely correct".

[0087] Modification information: If the recognition result is incorrect, the target object can provide the modified text content. According to the modification information in the feedback information, the recognition result is recorded and updated, and the problem is sent to the operation and maintenance object (such as a system administrator or a development team). The modification information provided by the target object is recorded together with the original recognition result for subsequent analysis or improvement of the model. If the feedback information carries modification information (i.e., the recognition result is incorrect), the original text to be recognized, the recognition result, and the modification information are sent to the operation and maintenance object. Through the feedback mechanism, the accuracy of the recognition result can be evaluated in real time, and the shortcomings of the model can be found.

[0088] The multi-scale text recognition method provided by the embodiments of the present application corrects image distortion and improves clarity through geometric correction and image enhancement technology, providing standardized input for subsequent recognition. By positioning the text area, the text content is extracted in combination with the character recognition model based on machine learning. The model is trained using multi-style samples (such as handwriting and printed matter), so that it has generalization ability, but may only recognize part of the text due to the influence of multi-scale. The preliminary recognition result is checked for language rules to correct incorrect characters or supplement missing parts, and finally a complete and logical text means is output, achieving the effect of accurately recognizing multi-scale text.

[0089] Figure 3 The structure diagram of the multi-scale text recognition device provided by the present application is shown in Figure 4 The multi-scale text recognition device 40 provided by the embodiments of the present application includes:

[0090] The acquisition module 401 is configured to acquire image information carrying a multi-scale text to be recognized, and perform optimization processing on the image information to obtain an optimized image. The optimization processing is used to correct distortion of the image information and / or adjust clarity of the image information.

[0091] The determination module 402 is configured to determine an image range in which text to be recognized exists in the optimized image, and identify the text to be recognized in the image range through a predetermined character recognition model to obtain a first recognition result. The character recognition model is obtained through machine learning of a plurality of sample data sets. Each data set in the plurality of sample data sets includes sample characters of different styles and entity characters corresponding to the sample characters. The first recognition result at least includes a recognition result of a part of the multi-scale text to be recognized.

[0092] The identification module 403 is configured to perform semantic identification on the first identification result, and determine a semantic identification result as the text identification result to be identified.

[0093] In a possible implementation, the acquisition module 401 is configured to adjust an image parameter of the image information according to a preset image parameter range, to obtain a preliminary optimized image, wherein the image parameter at least includes one of the following: image brightness, image contrast; perform noise reduction processing on the preliminary optimized image, to obtain a noise reduction image; detect a degree of inclination of the noise reduction image, and perform correction processing on the noise reduction image according to a detection result, to obtain a corrected image; and determine the corrected image as the optimized image corresponding to the image information.

[0094] In a possible implementation, the acquisition module 401 is configured to perform edge detection on the noise reduction image, to obtain an edge coordinate set of edge pixel points of the noise reduction image; calculate a discrete degree and a direction of edge pixel point distribution of the noise reduction image according to a coordinate mean corresponding to the edge coordinate set; the discrete degree and the direction of the edge pixel point distribution determine an inclination angle and a first inclination direction corresponding to the noise reduction image; and rotate the noise reduction image according to a second inclination direction by an angle identical to the inclination angle, wherein the second inclination direction is opposite to the first inclination direction.

[0095] In a possible implementation, the identification module 403 is configured to perform preliminary semantic identification on the first identification result, to obtain a plurality of entity names contained in the first identification result and an entity logical relationship chain between the plurality of entity names; in a case where the plurality of entity names pass entity format verification, match the entity logical relationship with a plurality of preset logical chains contained in a preset knowledge graph, to obtain a matching result; in a case where the matching result indicates that there is a preset logical chain in the preset knowledge graph, a matching degree of which with the entity logical chain is greater than a preset threshold, correct the entity logical chain according to the matched preset logical chain, and determine a correction result as the text identification result.

[0096] In a possible implementation, the identification module 403 is configured to determine an entity type corresponding to the entity name, and match a target entity type matching the entity type from the entity standard library; wherein the entity type is used to indicate an entity format corresponding to the entity; compare a format of the entity name with an entity format indicated by the target entity type, and in a case where there is a difference between the format of the entity name and the entity format indicated by the target entity type, determine to correct the format corresponding to the entity name according to the entity format indicated by the target entity type.

[0097] In a possible implementation, the determining module 402 is configured to identify a printed text region and a multi-scale text region in the optimized image by using a text detection model; and determine the multi-scale text region as an image range in which text to be recognized exists.

[0098] In a possible implementation, the identifying module 403 is configured to send the recognition result of the text to be recognized to a target object, receive feedback information issued by the target object based on the recognition result, where the feedback information is used to represent the accuracy of the recognition result of the text to be recognized; and in a case where the feedback information carries modification information, record the modification information and the recognition result of the text to be recognized, and send the text to be recognized to an operation and maintenance object.

[0099] The multi-scale text recognition apparatus provided in this embodiment can execute the method provided in the method embodiments, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0100] Figure 4 A structural schematic diagram of the multi-scale text recognition apparatus provided in this embodiment is shown in FIG. 5. ​ As shown in FIG. 5, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected through a bus.

[0101] In the specific implementation process, the at least one processor 501 executes the computer execution instructions stored in the memory 502, so that the at least one processor 501 executes the method described above.

[0102] The specific implementation process of the processor 501 can refer to the method embodiments described above, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0103] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as execution completed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0104] The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk memory.

[0105] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0106] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the above method.

[0107] The present application also provides a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the above method is implemented.

[0108] The above readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0109] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0110] The division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0111] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0112] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0113] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the present application that essentially contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0114] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The aforementioned program can be stored in a computer readable storage medium. The program executes the steps including the above-mentioned method embodiments when executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, and various program code storage media.

[0115] It should be understood that many of the materials and devices exemplified in this disclosure are articles of manufacture (i.e., articles of manufacture) according to this disclosure. The articles of manufacture can be manufactured as such or can be manufactured by combining the materials and devices exemplified in this disclosure. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It should be understood that, in some embodiments, equivalents to the specific electrode structures and / or methods described herein can be employed without departing from the scope of the application. Accordingly, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," "characterized by," "characterized into," and variations thereof herein, is meant to encompass the items listed thereafter, and equivalents thereof as well as additional items. Although the foregoing application has been described in some detail by way of illustration and example, it is not to be limited thereby, but rather, only by the scope of the appended claims.

Claims

1. A method for recognizing multi-scale text, characterized in that, The method comprises: obtaining image information carrying multi-scale text to be recognized, and performing optimization processing on the image information to obtain an optimized image, wherein the optimization processing is used to correct distortion of the image information and / or adjust the definition of the image information; determining an image range in which the text to be recognized exists in the optimized image, and identifying the text to be recognized in the image range through a preset character recognition model to obtain a first recognition result, wherein the character recognition model is obtained through machine learning of a plurality of sample data, wherein each set of data in the plurality of sample data comprises sample characters of different styles and entity characters corresponding to the sample characters, and the first recognition result at least comprises a recognition result of a middle part of the multi-scale text to be recognized; performing semantic recognition on the first recognition result, and determining a recognition result of the semantic recognition as a recognition result of the text to be recognized.

2. The method of claim 1, wherein, obtaining image information carrying multi-scale text to be recognized, and performing optimization processing on the image information to obtain an optimized image, comprising: adjusting image parameters of the image information according to a preset image parameter range to obtain a preliminary optimized image, wherein the image parameters at least comprise one of the following: image brightness, image contrast; performing noise reduction processing on the preliminary optimized image to obtain a noise reduction image; detecting the degree of inclination of the noise reduction image, and performing correction processing on the noise reduction image according to the detection result to obtain a corrected image; determining the corrected image as the optimized image corresponding to the image information.

3. The method of claim 2, wherein, detecting the degree of inclination of the noise reduction image, and performing correction processing on the noise reduction image according to the detection result to obtain a corrected image, comprising: performing edge detection on the noise reduction image to obtain an edge coordinate set of edge pixel points of the noise reduction image; calculating the dispersion degree and direction of the edge pixel point distribution of the noise reduction image according to the coordinate mean value corresponding to the edge coordinate set; the dispersion degree and direction of the edge pixel point distribution determine the inclination angle and the first inclination direction corresponding to the noise reduction image; rotating the noise reduction image according to the second inclination direction by an angle consistent with the inclination angle value, wherein the second inclination direction is opposite to the first inclination direction.

4. The method of claim 1, wherein, performing semantic recognition on the first recognition result, and determining a recognition result of the semantic recognition as a recognition result of the text to be recognized, comprising: performing preliminary semantic recognition on the first recognition result to obtain a plurality of entity names contained in the first recognition result and an entity logical relationship chain between the plurality of entity names; in a case where the plurality of entity names pass entity format verification, matching the entity logical relationship with a plurality of preset logical chains contained in a preset knowledge graph to obtain a matching result; in a case where the matching result indicates that there is a preset logical chain in the preset knowledge graph, the matching degree of which with the entity logical chain is greater than a preset threshold, correcting the entity logical chain according to the matched preset logical chain, and determining a correction result as the text recognition result.

5. The method of claim 4, wherein, after obtaining the plurality of entity names contained in the first recognition result and the entity logical relationship chain between the plurality of entity names, the method further comprises: determine an entity type corresponding to the entity name, and match a target entity type from the entity standard library that matches the entity type; the entity type is used to indicate an entity format corresponding to the entity; compare the format of the entity name with the entity format indicated by the target entity type, and in the case of a difference between the format of the entity name and the entity format indicated by the target entity type, determine to correct the format corresponding to the entity name according to the entity format indicated by the target entity type.

6. The method of claim 1, wherein, determine an image range in which the to-be-recognized text exists in the optimized image, including: identify a printed text region and a multi-scale text region in the optimized image through a text detection model; determine the multi-scale text region as the image range in which the to-be-recognized text exists.

7. The method according to any one of claims 1 to 6, characterized in that, After performing semantic recognition on the first recognition result and determining the recognition result of the semantic recognition as the recognition result of the to-be-recognized text, the method further includes: sending the recognition result of the to-be-recognized text to a target object and receiving feedback information issued by the target object based on the recognition result, wherein the feedback information is used to represent the correctness of the recognition result of the to-be-recognized text; in the case that the feedback information carries modification information, recording the modification information and the recognition result of the to-be-recognized text, and sending the to-be-recognized text to an operation and maintenance object.

8. A multi-scale text recognition device, characterized in that, including: an acquisition module, configured to acquire image information carrying a to-be-recognized multi-scale text, and perform optimization processing on the image information to obtain an optimized image, wherein the optimization processing is used to correct distortion of the image information and / or adjust the clarity of the image information; a determination module, configured to determine an image range in which to-be-recognized text exists in the optimized image, and identify the to-be-recognized text in the image range through a preset character recognition model to obtain a first recognition result, wherein the character recognition model is obtained through machine learning of a plurality of sample data, and each group of data in the plurality of sample data includes a sample character of different styles and an entity character corresponding to the sample character, and the first recognition result at least includes a recognition result of a part of the to-be-recognized multi-scale text; an identification module, configured to perform semantic recognition on the first recognition result, and determine a recognition result of the semantic recognition as a recognition result of the to-be-recognized text.

9. A device for recognizing multi-scale text, characterized by including: a memory and a processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the processor executes the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of any one of claims 1-7.