Character recognition method, character recognition device, electronic equipment and storage medium

By acquiring image data for text extraction, feature extraction, and missing detection, the problem of identifying incomplete or blurred text in images is solved, achieving higher recognition accuracy and comprehensiveness.

CN120808354APending Publication Date: 2025-10-17PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510860108.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively recognizing incomplete or blurred text in images, resulting in insufficient recognition accuracy.

Method used

By acquiring image data, performing text extraction and feature extraction, detecting missing text information, and completing it based on the features, it finally performs semantic recognition.

Benefits of technology

It improves the accuracy and comprehensiveness of text recognition, can effectively repair incomplete text, and is suitable for text recognition in a variety of scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808354A_ABST
    Figure CN120808354A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a character recognition method, a character recognition device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is applied to the fields of finance and medical treatment. The method comprises the following steps: acquiring current image data; performing character extraction on the current image data to obtain preliminary character information; performing feature extraction based on the preliminary character information to obtain preliminary character features; performing character missing detection based on the preliminary character information to obtain character missing information; complementing the character missing information according to the preliminary character features to obtain complemented character information; and performing semantic recognition according to the preliminary font features and the complemented character information, thereby improving the accuracy and comprehensiveness of character recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and is applied to the fields of financial technology and medical treatment, and in particular relates to a character recognition method, a character recognition device, an electronic device and a storage medium. BACKGROUND

[0002] Character recognition technology is a technology for automatically recognizing characters, and can be used to recognize characters in an image. Character recognition technology can be applied to many fields. For example, in the field of financial technology, it can be used to recognize the content of an insurance policy, in the field of medical treatment, it can be used to recognize case information, in other application scenarios, it can be used to recognize the character content of cultural relics images to understand the history of human development, and it can also be used to recognize various drug instructions, food packaging bags, and appliance packaging bags. At present, the difficulty of image-based character recognition technology lies in that the characters in the image are incomplete or blurred, and it is difficult to restore the complete content, resulting in great challenges in recognition. Therefore, how to improve the accuracy of character recognition has become a problem to be solved. SUMMARY

[0003] The main purpose of the embodiments of the present application is to provide a character recognition method, a character recognition device, an electronic device and a storage medium, which aims to improve the accuracy of character recognition.

[0004] To achieve the above-mentioned purpose, a first aspect of the embodiments of the present application provides a character recognition method, which comprises:

[0005] obtaining current image data;

[0006] extracting characters from the current image data to obtain preliminary character information;

[0007] extracting features based on the preliminary character information to obtain preliminary character features;

[0008] detecting character missing based on the preliminary character information to obtain character missing information;

[0009] completing the character missing information according to the preliminary character features to obtain completed character information;

[0010] performing semantic recognition according to the preliminary character features and the completed character information.

[0011] In some embodiments, the semantic recognition according to the preliminary character features and the completed character information comprises:

[0012] performing syntax recognition according to the completed character information to obtain a syntax structure;

[0013] performing sentence pattern recognition according to the completed character information to obtain a sentence pattern and a rhetorical device.

[0014] performing style recognition according to the preliminary character feature, to obtain a language style feature;

[0015] performing translation according to the syntax structure, the sentence pattern, the rhetoric method, and the language style feature.

[0016] In some embodiments, the feature extraction based on the preliminary text information obtains a preliminary text feature, including:

[0017] performing stroke indentation recognition based on the preliminary text information, to obtain a stroke indentation feature;

[0018] performing text carrier element recognition based on the preliminary text information, to obtain a text carrier element feature;

[0019] performing character extraction based on the preliminary text information, to obtain a preliminary character feature;

[0020] performing feature fusion according to the stroke indentation feature, the text carrier element feature, and the preliminary character feature, to obtain the preliminary text feature.

[0021] In some embodiments, the text missing detection based on the preliminary text information obtains text missing information, including:

[0022] performing abnormal interruption detection based on the preliminary text information, to obtain abnormal interruption information;

[0023] performing text structure detection based on the preliminary text information, to obtain a partial radical missing information;

[0024] performing character detection based on the preliminary text information, to obtain a character missing information;

[0025] performing information splicing according to the abnormal interruption information, the partial radical missing information, and the character missing information, to obtain the text missing information.

[0026] In some embodiments, the text extraction on the current image data obtains preliminary text information, including:

[0027] performing text detection on the current image data, to obtain a preliminary text region;

[0028] performing text positioning on the preliminary text region, to obtain text position information;

[0029] performing text extraction on the preliminary text region based on the text position information, to obtain a text sequence;

[0030] performing semantic analysis according to the text sequence, to obtain the preliminary text information.

[0031] In some embodiments, the missing information is completed according to the preliminary character feature to obtain completed character information, including:

[0032] The missing area is located according to the preliminary character feature to obtain a missing position;

[0033] Context analysis is performed according to the missing position and the preliminary character feature to obtain keyword information;

[0034] Missing character prediction is performed based on the keyword information to obtain missing content;

[0035] Information completion is performed according to the missing content to obtain the completed character information.

[0036] In some embodiments, the current image data is obtained, including:

[0037] Original image data is obtained; wherein the original image data is image data;

[0038] The original image data is denoised to obtain denoised image data;

[0039] The denoised image data is contrast-adjusted to obtain the preliminary image data;

[0040] The preliminary image data is binarized to obtain the current image data.

[0041] To achieve the above object, a second aspect of the embodiment of the present application proposes a character recognition device, the device comprising:

[0042] An image data acquisition module acquires current image data;

[0043] A character recognition module is configured to extract characters from the current image data to obtain preliminary character information;

[0044] A feature extraction module is configured to extract features based on the preliminary character information to obtain preliminary character features;

[0045] A missing detection module is configured to detect missing characters based on the preliminary character information to obtain character missing information;

[0046] A missing completion module is configured to complete the character missing information according to the preliminary character features to obtain completed character information;

[0047] A semantic recognition module is configured to perform semantic recognition according to the preliminary character features and the completed character information.

[0048] To achieve the above object, a third aspect of the embodiment of the present application provides an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.

[0049] To achieve the above object, a fourth aspect of the embodiment of the present application provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.

[0050] The character recognition method, character recognition device, electronic device and storage medium provided by the embodiment of the present application can improve the accuracy of character recognition by obtaining current image data, extracting characters from the current image data, extracting features based on the preliminary character information, detecting character missing based on the preliminary character information, completing the character missing information according to the preliminary character features, and performing semantic recognition according to the preliminary character features and the completed character information. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 is a flowchart of the character recognition method provided by the embodiment of the present application;

[0052] Figure 2 is a flowchart of step 102 in Figure 1

[0053] Figure 3 is a flowchart of step 103 in Figure 1

[0054] Figure 4 is a flowchart of step 104 in Figure 1

[0055] Figure 5 is a flowchart of step 105 in Figure 1

[0056] Figure 6 is a flowchart of step 106 in Figure 1

[0057] Figure 7 is a structural schematic diagram of the character recognition device provided by the embodiment of the present application;

[0058] Figure 8 is a hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0059] ​​​​​In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.

[0060] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the device or the sequence in the flowchart. The terms "first", "second", and the like in the specification and claims and the above-described drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of the present application only and is not intended to limit the present application.

[0062] First, the terms involved in the present application are analyzed:

[0063] Artificial intelligence (AI): is a new technical science that studies, develops and applies systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, artificial intelligence aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The field of research includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0064] Natural language processing (NLP): NLP uses computers to process, understand and use human language (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and is an interdisciplinary subject of computer science and linguistics, also commonly known as computational linguistics. Natural language processing includes syntax analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining, etc. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and language computing related linguistic research.

[0065] Optical Character Recognition (OCR): OCR can be used for electronic devices (such as scanners or digital cameras) to examine characters printed on paper, determine their shape by detecting dark and light patterns, and character recognition methods translate the shape into computer text. For printed characters, use optical methods to convert the text in paper documents into black and white bitmap image files, and convert the text in the image into text format through recognition software for further editing and processing by word processing software.

[0066] Text recognition technology can be applied to many fields. For example, in the financial technology field, it can be used to identify insurance content. In the medical and health field, it can be used to identify case information. In other application scenarios, it can also be used to identify the text content of cultural relics images to understand the history of human development, and to identify various drug instructions, food packaging bags, and appliance packaging bags. Currently, the difficulty of image-based text recognition technology lies in that the text in the image is incomplete or blurred, making it difficult to restore the complete content, resulting in great challenges in recognition. In the field of archaeology, ancient characters are often engraved on stone, pottery, metal, or written on papyrus, bamboo tubes and other perishable materials. After thousands of years, they may be incomplete or blurred. Many text materials are unearthed in the form of fragments, making it difficult to restore the complete content. At the same time, due to natural erosion, human destruction or improper preservation, text information is lost, so there are great challenges in recognition.

[0067] Therefore, the embodiments of the present application provide a text recognition method, a text recognition device, an electronic device, and a storage medium, which aim to improve the efficiency and accuracy of text recognition.

[0068] The text recognition method, the text recognition device, the electronic device, and the storage medium provided by the embodiments of the present application are specifically described by the following embodiments. First, the text recognition method in the embodiments of the present application is described.

[0069] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0070] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0071] The character recognition method provided by the embodiments of the present application relates to the technical field of artificial intelligence. The character recognition method provided by the embodiments of the present application can be applied to a terminal, can be applied to a server side, and can also be software running in the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as a stand-alone physical server, can be configured as a server cluster or a distributed system composed of multiple physical servers, can also be configured as a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform; and the software can be an application that implements the character recognition method, but is not limited to the above forms.

[0072] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0073] It should be noted that in each specific embodiment of the present application, when it is necessary to process relevant data related to the identity or characteristics of the user according to the user's personal information, user attribute information, user historical data, etc., the user's permission or consent will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0074] Figure 1 is an optional flowchart of the character recognition method provided by the embodiments of the present application, Figure 1 The method in can include but is not limited to steps 101 to 106.

[0075] Step 101, obtaining current image data;

[0076] Step 102, performing text extraction on the current image data to obtain preliminary text information;

[0077] Step 103, performing feature extraction based on the preliminary text information to obtain preliminary text features;

[0078] Step 104, performing text missing detection based on the preliminary text information to obtain text missing information;

[0079] Step 105, completing the text missing information according to the preliminary text features to obtain completed text information;

[0080] Step 106, performing semantic recognition according to the preliminary text features and the completed text information.

[0081] The steps 101 to 106 shown in the embodiments of the present application, by obtaining current image data, and performing text extraction on the current image data, then performing feature extraction based on the preliminary text information, and performing text missing detection based on the preliminary text information, and then completing the text missing information according to the preliminary text features, so as to perform semantic recognition according to the preliminary text features and the completed text information, thereby improving the accuracy and comprehensiveness of text recognition.

[0082] In step 101 of some embodiments, it can include but is not limited to including: pre-processing the image data, which can specifically include:

[0083] Obtaining original image data; wherein the original image data is image data;

[0084] Performing denoising processing on the original image data to obtain denoised image data;

[0085] Performing contrast adjustment on the denoised image data to obtain preliminary image data;

[0086] Performing binaryzation processing on the preliminary image data to obtain current image data.

[0087] In some embodiments, the original image data can be extracted from a preset database, or the original image data can be obtained by photographing or scanning related items based on relevant scanning equipment. In a medical and health scenario, a case image can be obtained by photographing a case, and the case image can be used as the original image data. A drug image can also be obtained by photographing a drug description or drug packaging, and the drug image can be used as the original image data. In a financial scenario, an insurance policy image can be obtained by photographing an insurance policy, and the policy image can be used as the original image data. Alternatively, the image data of the electronic insurance policy file can be directly obtained. In other application scenarios, a polarized light imager can be used to scan the front and back of an oracle bone and capture the surface micro-trace structure to obtain the original image data. A hyperspectral imager can also be used for scanning, and a hyperspectral imager can penetrate surface contaminants for scanning and identification. In another application scenario, the original image data can also be data that has been pre-acquired and stored in a preset database, so that the original image data can be directly extracted from the preset database. In addition, the original image data uploaded by the user can also be obtained. The embodiment of the present application does not limit the method for obtaining the original image data.

[0088] In one embodiment, denoising raw image data may include frequency-domain filtering, thereby enabling text recognition despite interference from moss and tourist scratches on the stone tablet surface. For example, the raw image data may contain moss covering the characters "大唐" (Tang Dynasty), three unclear characters in the case diagnosis "pulmonary nodules," and unclear gender information in the insurance policy. Frequency-domain filtering significantly improves the clarity of the text strokes in the denoised image data. Furthermore, adaptive median filtering can be used to address rust on bronze artifacts. Furthermore, Mask R-CNN segmentation technology can be used to distinguish text from biological stains, thereby improving text clarity.

[0089] In one embodiment, low-contrast text restoration is achieved by performing contrast adjustment on denoised image data. In one application scenario, where Tang Dynasty cinnabar ink has faded, resulting in insufficient contrast, a spectral ratio method (for weathered ink) can be used to calculate the ratio of near-infrared (780nm) to visible light (550nm) to highlight the remaining cinnabar areas, thereby improving text clarity.

[0090] In one embodiment, text can be separated from the background by binarizing the preliminary image data. For example, a local adaptive threshold can be used to account for uneven lighting, or a multispectral fusion binarization principle can be used to highlight cinnabar ink marks using a near-infrared and visible light ratio image. In one application scenario, when uneven lighting on a stone tablet causes the global threshold to fail, binarization can be used to separate the text from the background and highlight the text content.

[0091] Referring to Figure 2 In step 102 of some embodiments can include but not limited to including:

[0092] Step 201, text detection is performed on the current image data to obtain a preliminary text region;

[0093] Step 202, text positioning is performed on the preliminary text region to obtain text position information;

[0094] Step 203, text extraction is performed on the preliminary text region based on the text position information to obtain a text sequence;

[0095] Step 204, semantic analysis is performed according to the text sequence to obtain preliminary text information.

[0096] In step 201 of some embodiments, the text region can be positioned based on a target monitoring model, for example, using a YOLO (You Only Look Once), Faster R-CNN, etc. model, wherein YOLO is a real-time object detection algorithm that can identify the category of an object and mark the position of the object in a picture; Faster R-CNN is a target detection framework based on deep learning. The implementation of the text region positioning of the embodiments of the present application is not limited. In an application scenario, the text region can be framed in a stone tablet image.

[0097] In step 202 of some embodiments, YOLO, Mask R-CNN, Faster R-CNN, etc. model can be used for text positioning, for example, YOLO can be used for text positioning of murals, manuscripts, cases, insurance policies, etc., Mask R-CNN can be used for segmentation of adhered text, and can be used for calligraphy, handwritten text positioning, etc. to obtain text position information.

[0098] In step 203 of some embodiments, OCR technology can be used for text extraction, character recognition and extraction are performed on the preliminary text region based on the text position information to obtain a text sequence, thereby facilitating step 204.

[0099] In step 204 of some embodiments, the text sequence can be semantically analyzed based on the knowledge graph principle to extract text information.

[0100] Through steps 201 to 204 of the above embodiments, text detection can be performed on the current image data, a text sequence is extracted based on the obtained text position information, and voice analysis is performed based on the text sequence to obtain preliminary text information.

[0101] In step 103 of some embodiments, feature extraction can be performed using a CNN model such as ResNet or EfficientNet. Please refer to Figure 3 In step 103 of some embodiments, feature extraction can include but is not limited to including:

[0102] Step 301: Based on the preliminary text information, stroke indentation recognition is performed to obtain stroke indentation features;

[0103] Step 302: Based on the preliminary text information, text carrier element recognition is performed to obtain text carrier element features;

[0104] Step 303: Based on the preliminary text information, character extraction is performed to obtain preliminary character features;

[0105] Step 304: Based on the stroke indentation features, text carrier element features, and preliminary character features, feature fusion is performed to obtain preliminary text features.

[0106] In step 301 of some embodiments, the preliminary text information is annotated with stroke features of the text. Based on the stroke features, the depth differences of different strokes are analyzed to obtain stroke indentation features, which include depth features of the stroke indentation.

[0107] In step 302 of some embodiments, the text carrier element includes chemical composition elements and physical property elements, etc. For example, high concentration of sulfur was detected in the Dunhuang Xuanquan site Hanjian, which confirmed that the "burning to remove insects" technology was used at that time.

[0108] In step 303 of some embodiments, based on the preliminary text information, character extraction is performed to obtain preliminary character features. Through the character features, the writing form, appearance shape, etc. of the text, as well as the specific writing method, stroke order, and structure layout, etc. of the text can be obtained. Each character has unique character features.

[0109] In step 304 of some embodiments, by fusing the stroke indentation features, text carrier element features, and preliminary character features, more comprehensive text features can be obtained, which facilitates subsequent semantic recognition.

[0110] Please refer to Figure 4 In step 104 of some embodiments, feature extraction can include but is not limited to including:

[0111] Step 401: Based on the preliminary text information, abnormal interruption detection is performed to obtain abnormal interruption information;

[0112] Step 402: Based on the preliminary text information, text structure detection is performed to obtain partial radical missing information;

[0113] Step 403: Based on the preliminary text information, character detection is performed to obtain character missing information;

[0114] Step 404: information is spliced ​​based on the abnormal interruption information, the radical incomplete information and the character missing information to obtain the text missing information.

[0115] In step 401 of some embodiments, the text stroke continuity is detected on the preliminary text information to determine whether there is an abnormal break point, such as whether there is a sudden change in direction. Specifically, it includes: identifying the stroke break caused by weathering, wear or human damage by calculating the gradient direction change of the stroke edge, for example, the Sobel operator, Canny edge detection and direction consistency can be used for analysis; in other embodiments, CNN can be used to predict the stroke direction instead of the Sobel operator. In the embodiment of the present application, the abnormal break detection includes: sudden change in direction and abnormal gradient amplitude. The detection principle of sudden change in direction is mainly: when the stroke is normal, the direction change is smooth; if a sudden change occurs (for example, the sudden change is greater than 30%), it is judged that a break point may occur; the detection principle of abnormal gradient amplitude is mainly: if the gradient amplitude suddenly drops (for example, greater than 50%), it is judged that there may be a stroke break. In one application scenario, the severely weathered oracle bone inscription "受" (receive) is used as an example. Some of the strokes of this character are broken. During abnormal interruption detection, the direction mutation point is found to be: the end of the right vertical stroke suddenly changes by 42 degrees, and the gradient amplitude drops by 68% at the break. This abnormal interruption information facilitates step 105 to complete the strokes at the interruption point. In another application scenario, the left part of the bamboo slip character "盗" (robber) is missing due to insect damage. The abnormal interruption information is: direction mutation point: the end of the left "次" part suddenly changes by 35 degrees, and the gradient amplitude drops by 72% at the missing part. Therefore, the suggestion for step 105 to complete the "次" part structure based on the context is: complete the "次" part structure based on the context.

[0116] By performing directional mutation analysis and gradient amplitude change in step 401 , the break position in the cultural relic text can be effectively identified, thereby facilitating the completion processing in step 105 .

[0117] In step 402 of some embodiments, by detecting the radicals of the character structure, incomplete information of the radicals of the character can be obtained, for example, it is detected that the character "明" is missing the "日" radical; the character "清" is missing the "氵" structure.

[0118] In step 403 of some embodiments, through character missing detection, it is possible to accurately detect the situation where the entire character is missing, for example, the missing text caused by the string holes of bamboo slips, the disappearance of characters caused by the weathering of stone tablets, etc., and accurate detection can be achieved by combining OCR model prediction confidence analysis and contextual semantic reasoning, where the confidence threshold can be set to: if the confidence is ≥0.9, it indicates that the character is complete; if the confidence is ≥0.7 and <0.9, it indicates that the character may be partially missing; if the confidence is <0.7, it indicates that the character is missing.

[0119] In step 404 of some embodiments, by splicing the abnormal interruption information, the radical incomplete information and the character missing information, the text missing information obtained integrates the abnormal interruption information, the radical incomplete information and the character missing information, and can more accurately determine the missing information of the text, thereby facilitating the completion processing in step 105.

[0120] See also Figure 5 In some embodiments, step 105 may include, but is not limited to:

[0121] Step 501: Locate the missing area of ​​the missing text information based on preliminary text features to obtain the missing position;

[0122] Step 502: Perform context analysis based on the missing position and preliminary text features to obtain keyword information;

[0123] Step 503: predict missing text based on keyword information to obtain missing content;

[0124] Step 504: Complete the information according to the missing content to obtain completed text information.

[0125] In step 501 of some embodiments, models such as YOLO, Mask R-CNN, and Faster R-CNN can be used to locate the missing region. For example, YOLO can be used to locate murals and manuscripts, while Mask R-CNN can be used to segment overlapping text and locate text in calligraphy, handwriting, medical records, insurance policies, and other text, thereby determining the missing region. In one application scenario, the missing regions of "__卜,_贞:雨?王曰:其雨" include the first, second, and fifth positions, with the character "卜" being the third position and the comma being the fourth position.

[0126] In steps 502 to 504 of some embodiments, models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers) can be used for contextual reasoning, so as to perform contextual analysis based on the missing position and preliminary text features to obtain keyword information, and perform missing text prediction based on the keyword information to obtain missing content. For example, the missing content of "__卜,_贞:雨?王曰:其雨" in the above embodiment includes: "桂" at the first position, "卯" at the second position, and "争" at the third position, so after completion based on step 504, the result is: "桂毛卜,争贞:雨?王曰:其雨", which facilitates further semantic recognition in the subsequent step 106. In another application scenario, the missing content of "Pinellia ___Licorice_______" includes: "three liang" at the third, fourth, and fifth positions, and "two liang, five slices of ginger" at the eighth, ninth, tenth, eleventh, twelfth, thirteenth, and fourteenth positions. Therefore, after completion based on step 504, the result is: "three liang of Pinellia, two liang of Licorice, five slices of ginger", which facilitates further semantic recognition in the subsequent step 106 and obtains "a prescription for treating cough, Pinellia dries dampness and resolves phlegm, and Licorice harmonizes the medicines."

[0127] Through steps 501 to 504 , the missing text content can be completed and repaired, which facilitates semantic recognition in step 106 .

[0128] In step 106 of some embodiments, semantic recognition of text can be performed based on models such as BERT and GPT to determine semantic information. For example, in a financial scenario, the content of an insurance policy can be identified and interpreted to obtain information such as the subject matter of insurance, insurance liability, and exemption clauses; in a medical and health scenario, drug information can be identified, and the use and precautions of the drug can be interpreted. Case information can also be identified and the case can be interpreted.

[0129] Step 106 also allows the translation of relevant text information. Figure 6 In some embodiments, step 106 may include, but is not limited to:

[0130] Step 601, performing grammar recognition based on the completed text information to obtain a grammatical structure;

[0131] Step 602: Sentence pattern recognition is performed based on the completed text information to obtain sentence patterns and rhetoric;

[0132] Step 603, performing style recognition based on the preliminary glyph features to obtain language style features;

[0133] Step 604, according to the syntax structure, sentence pattern, rhetoric and language style features for translation.

[0134] In some embodiments of step 601, according to the completion of the text information for syntax recognition, the resulting syntax structure includes the basic syntax structure of subject-predicate-object. In some embodiments of step 602, according to the completion of the text information for sentence pattern recognition, the sentence pattern and rhetoric can be obtained, wherein the sentence pattern includes: through the attributive modification of the center, the adverbial modification of the predicate, etc., for example, in "fifty acres of good land", "good" (adjective) modifies "land" (center). In other embodiments, there are also special sentence patterns, such as judgment sentences, inverted sentences, passive sentences, etc. For example, "this imperial edict" is a judgment sentence, which can be translated as "this is the emperor's edict"; "the city was destroyed" is a passive sentence, which can be translated as "the city was destroyed". Rhetoric can include: metaphor, metaphor, metaphor, exaggeration, contrast, etc. For example, "hands like soft silk, skin like condensed cream" is a metaphor, "the gentleman's virtue is the wind, and the small man's virtue is the grass" is a metaphor, "Xiang Zhuang dances sword, and intends to be the king" is a metaphor, "flying water straight down three thousand feet" is an exaggeration, and "yin and yang are born, and difficult and easy are compared" is a contrast.

[0135] In some embodiments of step 603, the style recognition of the preliminary character shape features obtains language style features, which can include but are not limited to: the character features of the same type of cultural relics of the same period, the font style features (such as the thin strength / fat of strokes, etc.), the writing habit features (such as the way of starting and ending writing, etc.), the word style of the same doctor describing the condition, etc.

[0136] In some embodiments of step 604, the translation is implemented according to the syntax structure, sentence pattern, rhetoric and language style features. For example, in "thousands of miles", "thousands of miles" is a quantity word phrase, which needs to be supplemented to explain the degree of "walking", so the quantity word expression method needs to be preserved during translation, and it is translated as: walking thousands of miles. In other application scenarios, it is not limited to translating the text content itself, but also includes further interpretation of the text content, such as the above embodiment "half summer three two, licorice two two, ginger five pieces" is translated as: "treatment of cough prescription, half summer dryness and phlegm, licorice harmonizes various drugs".

[0137] In an application scenario, step 604 can be associated with historical data for interpretation, interpreting specific materials and associated historical time, and identifying specific characters in a relative time range. For the same text may appear a variety of interpretations, the embodiments of the application can provide corresponding reference literature and reference data based on each interpretation to ensure the logical rationality of data interpretation and avoid subjective interpretation of human interpreters.

[0138] Through steps 601 to 604, semantic recognition can be performed according to the preliminary character feature and the completed character information, and further translation can be implemented, so as to realize the interpretation of the grammar and semantics of the characters, understand the meaning of the ancient literature. In addition, the rhetorical devices such as metaphors and symbols in ancient characters can also be recognized, the social structure and economic activities reflected thereby can be interpreted, the development context of historical events can be restored, and the structure of medical prescriptions can also be interpreted, and further the compatibility relationship of medicinal materials can be obtained.

[0139] In other application scenarios, the preference input of the user can also be received, and the grammar structure, sentence pattern, rhetorical device and language style feature preferred by the user are received for interpretation and translation, for example, a professional and complex case diagnosis is translated into a popular and easy-to-understand vernacular, so that non-medical personnel can understand it, and professional insurance knowledge is translated into a popular and easy-to-understand vernacular, so that ordinary users can understand it.

[0140] The character recognition method provided by the embodiment of the present application can improve the accuracy of character recognition by obtaining current image data, extracting characters from the current image data, extracting features based on preliminary character information, detecting missing characters based on preliminary character information, and completing missing character information according to preliminary character features.

[0141] The embodiment of the present application can combine image recognition technology (such as OCR) to recognize and complete blurred and incomplete characters, so as to realize repair, and can be applied to financial scenarios, medical and health scenarios and other scenarios, and can process various ancient languages and character systems, help readers quickly identify and classify character materials, and for incomplete character fragments, the possible situation can be inferred according to the context, and relevant translation can be provided, so that the reader can more intuitively understand the character content. The technical means of the embodiment of the present application can be used as a translation tool for ancient languages and modern languages, helping to interpret oracle bone inscriptions, cuneiform characters, Maya characters, etc. By interpreting the grammar and semantics of characters, the related meanings of ancient literature can be understood. In addition, the cultural background and context of the characters can also be interpreted by the embodiment of the present application, helping the reader to understand the actual use and meaning of the characters; the embodiment of the present application can also recognize metaphors, symbols and rhetorical devices in ancient characters, interpret the social structure, economic activities, historical development context and other information reflected thereby, compare different cultural character systems, and reveal the rules of cultural exchange and communication.

[0142] In addition, the embodiments of the present application can also automatically classify and organize a large number of documents according to characteristics such as content, language, and period. By extracting keywords from the documents, the embodiments of the present application can help readers quickly locate important information and suggest potential historical contexts. The embodiments of the present application can integrate knowledge from multiple disciplines such as linguistics, history, and archaeology to provide a more comprehensive research perspective and improve the quality of textual verification.

[0143] The embodiments of the present application have significant advantages in supporting textual research, which can improve research efficiency, expand research scope, and provide new ideas for solving complex problems. The embodiments of the present application can quickly process and interpret a large amount of textual data, significantly shorten the research time, and can simultaneously process multiple documents or text fragments to improve research efficiency and automatically complete tasks such as text recognition, classification, and translation, reducing manual workload. The embodiments of the present application can infer the possible content of incomplete text based on context, assist in repairing and interpreting documents, and combine multi-modal data such as images and text to improve the accuracy of text recognition. The embodiments of the present application can deeply interpret the grammar and semantics of the text, help understand the content of ancient documents, infer the actual use and cultural background of the text by interpreting the context association, identify metaphors, symbols, and rhetorical devices in the text, and reveal the deep meaning. By interpreting massive data, potential rules and patterns in the use of text are discovered, and abnormal phenomena (such as rare vocabulary and special grammar) in the text are identified, providing new clues for archaeological research and revealing cultural transmission and historical changes.

[0144] Please refer to Figure 7 The embodiments of the present application also provide a text recognition device that can implement the above-mentioned text recognition method. The device includes:

[0145] An image data acquisition module acquires current image data.

[0146] A text recognition module is configured to extract text from the current image data to obtain preliminary text information.

[0147] A feature extraction module is configured to extract features based on the preliminary text information to obtain preliminary text features.

[0148] A missing detection module is configured to detect text missing based on the preliminary text information to obtain text missing information.

[0149] A missing completion module is configured to complete the text missing information based on the preliminary text features to obtain completed text information.

[0150] A semantic recognition module is configured to recognize semantics based on the preliminary text features and the completed text information.

[0151] In some embodiments, the text recognition module can be specifically configured to implement:

[0152] performing text detection on the current image data to obtain a preliminary text region;

[0153] performing text positioning on the preliminary text region to obtain text position information;

[0154] performing text extraction on the preliminary text region based on the text position information to obtain a text sequence;

[0155] performing semantic analysis according to the text sequence to obtain preliminary text information.

[0156] Specifically, the text recognition module can be used to implement steps 201 to 204 described above, and will not be described here again.

[0157] In some embodiments, the feature extraction module, in particular, can be used to implement:

[0158] performing stroke indentation recognition based on the preliminary text information to obtain stroke indentation features;

[0159] performing text carrier element recognition based on the preliminary text information to obtain text carrier element features;

[0160] performing character shape extraction based on the preliminary text information to obtain preliminary character shape features;

[0161] performing feature fusion according to the stroke indentation features, the text carrier element features, and the preliminary character shape features to obtain preliminary text features.

[0162] Specifically, the feature extraction module can be used to implement steps 301 to 304 described above, and will not be described here again.

[0163] In some embodiments, the missing detection module, in particular, can be used to implement:

[0164] performing abnormal interruption detection based on the preliminary text information to obtain abnormal interruption information;

[0165] performing text structure detection based on the preliminary text information to obtain information about missing components of radicals and / or phonetic elements;

[0166] performing character detection based on the preliminary text information to obtain character missing information;

[0167] performing information splicing according to the abnormal interruption information, the information about missing components of radicals and / or phonetic elements, and the character missing information to obtain text missing information.

[0168] Specifically, the missing detection module can be used to implement steps 401 to 404 described above, and will not be described here again.

[0169] In some embodiments, the missing completion module, in particular, can be used to implement:

[0170] According to the preliminary character features, the missing region of the missing character information is located to obtain a missing position;

[0171] According to the missing position and the preliminary character features, context analysis is performed to obtain keyword information;

[0172] Based on the keyword information, missing character prediction is performed to obtain missing content;

[0173] According to the missing content, information completion is performed to obtain completed character information.

[0174] Specifically, the text missing completion module can be used to implement the steps 501 to 504, and details are not described herein.

[0175] In some embodiments, the semantic recognition module can be specifically used to implement:

[0176] According to the completed character information, syntax recognition is performed to obtain a syntax structure;

[0177] According to the completed character information, sentence pattern recognition is performed to obtain a sentence pattern and a rhetoric method;

[0178] According to the preliminary character features, style recognition is performed to obtain language style features;

[0179] According to the syntax structure, the sentence pattern, the rhetoric method and the language style features, translation is performed.

[0180] Specifically, the semantic recognition module can be used to implement the steps 601 to 604, and details are not described herein.

[0181] The specific implementation of the character recognition device is basically the same as the specific embodiments of the character recognition method described above, and details are not described herein.

[0182] Embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the character recognition method described above when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0183] Please refer to Figure 8 , Figure 8 The hardware structure of the electronic device of another embodiment is illustrated, which includes:

[0184] The processor 801 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0185] The memory 802 can be implemented by a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), etc. The memory 802 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 802 and are called and executed by the processor 801 to implement the character recognition method of the embodiments of the present application.

[0186] The input / output interface 803 is configured to implement information input and output.

[0187] The communication interface 804 is configured to implement the communication interaction between the device and other devices. The communication can be implemented by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0188] The bus 805 is configured to transmit information between various components (for example, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804) of the device.

[0189] The processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are connected to each other by the bus 805 to realize the communication connection between the device.

[0190] The embodiments of the present application further provide a storage medium, which is a computer readable storage medium. The storage medium stores a computer program. The computer program is executed by a processor to implement the character recognition method.

[0191] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0192] The character recognition method, character recognition device, electronic equipment and storage medium provided by the embodiments of the present application can improve the accuracy and comprehensiveness of character recognition by obtaining current image data, extracting characters from the current image data, extracting features based on preliminary character information, detecting character missing information based on the preliminary character information, and completing the character missing information according to the preliminary character features, and then performing semantic recognition according to the preliminary character features and the completed character information. The embodiments of the present application can be associated with historical data for interpretation, interpret specific materials and associated historical time, and identify specific characters in a relative time range. For multiple interpretations of the same text, the embodiments of the present application can provide corresponding reference literature and reference data based on each interpretation to ensure the logical rationality of data interpretation and avoid subjective interpretation by human interpreters.

[0193] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0194] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than shown in the figures, or combine certain steps or different steps.

[0195] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, that is, can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0196] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the function modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.

[0197] The terms "first", "second", "third", "fourth", and the like in the description and in the claims of this application, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed is interchangeable under appropriate circumstances such that the embodiments of the application described herein are, for example, capable of orderly or chronological mundane operation, reverse order operation, based on circuitry availability, based on stated preference or the like, and that "default" or other orderings are thus permissible. Further, the terms "comprise", "comprising", "include", "including", and the like, are specifically intended to be open-ended. That is, references to individual steps and the like do not suhstantially exclude the presence of two or more of a given step or its integral presence in the process, method, system, article, or apparatus having been made with a wider scope. The use of notation such as "first", "second", "third", etc. does not generally limit the areas, but is used to connect like elements or to distinguish one claim from another. These terms can be used interchangeably when appropriate. Terms concerning the relative position of elements can be used to describe preferred embodiments and variations thereof, but such terms are used only to describe relative position, and do not suhstantially limit or otherwise restrict the scope of the application to a specific spatial arrangement.

[0198] It should be understood that, in this application, "at least one" means one or more, "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0199] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the above-mentioned units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed objects can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0200] The units described above as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0201] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0202] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store programs.

[0203] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not intended to limit the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A method for character recognition, characterized in that: The method comprises: Get current image data; Performing text extraction on the current image data to obtain preliminary text information; Perform feature extraction based on the preliminary text information to obtain preliminary text features; Performing text missing detection based on the preliminary text information to obtain text missing information; Completing the missing text information according to the preliminary text features to obtain completed text information; Semantic recognition is performed based on the preliminary glyph features and the completed text information.

2. The method according to claim 1, characterized in that The performing semantic recognition according to the preliminary glyph features and the supplementary text information includes: Performing grammar recognition according to the completed text information to obtain a grammatical structure; Perform sentence pattern recognition based on the completed text information to obtain sentence patterns and rhetoric; Performing style recognition based on the preliminary glyph features to obtain language style features; The translation is performed according to the grammatical structure, the sentence pattern, the rhetoric and the language style features.

3. The method according to claim 1, characterized in that The extracting features based on the preliminary text information to obtain preliminary text features includes: Performing stroke and mark recognition based on the preliminary text information to obtain stroke and mark features; Performing text carrier element recognition based on the preliminary text information to obtain text carrier element features; Extracting glyphs based on the preliminary text information to obtain preliminary glyph features; The preliminary character features are obtained by performing feature fusion based on the stroke mark features, the character carrier element features and the preliminary character shape features.

4. The method according to claim 1, wherein The performing of character missing detection based on the preliminary character information to obtain character missing information includes: Performing abnormal interruption detection based on the preliminary text information to obtain abnormal interruption information; Performing text structure detection based on the preliminary text information to obtain incomplete information of radicals; Performing character detection based on the preliminary text information to obtain character missing information; Information splicing is performed according to the abnormal interruption information, the radical incomplete information and the character missing information to obtain the character missing information.

5. The method according to claim 1, wherein The extracting text from the current image data to obtain preliminary text information includes: Performing text detection on the current image data to obtain a preliminary text area; Performing text positioning on the preliminary text area to obtain text position information; Extracting text from the preliminary text area based on the text position information to obtain a text sequence; Semantic analysis is performed based on the text sequence to obtain the preliminary text information.

6. The method according to claim 1, characterized in that The step of completing the missing text information according to the preliminary text features to obtain completed text information includes: Locating the missing area of ​​the missing text information according to the preliminary text features to obtain the missing position; Performing context analysis based on the missing position and the preliminary text features to obtain keyword information; Predict missing words based on the keyword information to obtain missing content; Information is completed according to the missing content to obtain the completed text information.

7. The method according to any one of claims 1 to 6, characterized in that The obtaining of current image data includes: Acquire original image data; wherein the original image data is image data; Performing denoising processing on the original image data to obtain denoised image data; performing contrast adjustment on the denoised image data to obtain the preliminary image data; The preliminary image data is binarized to obtain the current image data.

8. A text recognition device, characterized in that: The device comprises: An image data acquisition module, which acquires current image data; wherein the current image data includes current image data; A text recognition module, configured to extract text from the current image data to obtain preliminary text information; A feature extraction module, configured to extract features based on the preliminary text information to obtain preliminary text features; a missing character detection module, configured to perform missing character detection based on the preliminary character information to obtain missing character information; A missing completion module, configured to complete the missing text information according to the preliminary text features to obtain completed text information; A semantic recognition module is used to perform semantic recognition based on the preliminary glyph features and the completed text information.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.