Text input method and device, electronic equipment and storage medium

By acquiring structured text images for character recognition and text frame segmentation, the error problem caused by uneven image placement during text entry is solved, achieving highly accurate and efficient text entry.

CN120706377APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510780507.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing text entry methods, if the input text image is not placed flat, scanning errors are likely to occur, resulting in incorrect text information content and low accuracy.

Method used

By acquiring structured text images, performing character recognition, extracting text box position information and content information, restoring the format based on the text box position information, segmenting the text box using the text box adhesion information, marking and merging the first and last texts, a structured merged text is formed and entered.

Benefits of technology

It improves the accuracy, efficiency and structuring of text entry, ensuring the integrity and logic of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706377A_ABST
    Figure CN120706377A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text input method and device, electronic equipment and a storage medium, belongs to the technical field of text processing, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: acquiring a structured text image; character recognition is conducted on the structured text image, textbox position information and text content information are obtained, and the textbox position information comprises textbox adhesion information; based on the textbox position information, performing plate-type reduction on the text content information to obtain a target structured text; based on the textbox adhesion information, performing textbox segmentation on the target structured text to obtain a standard text; performing head and tail text dotting on the standard text to obtain a text content combination mark point; on the basis of the text content merging mark points, text merging is conducted on the standard text, and a structured merged text is obtained; and carrying out text input on the structured combined text. According to the embodiment of the invention, the accuracy of text input can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing technology, and is applicable to the fields of financial technology and medicine, and in particular to a text entry method and device, an electronic device, and a storage medium. Background Art

[0002] Text entry is used to enter text information from an image into a designated system. For example, in the fintech sector, text entry of insurance policies can be used to store them long-term, allowing business personnel to access policy information at any time. In the healthcare sector, text entry of medical records can be used to store them long-term, allowing doctors to quickly access patient information.

[0003] Currently, the method of text entry usually involves scanning the input text image, generating text content, and finally entering the generated text content directly into a specific system. However, in this text entry method, if the input text image is not placed flat, scanning errors are likely to occur, resulting in content errors in the entered text information. Therefore, how to improve the accuracy of text entry has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a text entry method and device, an electronic device and a storage medium, aiming to improve the accuracy of text entry.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a text entry method, the method comprising:

[0006] Get structured text image;

[0007] Performing character recognition on the structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information;

[0008] Based on the text box position information, the text content information is restored to a plate format to obtain a target structured text;

[0009] Based on the text box adhesion information, the target structured text is segmented into text boxes to obtain a standard text;

[0010] Marking the beginning and end of the standard text to obtain text content merging mark points;

[0011] Based on the text content merging mark points, the standard text is merged to obtain a structured merged text;

[0012] Text entry is performed on the structured merged text.

[0013] In some embodiments, marking the beginning and end of the standard text to obtain text content merging mark points includes:

[0014] Performing similarity comparison on the standard texts to obtain a target merged text;

[0015] Performing text element position recognition on the target merged text to obtain the text element start position and the text element end position;

[0016] Marking the starting position of the text element to obtain a text merging starting mark point;

[0017] Marking the end position of the text element to obtain a text merging end mark point;

[0018] Mark point integration is performed on the text merging start mark point and the text merging end mark point to obtain the text content merging mark point.

[0019] In some embodiments, performing similarity comparison on the standard text to obtain the target merged text includes:

[0020] Comparing the content similarity of the standard text to obtain text content similarity;

[0021] Performing typesetting similarity comparison on the standard text to obtain text typesetting similarity;

[0022] Based on the text content similarity and the text layout similarity, the standard text is screened for merging areas to obtain the target merged text.

[0023] In some embodiments, performing content similarity comparison on the standard text to obtain text content similarity includes:

[0024] Performing semantic analysis on the standard text to obtain lexical semantics;

[0025] Calculating the similarity of the vocabulary semantics to obtain vocabulary similarity;

[0026] Based on the vocabulary similarity, context understanding is performed on the standard text to obtain the text content similarity.

[0027] In some embodiments, performing format restoration on the text content information based on the text box position information to obtain a target structured text includes:

[0028] Associating the text box position information with the text content information to obtain a text block;

[0029] Based on the text box position information, the text blocks are arranged and combined to obtain the target structured text.

[0030] In some embodiments, performing text box segmentation on the target structured text based on the text box adhesion information to obtain standard text includes:

[0031] Extracting features of the text box adhesion information to obtain adhesion text box features;

[0032] Based on the adhesion text box feature, the target structured text is segmented to obtain a segmented text box;

[0033] Based on the segmented text frames, the target structured text is subjected to association processing to obtain a standard text.

[0034] In some embodiments, performing character recognition on the structured text image to obtain text box position information and text content information includes:

[0035] Performing text region detection on the structured text image to obtain a text region;

[0036] Based on the text area, extracting text boxes from the structured text image to obtain text box position information;

[0037] Perform text extraction on the text area to obtain the text content information.

[0038] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a text entry device, comprising:

[0039] A data acquisition module, used to acquire structured text images;

[0040] a character recognition module, configured to perform character recognition on the structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information;

[0041] A format restoration module, configured to restore the format of the text content information based on the text box position information to obtain a target structured text;

[0042] A text box segmentation module, configured to segment the target structured text into text boxes based on the text box adhesion information to obtain a standard text;

[0043] A text marking module is used to mark the beginning and end of the standard text to obtain text content merging mark points;

[0044] A text merging module, configured to merge the standard texts based on the text content merging mark points to obtain a structured merged text;

[0045] The text input module is used to input the structured merged text.

[0046] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0047] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0048] The text entry method and device, electronic device and storage medium proposed in the present application can accurately extract text box position information and content information by acquiring structured text images and performing character recognition, wherein the text box position information includes text box adhesion information. Then, based on the text box position information, the format restoration is performed to reconstruct the original structure of the document and obtain the target structured text. The text box adhesion information is further used to realize text box segmentation to obtain standard text, and the standard text is marked at the beginning and end of the text to determine the text content merging mark point to realize text merging, and finally a structured merged text is formed and entered, which effectively improves the accuracy, efficiency and structured degree of text entry. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a flowchart of the text entry method provided by an embodiment of the present application;

[0050] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.

[0051] Figure 3 yes Figure 1 Flowchart of step S103 in FIG.

[0052] Figure 4 yes Figure 1 Flowchart of step S104 in FIG.

[0053] Figure 5 yes Figure 1 Flowchart of step S105 in FIG.

[0054] Figure 6 yes Figure 5 Flowchart of step S501 in FIG.

[0055] Figure 7 yes Figure 6 Flowchart of step S601 in FIG.

[0056] Figure 8 It is a structural diagram of the text entry device provided in an embodiment of the present application;

[0057] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0059] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0061] First, let’s analyze some of the terms used in this application:

[0062] Text Entry System: A text entry system is a tool or software used to enter text content into an editable, storable, and processable system or platform. Text entry systems digitize text information for subsequent editing, storage, retrieval, and analysis. Common text entry systems include document editing software, spreadsheet software, database systems, content management systems, and various specialized business systems.

[0063] Optical Character Recognition (OCR): OCR technology uses electronic devices to analyze and process printed images, identifying the characters within them and converting them into editable, searchable text. OCR can quickly digitize text within paper documents and images, facilitating subsequent editing, storage, and retrieval.

[0064] Text entry is used to enter text information from an image into a designated system. For example, in the fintech sector, text entry of insurance policies can be used to store them long-term, allowing business personnel to access policy information at any time. In the healthcare sector, text entry of medical records can be used to store them long-term, allowing doctors to quickly access patient information.

[0065] Currently, the method of text entry usually involves scanning the input text image, generating text content, and finally entering the generated text content directly into a specific system. However, in this text entry method, if the input text image is not placed flat, scanning errors are likely to occur, resulting in content errors in the entered text information. Therefore, how to improve the accuracy of text entry has become a technical problem that needs to be solved urgently.

[0066] Based on this, embodiments of the present application provide a text entry method and device, an electronic device, and a storage medium, aiming to improve the accuracy of text entry.

[0067] The text entry method and device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the text entry method in the embodiments of the present application is described.

[0068] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0069] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0070] The text entry method provided in the embodiment of the present application relates to the field of text processing technology and is applicable to the fields of financial technology and medicine. The text entry method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the text entry method, etc., but is not limited to the above forms.

[0071] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0072] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0073] Figure 1 This is an optional flowchart of the text entry method provided in the embodiment of the present application, which can be used in a text entry system. Figure 1The method may include but is not limited to steps S101 to S107.

[0074] Step S101, obtaining a structured text image;

[0075] Step S102: performing character recognition on the structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information;

[0076] Step S103: Based on the text box position information, the text content information is restored to a format to obtain a target structured text;

[0077] Step S104: Based on the text box adhesion information, the target structured text is segmented into text boxes to obtain standard text;

[0078] Step S105, marking the beginning and end of the standard text to obtain text content merging mark points;

[0079] Step S106, merging the standard text based on the text content merging mark points to obtain a structured merged text;

[0080] Step S107: input text into the structured merged text.

[0081] Steps S101 to S107 shown in the embodiment of the present application can accurately extract text box position information and content information by acquiring a structured text image and performing character recognition, wherein the text box position information includes text box adhesion information. Then, based on the text box position information, format restoration is performed to reconstruct the original structure of the document and obtain the target structured text. The text box adhesion information is further used to implement text box segmentation to obtain standard text, and by marking the beginning and end of the standard text, the text content merging mark point is determined to achieve text merging, and finally a structured merged text is formed and entered, which effectively improves the accuracy, efficiency and structured degree of text entry.

[0082] In step S101 of some embodiments, the structured text image refers to a text image with a clear format and layout, wherein the text content of the text image is arranged according to a certain structure, such as a bank account statement image, a hospital test report image, etc.

[0083] The embodiments of the present application can generate a structured text image by scanning a paper document specified by the user, or a document in image format downloaded from the Internet. In addition, a structured text image can be obtained by storing a document in the form of a screenshot. For example, in the financial field, a structured text image can be obtained by scanning a bank account statement or downloading an account statement document from an online bank. In the medical field, a structured text image can be obtained by collecting a scanned copy of a hospital test report or a screenshot in an electronic medical record system.

[0084] In step S102 of some embodiments, text box position information refers to the coordinates and size information of the text box in the structured text image, for example, in a bank statement image, the text box position information of each transaction record item, or in a medical prescription image, the position information of the text box such as the drug name, dosage, and usage. Text content information refers to the text content of the structured text image. Text box adhesion information refers to information about adhesion between text boxes. It should be noted that text box adhesion information may cause multiple independent text boxes to be mistakenly identified as one text box. For example, in a financial report image, two lines of text may be adhered to one text box due to the small line spacing. In a medical record image, the diagnosis results and treatment recommendations may be adhered to one text box due to the close layout.

[0085] The embodiment of the present application can determine the text area in the structured text image by identifying the text elements in the structured text image, and then extract the text box in the structured text image based on the text area, thereby obtaining the text box position information. Furthermore, text recognition can be performed on the text area to obtain text content information.

[0086] For details, see Figure 2 In some embodiments, step S102 may include but is not limited to steps S201 to S203:

[0087] Step S201, performing text region detection on the structured text image to obtain a text region;

[0088] Step S202: extracting text boxes from the structured text image based on the text area to obtain text box location information;

[0089] Step S203: extract text from the text area to obtain text content information.

[0090] In step S201 of some embodiments, the text region refers to a region in the structured text image that is identified as containing text content, and is usually defined by coordinates and size.

[0091] The embodiments of the present application can identify the area containing text in the structured text image through image processing and pattern recognition technology, and exclude non-text parts such as pictures and charts, thereby obtaining the text area.

[0092] In step S202 of some embodiments, each independent text box can be identified by analyzing the layout features of the text area, such as arrangement, spacing, and alignment, and then text box position information is generated based on the position of the text box in the structured text image.

[0093] In step S203 of some embodiments, by performing optical character recognition on each text element in the text area, the text content in the text area can be converted into editable and searchable text information, thereby obtaining text content information.

[0094] In steps S201 to S203 shown in the embodiment of the present application, by performing text area detection on the structured text image, the text area can be accurately located and the interference of irrelevant content can be eliminated, laying the foundation for subsequent text merging processing. Then, based on the detected text area, the text box position information is extracted to clearly define the text layout, which facilitates the understanding of the text structure. Finally, text extraction is performed in the located text area to obtain the text content, thereby realizing the conversion of image to editable text.

[0095] In step S103 of some embodiments, the target structured text refers to structured text content whose original format has been restored, and the logical structure and layout information of the text in the structured text image are retained.

[0096] The embodiment of the present application can obtain multiple text blocks by associating text box position information with text content information. Furthermore, the target structured text can be obtained by combining the text blocks according to the permutations and combinations in the text box position information.

[0097] For details, see Figure 3 In some embodiments, step S103 may include but is not limited to steps S301 to S302:

[0098] Step S301, associating the text box position information with the text content information to obtain a text block;

[0099] Step S302: Based on the text box position information, the text blocks are arranged and combined to obtain the target structured text.

[0100] In step S301 of some embodiments, a text block refers to a logical unit containing text content and its location information. For example, in a bank statement, a text area containing text content such as transaction date, amount, balance and its corresponding text box location information can be a transaction record text block. In a medical test report, a text area containing text content such as test item name, results, reference range and its corresponding text box location information can be a test item text block.

[0101] In the embodiment of the present application, each text box can be associated with the text content information at its corresponding position according to the text box position information, so as to obtain a text block.

[0102] In step S302 of some embodiments, the above text blocks are spliced ​​according to the coordinates provided in the text box position information to obtain a complete text, that is, a target structured text.

[0103] In steps S301 and S302 shown in the embodiment of the present application, the text content is accurately located at its location by associating the text box position information with the text content information, ensuring that the geometric relationship of the information is accurate. Subsequently, the text blocks are arranged and combined according to the text box position information, effectively restoring the structural layout of the original document in the structured text image, thereby greatly improving the usability and readability of the information.

[0104] In step S104 of some embodiments, standard text refers to text content that complies with standard formats and specifications.

[0105] The embodiment of the present application extracts features from text box adhesion information to obtain adhesion text box features, then performs image segmentation operations on the target structured text based on the features of each adhesion text box to obtain segmented text boxes, and finally uses the segmented text boxes to perform association processing on the target structured text, and ultimately obtains text that conforms to the standard format, that is, standard text.

[0106] For details, see Figure 4 In some embodiments, step S104 may include but is not limited to steps S401 to S403:

[0107] Step S401, extracting features of text box adhesion information to obtain adhesion text box features;

[0108] Step S402: performing image segmentation on the target structured text based on the adhesion text box feature to obtain segmented text boxes;

[0109] Step S403: performing association processing on the target structured text based on the segmented text frames to obtain a standard text.

[0110] In step S401 of some embodiments, the sticky text box feature refers to a feature set that can reflect the sticky condition of the text box, for example, the shape irregularity, color similarity, color distribution, i.e., text density, and other features of the sticky text box.

[0111] In the embodiment of the present application, the text box adhesion information can be parsed to obtain the text box edge, shape, area, color distribution and other features, namely, adhesion text box features.

[0112] In step S402 of some embodiments, the split text box refers to an independent text box that is no longer attached.

[0113] In the embodiment of the present application, based on the above-mentioned sticky text box feature, the sticky text box in the target structured text is separated from the overall target structured text to obtain a segmented text box.

[0114] In step S403 of some embodiments, the text content of the target structured text is reintegrated and associated based on the segmented text frames, so that the generated text conforms to the standard text format, ie, standard text.

[0115] In steps S401 to S403 shown in the embodiment of the present application, by extracting features of the text box adhesion information, the adhesion text box features are accurately identified, providing a key basis for subsequent image segmentation processing. On this basis, the target structured text is segmented according to the extracted adhesion text box features, the adhesion text box is effectively separated, and an independent segmented text box is generated, thereby significantly improving the recognition accuracy of the target structured text. Finally, the target structured text is associated with the segmented text box to restore the integrity and logic of the target structured text, and obtain a standard text that conforms to the standard format, which greatly improves the accuracy and efficiency of document entry.

[0116] In step S105 of some embodiments, the text content merging mark point refers to a mark point that identifies a text position that needs to be merged in the standard text.

[0117] The embodiment of the present application compares the similarity of the standard text to identify the part of the standard text that needs to be merged, that is, the target merged text. Thereafter, the position of the text elements in the target merged text is identified to determine the starting position and ending position of each text element. Then, marks are added to the starting position and ending position of each text element respectively to form a text merge starting mark point and a text merge ending mark point. Finally, the text merge starting mark point and the text merge ending mark point are integrated together to obtain a text content merge mark point.

[0118] For details, see Figure 5 In some embodiments, step S105 may include but is not limited to steps S501 to S505:

[0119] Step S501, performing similarity comparison on the standard text to obtain the target merged text;

[0120] Step S502: identifying the position of text elements in the target merged text to obtain the starting position and the ending position of the text elements;

[0121] Step S503: Mark the starting position of the text element to obtain a text merging starting mark point;

[0122] Step S504: Mark the end position of the text element to obtain the text merging end mark point;

[0123] Step S505 , integrating the text merging start mark point and the text merging end mark point to obtain the text content merging mark point.

[0124] In step S501 of some embodiments, the target merged text refers to the text content that needs to be merged in the standard text after similarity comparison.

[0125] The embodiment of the present application compares the text contents, ie, the typesetting formats, of various regions in the standard text, calculates the similarity between them, and selects the target merged text based on the similarity.

[0126] For details, see Figure 6 In some embodiments, step S501 may include but is not limited to steps S601 to S603:

[0127] Step S601, performing content similarity comparison on the standard text to obtain text content similarity;

[0128] Step S602, performing typesetting similarity comparison on the standard text to obtain text typesetting similarity;

[0129] Step S603 : Based on the text content similarity and text layout similarity, the standard text is screened for merging areas to obtain a target merged text.

[0130] In step S601 of some embodiments, the text content similarity is used to measure the similarity between two text segments in terms of content.

[0131] The embodiment of the present application obtains text content similarity by comparing the semantic similarity of adjacent texts in the standard text, and can determine whether there is text content that can be merged between the adjacent texts.

[0132] For details, see Figure 7 In some embodiments, step S601 may include but is not limited to steps S701 to S703:

[0133] Step S701, semantically analyze the standard text to obtain lexical semantics;

[0134] Step S702, calculating the similarity of the vocabulary semantics to obtain the vocabulary similarity;

[0135] Step S703: Based on the vocabulary similarity, context understanding is performed on the standard text to obtain text content similarity.

[0136] In step S701 of some embodiments, lexical semantics refers to the meaning and usage of a word in a specific context. For example, in a bank statement, the semantics of the word "transfer" indicates the transfer of funds from one account to another. In a medical report, the semantics of the word "blood sugar" indicates the glucose content in the blood.

[0137] The embodiment of the present application can obtain lexical semantics by deeply interpreting the text content of the standard text and then analyzing the meaning and contextual relationship of the words and phrases in the standard text.

[0138] In step S702 of some embodiments, the vocabulary similarity is used to measure the semantic similarity between two or more words.

[0139] In the embodiment of the present application, the degree of similarity of lexical semantics between two or more words can be calculated by using a cosine similarity algorithm, or the degree of similarity of lexical semantics between two or more words can be calculated by using a lexical distance algorithm, which is not limited in the present application.

[0140] In step S703 of some embodiments, the standard text is contextually understood based on the vocabulary similarity, and the usage and context of each vocabulary in the context of the standard text are comprehensively considered. Then, the text content similarity is calculated based on the usage and context of each vocabulary in the context of the standard text.

[0141] In steps S701 to S703 illustrated in the embodiment of this application, semantic parsing of the standard text extracts lexical meanings, i.e., understanding the basic meaning of each word in the text. The similarity between these lexical meanings is then calculated to obtain lexical similarity, thereby identifying semantically similar words in the text. Finally, based on lexical similarity, the standard text is contextually understood, comprehensively considering the usage and context of the words in the context to obtain text content similarity, which accurately measures the similarity of text content.

[0142] In step S602 of some embodiments, the text layout similarity is used to measure the similarity between two text segments in terms of layout features.

[0143] This application example can obtain the text layout similarity between adjacent texts by comparing the fonts, font sizes, line spacing, alignment, etc. of adjacent texts in the standard text.

[0144] In step S603 of some embodiments, the text content similarity and the text layout similarity may be weighted and summed according to a preset similarity weight ratio, thereby determining whether adjacent texts need to be merged.

[0145] In steps S601 to S603 shown in the embodiment of the present application, by comparing the content similarity of the standard text, the similarity of the text content can be identified, and the text content similarity can be obtained, which provides a key basis for subsequent merging. At the same time, the standard text is compared for typesetting similarity to obtain text typesetting similarity, further improving the accuracy of the merging from the typesetting dimension. Finally, based on these two similarities, the standard text is screened for merging areas to determine the target merged text, which can achieve efficient and accurate merging of texts and effectively improve the accuracy and efficiency of text entry.

[0146] In step S502 of some embodiments, the text element start position refers to the position where the text element starts in the text, and the text element end position refers to the position where the text element ends in the text.

[0147] The embodiment of the present application locates text elements such as characters, words, and numbers in the target merged text, and determines the starting position and ending position of the target merged text based on the positioning information, thereby obtaining the starting position and ending position of the text element.

[0148] In steps S503 and S504 of some embodiments, the text merge start mark point refers to the point at which the standard text merge begins, and the text merge end mark point refers to the point at which the standard text merge ends.

[0149] The embodiment of the present application adds a start mark at the starting position of the text element of the target merged text to obtain a text merge start mark point, and adds an end mark at the ending position of the text element of the target merged text to obtain a text merge end mark point.

[0150] In step S505 of some embodiments, a complete text content merging mark point may be formed by embedding the text merging start mark point and the text merging end mark point into a preset empty set.

[0151] In steps S501 to S505 shown in the embodiment of the present application, by performing a similarity comparison on the standard text, the text portion with similar content, i.e., the target merged text, can be identified. Then, the text element position is identified on the target merged text, and the starting position and ending position of each text element are clarified, thereby providing accurate position information for the text merging operation. Next, the starting and ending positions of the text elements are marked respectively to obtain the text merging starting mark point and the text merging ending mark point, thereby ensuring the accuracy and traceability of the merging operation. Finally, the text merging starting mark point and the text merging ending mark point are integrated to form a complete text content merging mark point, thereby realizing efficient merging of texts, ensuring the integrity and logic of the text content, and greatly improving the efficiency and quality of text entry.

[0152] In step S106 of some embodiments, the structured merged text refers to structured text content with complete semantics and logical structure, for example, a complete transaction record text after merging different parts of the same transaction, or a complete medical record text after merging multiple medical records of the same patient.

[0153] The embodiment of the present application identifies the text content merging mark point, determines the text merging start mark point and the text merging end mark point, and then performs text splicing on the text content between the text merging start mark point and the text merging end mark point to obtain a merged text, and then replaces the text content between the text merging start mark point and the text merging end mark point with the merged text to obtain a structured merged text.

[0154] In step S107 of some embodiments,

[0155] After obtaining the structured merged text, the embodiment of the present application can implement text entry of the structured merged text through text storage, thereby improving the accuracy of text entry.

[0156] This application can accurately extract text box position information and content information by acquiring structured text images and performing character recognition, wherein the text box position information includes text box adhesion information. Then, based on the text box position information, the format restoration is performed to reconstruct the original structure of the document and obtain the target structured text. The text box adhesion information is further used to realize text box segmentation to obtain standard text, and by marking the beginning and end of the standard text, the text content merging mark point is determined to realize text merging, and finally a structured merged text is formed and entered, which effectively improves the accuracy, efficiency and structured degree of text entry.

[0157] See also Figure 8 The present application also provides a text entry device that can implement the above-mentioned text entry method. The device includes:

[0158] The data acquisition module 801 is used to acquire a structured text image;

[0159] A character recognition module 802 is configured to perform character recognition on a structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information;

[0160] The format restoration module 803 is used to restore the format of the text content information based on the text box position information to obtain the target structured text;

[0161] The text box segmentation module 804 is used to segment the target structured text into text boxes based on the text box adhesion information to obtain standard text;

[0162] The text marking module 805 is used to mark the beginning and end of the standard text to obtain the text content merging mark points;

[0163] The text merging module 806 is used to merge the mark points based on the text content, perform text merging on the standard text, and obtain a structured merged text;

[0164] The text input module 807 is used to input the structured merged text.

[0165] The specific implementation of the text entry device is basically the same as the specific embodiment of the above-mentioned text entry method, and will not be repeated here.

[0166] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned text entry method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0167] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0168] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0169] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the text entry method of the embodiments of this application.

[0170] Input / output interface 903, used to implement information input and output;

[0171] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0172] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0173] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0174] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned text entry method is implemented.

[0175] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0176] The text entry method, text entry device, electronic device and storage medium provided in the embodiments of the present application obtain a structured text image, perform character recognition on the structured text image, obtain text box position information and text content information, wherein the text box position information includes text box adhesion information, based on the text box position information, perform format restoration on the text content information to obtain the target structured text, based on the text box adhesion information, perform text box segmentation on the target structured text to obtain standard text, perform text marking on the first and last text of the standard text to obtain text content merging mark points, perform text merging on the standard text based on the text content merging mark points to obtain structured merged text, and perform text entry on the structured merged text.

[0177] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0178] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0180] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0181] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0182] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0183] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0184] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0185] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0186] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0187] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A text entry method, characterized in that: The method comprises: Get structured text image; Performing character recognition on the structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information; Based on the text box position information, the text content information is restored to a plate format to obtain a target structured text; Based on the text box adhesion information, the target structured text is segmented into text boxes to obtain a standard text; Marking the beginning and end of the standard text to obtain text content merging mark points; Based on the text content merging mark points, the standard text is merged to obtain a structured merged text; Text entry is performed on the structured merged text.

2. The method according to claim 1, characterized in that Marking the beginning and end of the standard text to obtain text content merging mark points includes: Performing similarity comparison on the standard texts to obtain a target merged text; Performing text element position recognition on the target merged text to obtain the text element start position and the text element end position; Marking the starting position of the text element to obtain a text merging starting mark point; Marking the end position of the text element to obtain a text merging end mark point; Mark point integration is performed on the text merging start mark point and the text merging end mark point to obtain the text content merging mark point.

3. The method according to claim 2, characterized in that The performing similarity comparison on the standard text to obtain the target merged text includes: Comparing the content similarity of the standard text to obtain text content similarity; Performing typesetting similarity comparison on the standard text to obtain text typesetting similarity; Based on the text content similarity and the text layout similarity, the standard text is screened for merging areas to obtain the target merged text.

4. The method according to claim 3, characterized in that Comparing the content similarity of the standard text to obtain the text content similarity includes: Performing semantic analysis on the standard text to obtain lexical semantics; Calculating the similarity of the vocabulary semantics to obtain vocabulary similarity; Based on the vocabulary similarity, context understanding is performed on the standard text to obtain the text content similarity.

5. The method according to claim 1, wherein The step of performing format restoration on the text content information based on the text box position information to obtain a target structured text includes: Associating the text box position information with the text content information to obtain a text block; Based on the text box position information, the text blocks are arranged and combined to obtain the target structured text.

6. The method according to any one of claims 1 to 5, characterized in that The step of performing text frame segmentation on the target structured text based on the text frame adhesion information to obtain a standard text includes: Extracting features of the text box adhesion information to obtain adhesion text box features; Based on the adhesion text box feature, the target structured text is segmented to obtain a segmented text box; Based on the segmented text frames, the target structured text is subjected to association processing to obtain a standard text.

7. The method according to any one of claims 1 to 5, characterized in that The performing of character recognition on the structured text image to obtain text box position information and text content information includes: Performing text region detection on the structured text image to obtain a text region; Based on the text area, extracting text boxes from the structured text image to obtain text box position information; Perform text extraction on the text area to obtain the text content information.

8. A text entry device, characterized in that: The device comprises: A data acquisition module, used to acquire structured text images; a character recognition module, configured to perform character recognition on the structured text image to obtain text box position information and text content information, wherein the text box position information includes text box adhesion information; A format restoration module, configured to restore the format of the text content information based on the text box position information to obtain a target structured text; A text box segmentation module, configured to segment the target structured text into text boxes based on the text box adhesion information to obtain a standard text; A text marking module is used to mark the beginning and end of the standard text to obtain text content merging mark points; A text merging module, configured to merge the standard texts based on the text content merging mark points to obtain a structured merged text; The text input module is used to input the structured merged text.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the text entry method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text entry method according to any one of claims 1 to 7 is implemented.