Method for digital processing of information

By classifying and reading document information data and analyzing its components, the problem of inaccurate text data recognition was solved, and accurate digital processing of text data was achieved.

CN119206731BActive Publication Date: 2026-04-10HUNAN INST OF APPLIED TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN INST OF APPLIED TECH
Filing Date
2024-07-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify text content in non-text data when recognizing document information, especially due to the personalized writing styles of text, such as cursive and semi-cursive scripts, leading to numerous text recognition errors.

Method used

By collecting information data, dividing it into directly identifiable and non-directly identifiable data, scanning non-text data areas, obtaining initial identifiable text data, querying radical information and performing verification analysis, determining the accuracy of the text, and adjusting or marking the output of the final result.

Benefits of technology

It achieves accurate recognition of text data in different writing styles, reduces text recognition errors, and ensures the accuracy of information digitization processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206731B_ABST
    Figure CN119206731B_ABST
Patent Text Reader

Abstract

The application provides an information digitization processing method, and relates to the technical field of data processing, and comprises the following steps: S1, collecting information data, directly identifying and reading obtainable text data; S2, scanning non-text data, obtaining initial identified text data corresponding to the corresponding scanning position; S3, inquiring the radical information of each character in the initial identified text data, checking and analyzing each radical information at the scanning position, and obtaining matched confirmed text data; and S4, inserting the confirmed text data into the obtainable text data, and obtaining digitization processing result data. The application can more accurately realize text digitization identification processing of information data under different types of conditions, and can also accurately inquire whether the radical is accurate according to individualization of writing, and ensures the accuracy of information conversion to characters.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an information digitization processing method. BACKGROUND

[0002] With the continuous development of document processing technology, a large amount of readable text content and some other formats, such as pictures, formula formats, table formats, and the like, are contained in document data, which are displayed in the form of picture text, picture table, picture table, and formulas and tables that cannot be directly recognized, thereby causing difficulty in recognizing document information data and inaccuracy.

[0003] Therefore, when the existing document information data is recognized and processed, it is difficult to accurately confirm the text content in the non-text data, and due to the individualization of text writing, such as running script and cursive script, the recognized radicals and components may not be accurate, thereby causing a large number of text recognition errors in the recognition result. SUMMARY

[0004] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide an information digitization processing method to solve the problem that it is difficult to accurately confirm the text content in the non-text data when the document information data is recognized and processed in the prior art, and a large number of text recognition errors are caused due to the individualization of text writing.

[0005] To achieve the above-mentioned purpose and other related purposes, the present application provides an information digitization processing method, comprising the following steps:

[0006] S1: collecting information data, directly recognizing and reading recognizable text data;

[0007] S2: scanning non-text data to obtain initial recognized text data corresponding to the scanning position;

[0008] S3: querying the radical information of each character in the initial recognized text data, and analyzing and checking each radical information at the scanning position to obtain matched confirmed text data;

[0009] S4: inserting the confirmed text data into the recognizable text data to obtain digitization processing result data.

[0010] In an embodiment of the present application, in S1, the information data is collected, comprising:

[0011] obtaining information data to be read, and dividing the information data into directly recognizable character data and other data that cannot be directly recognized.

[0012] In an embodiment of the present application, the S1, the directly recognizable text data obtained by reading includes:

[0013] The directly recognizable text data is read according to the sequence in the information data, and the directly recognizable text data is obtained.

[0014] In an embodiment of the present application, the S2, the non-text data is scanned, and includes:

[0015] When the non-directly recognizable text data is divided, the coordinate region of the corresponding position of the non-directly recognizable text data is obtained as a non-text data region, and the image of the non-text data region is controlled to be intercepted and scanned.

[0016] In an embodiment of the present application, the S2, the initial recognition text data corresponding to the scanning position is obtained, and includes:

[0017] The processing number corresponding to each scanning is established, and the processing number is bound to the coordinate region and the initial recognition text data obtained by scanning.

[0018] In an embodiment of the present application, the S3, the information of the radicals of each character in the initial recognition text data is queried, and includes:

[0019] Each character in the initial recognition text data is called out, and the information of the radicals of each character is analyzed, and all writing image information corresponding to each radical information is obtained.

[0020] In an embodiment of the present application, the S3, each radical information is analyzed by checking at the scanning position, and includes:

[0021] According to the arrangement sequence of the radicals in the initial recognition text data, the image of the intercepted non-text data region corresponding to the scanning position is scanned again, it is judged whether the graphic information obtained by scanning again corresponds to the radical information of the character, if it corresponds, the initial recognition text data is output, if it does not correspond, the radical corresponding to the scanning again is controlled to be queried again.

[0022] In an embodiment of the present application, it is judged whether the graphic information obtained by scanning again corresponds to the radical information of the character, and includes:

[0023] The graphic information obtained by scanning again is compared with each writing image information corresponding to the radical information, and it is judged whether the similarity of the two satisfies a preset requirement.

[0024] In an embodiment of the present application, the S3, the radical corresponding to the scanning again is controlled to be queried again, and includes:

[0025] The corresponding partial character is scanned again and compared in the partial character database, and a partial character reaching a preset similarity is found, the found partial character is combined with the rest of the character, and the character is determined according to the combination.

[0026] In an embodiment of the present application, the character is determined according to the combination, comprising:

[0027] It is judged whether the pattern obtained by the combination is a pre-stored character, if not, the control mark is displayed, and if yes, the character obtained by the combination is directly generated.

[0028] As described above, the information digital processing method has the following beneficial effects: the read data is directly classified and read according to the document information data to be analyzed, i.e. the recognized text data is obtained. As another part of the document information data, i.e. the non-text data, the region of the non-text data needs to be analyzed by image interception. Specifically, the intercepted image is initially recognized to obtain initial recognition text data. The initial recognition text data is analyzed by partial character to determine all writing methods of the current partial character, such as the image of cursive script, clerical script, etc. Then the intercepted region is scanned again, and the partial characters are compared. The initial recognition text data can be accurately analyzed. If not accurate, the partial character is adjusted and replaced, and it is determined whether the combined image after replacement is a normal character. If not, the control mark is output and displayed. If it is a normal character, the normal character is directly output to form an accurate information digital processing result. Through the above method, the text digital recognition processing of information data in different types of situations can be accurately realized, and the accuracy of the query of the partial character according to the individualization of writing can be ensured, and the accuracy of the conversion of information to characters is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 The flowchart of the digital processing method of the present application. DETAILED DESCRIPTION

[0030] Following, the embodiments of the present application are illustrated by specific examples, and other advantages and effects of the present application can be easily understood by those skilled in the art from the disclosure of the present specification. The present application can also be implemented or applied by other different specific embodiments, and various modifications or changes can be made to the details in the present specification based on different views and applications without departing from the spirit of the present application. It should be noted that the following examples and features in the examples can be combined with each other without conflict. It should also be understood that the terms used in the embodiments of the present application are for describing specific specific embodiments, not for limiting the protection scope of the present application. The test methods in the following examples are not specified, and are generally carried out under conventional conditions or under the conditions recommended by the manufacturers.

[0031] It should be understood that the structures, proportions, sizes, etc. shown in the drawings attached to the present specification are only used to illustrate the content disclosed in the present specification, to be understood and read by those skilled in the art, and do not have technical significance to limit the conditions under which the present application can be implemented. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effects and purposes that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" in the present specification are only for the convenience of clear description, not for limiting the scope of the present application, and the change or adjustment of the relative relationship without substantially changing the technical content is also considered as the scope of the present application.

[0032] Please refer to Figure 1 The present application provides an information digitization processing method, comprising the following steps:

[0033] S1: collecting information data, directly identifying and reading obtainable text data;

[0034] S2: scanning non-text data to obtain initial identification text data corresponding to the corresponding scanning position;

[0035] S3: querying the radical information of each character in the initial identification text data, and checking and analyzing each radical information at the scanning position to obtain matched confirmation text data;

[0036] S4: inserting the confirmation text data into the obtainable text data to obtain digitization processing result data.

[0037] It can be found from the above that in the process of digital recognition of the data content of the document data, the document data needs to be converted into a corresponding processing format first. Then the information data in the document data is collected and divided into recognizable text data and non-text data. The non-text data includes unrecognizable character data, formula data, table data, etc. The information data is directly read first to obtain the recognizable text data. In the reading process, non-text data is also read. When the non-text data is scanned, the initial recognition of the scanning position and the corresponding initial recognition text data is recorded. In order to verify the accuracy of the initial recognition text data, the radical information of the initial recognition text data is queried. According to the radical information, the scanning position is scanned again to determine whether the initial recognition text data obtained by the initial scanning is accurate. After determining whether it is accurate, the confirmed text data is obtained and inserted into the corresponding readable text data, so as to realize the digital processing result of the information data.

[0038] In the step S1, the information data is collected, including:

[0039] The information data to be read is obtained, and the information data is divided into directly recognizable character data and other data which cannot be directly recognized.

[0040] In an embodiment of the present application, when the information data is collected, there may be various types of information data in the document. That is, the directly recognizable character data corresponding to the processing system, and other data including unrecognizable character data, formula data and table data, etc. All the data in the information data are recognized and processed.

[0041] Further, in the step S1, the directly recognizable text data is obtained by directly recognizing and reading, including:

[0042] The directly recognizable character data is read according to the order in the information data to obtain the recognizable text data.

[0043] In an embodiment of the present application, when the directly recognizable text data is read, in order to ensure that the finally obtained digital processing data can maintain the order of the original information data, the directly recognizable character data in each region needs to be read in order to obtain the recognizable text data. It is worth noting that when the initial recognition text data is inserted in the directly recognizable character data arranged in order, the position of the initial recognition text data corresponding to the recognizable text data is also recorded to be inserted in the corresponding order after the initial recognition text data is obtained.

[0044] Specifically, in the step S2, the non-text data is scanned, including:

[0045] When the data that cannot be directly recognized is divided, a coordinate region of a corresponding position that cannot be directly recognized is obtained as a non-text data region, and an image of the non-text data region is controlled to be intercepted and scanned.

[0046] In an embodiment of the present application, when non-text data is scanned, the data that cannot be directly recognized needs to be divided. That is, when data that cannot be directly recognized is scanned, a coordinate region of a corresponding region that is recognized as data that cannot be directly recognized is obtained, and then is used as a non-text data region. That is, when the non-text data region is obtained, the non-text data region needs to be image-intercepted, and in order to realize recognition, the image is scanned again to obtain initial recognized text data.

[0047] Further, in the step S2, the initial recognized text data corresponding to the scanning position is obtained, and includes:

[0048] A processing number corresponding to each scanning is established, and the processing number is bound to the coordinate region and the initial recognized text data obtained by scanning.

[0049] In an embodiment of the present application, when initial recognized text data is obtained, in order to better realize processing of the data, by establishing a processing number corresponding to each image scanning when scanning, and binding each processing number, a coordinate region, and initial recognized text data obtained by scanning, a bound processing number-coordinate region-initial recognized text data is obtained. Then, according to one of the three, the other two types of data information can be quickly obtained, so as to facilitate quick query.

[0050] Next, in the step S3, the information of radicals of each character in the initial recognized text data is queried, and includes:

[0051] Each character in the initial recognized text data is called out, and the information of radicals of each character is analyzed, and all writing image information corresponding to each radical information is obtained.

[0052] In an embodiment of the present application, when the radicals of characters in the initial recognized text data are queried, in order to better realize analysis, each character in the initial recognized text data obtained by scanning is called out first, and then the radical information of each character is analyzed. After the radical information is obtained, according to the calligraphy writing characteristics corresponding to the radical, all image information of writing calligraphy corresponding to the corresponding radical information is called out, so as to determine whether the image intercepted in the current information data conforms to the radical corresponding to the initial recognized text data.

[0053] Further, in the step S3, the checking and analyzing of each of the radicals in the scanning position includes:

[0054] According to the arrangement order of the radicals in the initial recognized text data, the image of the intercepted corresponding non-text data region is scanned again, it is judged whether the graph information obtained by the re-scanning corresponds to the radical information of the character, if corresponding, the initial recognized text data is output, if not corresponding, the corresponding radical of the re-scanning is controlled to be re-inquired.

[0055] In an embodiment of the present application, when the radical information is analyzed and checked, the re-scanning mode of the radical is obtained according to the arrangement of the position of the radical corresponding to the character in the initial recognized text data. For example, when the radical is a left-right structure, the left radical is recognized from left to right or the right radical is recognized from right to left. When the radical is an up-down structure, the upper radical is recognized from top to bottom or the lower radical is recognized from bottom to top. Or when there is no radical, the corresponding stroke order is used to scan and recognize the whole image again. The graph information obtained by the re-scanning is compared with the position radical corresponding to the initial recognized text data, if corresponding, the initial recognized text data is output, if not corresponding, the corresponding radical of the re-scanning is controlled to be re-inquired.

[0056] Specifically, judging whether the graph information obtained by the re-scanning corresponds to the radical information of the character includes:

[0057] The graph information obtained by the re-scanning is compared with each of the writing image information corresponding to the radical, it is judged whether the similarity of the two satisfies the preset requirement.

[0058] In an embodiment of the present application, after the graph information is obtained by the re-scanning, the graph data is compared with the position radical corresponding to the initial recognized text data, and it is determined whether the similarity of the radical reaches the preset requirement. If reaching, the initial recognized text data is output, if not reaching, the corresponding radical of the re-scanning is controlled to be re-inquired.

[0059] Then, in the step S3, the corresponding radical of the re-scanning is controlled to be re-inquired, including:

[0060] The corresponding radical of the re-scanning is compared and analyzed in the radical database, the radical reaching the preset similarity is found, the found radical is combined with the remaining part of the character, and the character is determined according to the combination.

[0061] In an embodiment of the present application, when the partial character to be re-scanned does not reach the partial character that needs to be controlled to re-query the scanned image, the re-scanned partial character is queried against the partial character database. The partial character database stores all partial characters, such as left partial characters, right partial characters, upper partial characters, or lower partial characters. It is determined whether the re-scanned partial character has a preset similarity with the partial character data. If yes, the corresponding partial character is outputted, and the other re-scanned partial characters are re-fused to obtain a combined character. It is further determined whether the combined character is a normal character.

[0062] Further, the character is determined according to the combination, comprising:

[0063] It is determined whether the combined character is a pre-stored character. If no, a mark is displayed. If yes, the combined character is directly generated.

[0064] In an embodiment of the present application, when it is determined whether the character is a normal character, it is determined whether the combined position character is the same as the pre-stored character determined as the character. If yes, the combined character is outputted. If no, the character is incorrect, and the combined character is marked for display, such as highlighted display, to alert.

[0065] The present application further provides an information digitization processing system, comprising:

[0066] A text recognition unit is configured to collect information data, and directly recognize and read recognizable text data.

[0067] An initial recognition unit is configured to scan non-text data, and obtain initial recognition text data corresponding to a scanning position.

[0068] A text confirmation unit is configured to query partial character information of each character in the initial recognition text data, and analyze and check each partial character information at the scanning position to obtain matched confirmation text data.

[0069] A result output unit is configured to insert the confirmation text data into the recognizable text data to obtain digitization processing result data.

[0070] In summary, the present application directly classifies and reads the readable poetic sentences by analyzing the document information data as required, i.e. obtains the recognizable text data. As another part of the document information data, i.e. the non-text data, the image of the region where the non-text data is located needs to be intercepted and analyzed. Specifically, the intercepted image is initially recognized, i.e. obtains the initial recognition text data. The initial recognition text data is analyzed by radical and stroke analysis, all writing methods of the current radical and stroke are determined, such as the image of cursive script, clerical script, etc., then the intercepted region is scanned again, and the radicals and strokes are compared, whether the initial recognition text data is accurate can be accurately analyzed, if not, the radical and stroke are controlled to be adjusted and replaced, and whether the replaced combined image is a normal character is determined, if not, the output display is marked, if it is a normal character, the normal character is directly outputted, to constitute the accurate information digital processing result. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has high industrial utilization value.

[0071] The above embodiments only exemplarily illustrate the principles and effects of the present application, and are not used to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those skilled in the art without departing from the spirit and technical thought disclosed by the present application should be covered by the claims of the present application.

Claims

1. A method for digitizing information, characterized in that, The method includes the following steps: S1: Collect information data and directly identify and read recognizable text data; S2: Scan non-text data to obtain initial recognized text data corresponding to the corresponding scan position, including: When dividing the data into regions containing unrecognizable characters, the coordinate regions of the corresponding positions of the unrecognizable characters are obtained as non-text data regions. The image of the non-text data regions is controlled to be scanned. When data containing unrecognizable characters is scanned, the coordinate regions of the regions that are identified as unrecognizable characters are obtained and then used as non-text data regions. When the non-text data regions are obtained, the images of the non-text data regions need to be cropped. In order to achieve recognition, the images are scanned again to obtain the initial recognized text data. Establish a processing number corresponding to each scan, and bind the processing number to the coordinate region and the initial recognition text data obtained by the scan. By establishing a processing number corresponding to each image scan during the scan, and binding each processing number, coordinate region and the initial recognition text data obtained by the scan, the bound processing number-coordinate region-initial recognition text data is obtained. Then, based on one of the three, the other two types of data information can be obtained. S3: Query the radical information of each character in the initial recognized text data, and perform verification analysis on each radical information at the scanning position to obtain matching confirmation text data, including: According to the arrangement order of radicals in the initial recognized text data, the image of the corresponding non-text data area is scanned again, and it is determined whether the graphic information obtained by the rescan corresponds to the radical information of the text. If they correspond, the initial recognized text data is output. If they do not correspond, the corresponding radical is re-queried and re-scanned. When analyzing and verifying radical information, it is necessary to determine the re-scanning method of the radicals based on the arrangement of the radicals corresponding to the characters in the initial recognition text data. When the radical is a left-right structure, the left radical will be identified from left to right, or the right radical will be identified from right to left; when the radical is a top-bottom structure, the upper radical will be identified from top to bottom, or the lower radical will be identified from bottom to top; when there are no radicals, the entire image will be scanned and identified again according to the corresponding stroke order. S4: Insert the confirmed text data into the recognizable text data to obtain the digitized processing result data.

2. The information digitization processing method according to claim 1, characterized in that: In step S1, the collected information data includes: The information data to be read is obtained and divided into data with directly recognizable text and other data with non-recognizable text.

3. The information digitization processing method according to claim 2, characterized in that: In step S1, the recognizable text data is directly identified and read, including: The directly recognizable text data is read in the order of the information data to obtain the recognizable text data.

4. The information digitization processing method according to claim 1, characterized in that: In step S3, querying the radical information of each character in the initial recognized text data includes: Each character in the initial recognized text data is retrieved, and the radical information of each character is analyzed, as well as all the written image information corresponding to each radical information, are obtained.

5. The information digitization processing method according to claim 1, characterized in that: Determining whether the graphic information obtained from the second scan corresponds to the radical information of the text includes: The graphic information obtained from the rescan is compared with the written image information corresponding to the radical information to determine whether the similarity pairs meet the preset requirements.

6. The information digitization processing method according to claim 5, characterized in that: In S3, controlling the re-querying and re-scanning of the corresponding radicals includes: The corresponding radicals will be scanned again for comparison and analysis in the radical database. Radicals that reach a preset similarity will be found, and the found radicals will be combined with the rest of the text. The text will be determined based on the combination.

7. The information digitization processing method according to claim 6, characterized in that: The text is determined based on the combination, including: Determine whether the combined graphic is a pre-stored text. If not, control the display of the marker; if so, directly generate the combined text.

Citation Information

Patent Citations

  • Electronic file generation method and device, equipment and medium

    CN116416629A

  • Document identification method, document translation method and related equipment

    CN117935294A