A method, system, terminal and medium for ancient character recognition
By identifying and replacing traditional Chinese characters in ancient books with simplified Chinese characters, the problem of readers understanding the content of ancient books has been solved, the accuracy of recognition has been improved and the conversion process has been optimized, making it easier for readers to understand classical Chinese texts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-10
- Publication Date
- 2026-04-03
AI Technical Summary
Existing OCR text recognition technology is unable to effectively convert traditional Chinese characters into simplified Chinese characters, making it difficult for readers unfamiliar with traditional Chinese characters to understand the content in ancient books.
By acquiring images of ancient books, recognizing traditional Chinese characters and searching for corresponding simplified Chinese characters from a preset text database, the system replaces and generates converted characters, outputs and displays simplified Chinese character images, supports error correction and character accuracy calculation, and saves the converted characters for later use.
It simplifies traditional Chinese characters, making it easier for readers to understand classical Chinese texts, improving recognition accuracy, and simplifying the subsequent conversion process.
Smart Images

Figure CN116682118B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of character recognition technology, and in particular to a method, system, terminal and medium for recognizing ancient characters. Background Technology
[0002] OCR (Optical Character Recognition) refers to the process by which electronic devices (such as scanners or digital cameras) examine characters printed on paper and then use character recognition methods to translate the shapes into computer text; that is, scanning text materials and then analyzing and processing image files to obtain text and layout information.
[0003] Currently, ancient books serve a triple purpose: documents, archives, and books. The text recorded in ancient books is mostly in traditional Chinese characters. Existing OCR text recognition only converts the text in an image into text; the converted image of traditional characters remains in traditional characters. For readers unfamiliar with traditional characters, understanding the content of ancient books is difficult, hindering their comprehension of classical Chinese. Summary of the Invention
[0004] Firstly, to facilitate readers' understanding of classical Chinese, this application provides a method for recognizing ancient characters.
[0005] This application provides a method for recognizing ancient characters, which employs the following technical solution:
[0006] A method for identifying ancient characters includes:
[0007] Get text images;
[0008] Based on the text image, generate the original characters, which include traditional Chinese characters;
[0009] Based on the traditional Chinese characters, the corresponding simplified Chinese characters are searched from the preset text database to determine the simplified Chinese characters;
[0010] Replace traditional Chinese characters with simplified Chinese characters to generate converted characters;
[0011] Based on the original characters, output the first transformed image displaying the original characters;
[0012] Based on the converted characters, output the second converted image that displays the converted characters.
[0013] By adopting the above technical solution, original characters with the same simplified and traditional characters as the characters on the ancient text image are generated from the image. The simplified characters corresponding to the traditional characters in the original characters are searched in the character database. The traditional characters in the original characters can be replaced with simplified characters. The original characters are thus transformed into converted characters. A second converted image containing the original characters is then output based on the converted characters. A first converted image containing the original characters is also output based on the original characters. By comparing the first and second converted images, the simplified characters corresponding to the traditional characters can be observed. After the traditional characters in the ancient text are converted into simplified characters, it is easier for readers to understand the ancient text.
[0014] Preferably, the step of generating original characters based on text images specifically includes:
[0015] Based on text images, recognize text size and shape;
[0016] Based on the size and shape of the text, generate a character selection box, and select one character at a time.
[0017] Identify the text style of the text within the text capture box;
[0018] Determine the font based on the text style;
[0019] Generate original characters based on the shape and font of the text.
[0020] By adopting the above technical solution, character extraction frames can be generated to separate text based on text size and shape. Then, by determining the font of the text within the character extraction frame, the original character can be identified. This process of separating and individually recognizing a piece of text can improve the accuracy of the original character.
[0021] Preferably, after the step of outputting a first converted image displaying the original characters based on the original characters, the method further includes:
[0022] Obtain the error correction instruction corresponding to the original character;
[0023] Based on error correction instructions, generate and display modification controls on the human-computer interaction interface;
[0024] Based on the original character, multiple similar characters are searched from a preset text database. A modification control is then used to allow selection and input of similar characters.
[0025] Get similar characters and confirm the control modification command;
[0026] Modify the original characters based on the confirmation command of the control;
[0027] Update the original characters based on the modified original characters.
[0028] By adopting the above technical solution, when the original character is inconsistent with the character in the text image, the original character can be modified by modifying the control. After inputting a character that is similar in shape to the incorrect original character and the same as the character in the text image, the original character can be updated. After the original character is updated, the converted character can be updated accordingly, and the first converted image and the second converted image can also be updated accordingly.
[0029] Preferably, after the step of obtaining the confirmation instruction for similar characters and modifying the control, the method further includes:
[0030] Calculate the number of times the corresponding original character is generated;
[0031] Based on the confirmation command of the modification control, calculate the number of times the corresponding original character has been modified;
[0032] Based on the number of times the corresponding original character was generated and the number of times the corresponding original character was modified, the accuracy rate of the corresponding original character is generated and displayed.
[0033] By adopting the above technical solution, the accuracy rate of the original character can be determined by the number of times it is generated and modified. For original characters with a low accuracy rate, readers can pay special attention to whether the original character is misidentified so as to correct it in time.
[0034] Preferably, the step of generating a character extraction frame based on the character size and shape specifically includes:
[0035] Based on the size and shape of the text, multiple boundary points are generated, and each boundary point is distributed around the text.
[0036] Based on each boundary point, a model frame is generated, which is formed by connecting each boundary point in sequence;
[0037] Based on the model frame, a character capture frame is generated. The character capture frame is a rectangular frame, and there are at least four minimum gaps of a preset size between the model frame and the character capture frame.
[0038] By adopting the above technical solution, multiple boundary points are distributed around the characters based on their size and shape. These boundary points are connected to form a model frame, which can surround the characters. A character selection frame is generated outside the model frame, which can then be used to separate the characters so that one character selection frame can select one character.
[0039] Preferably, after the step of replacing traditional Chinese characters with simplified Chinese characters to generate converted characters, the process includes:
[0040] Based on the converted characters, generate converted text paragraphs;
[0041] Generate the original text paragraph based on the original characters;
[0042] Save the converted text paragraphs and their corresponding original text paragraphs to a preset text database;
[0043] After the step of generating the original characters based on the text image, the following steps are included:
[0044] Generate the original text paragraph based on the original characters;
[0045] Based on the original text paragraph, determine whether there is an original text paragraph in the preset text database with a similarity higher than the preset similarity value;
[0046] If so, retrieve and display the converted text paragraph corresponding to the original text paragraph;
[0047] Generate converted characters based on the text paragraphs to be converted;
[0048] Based on the converted characters, output and display the second converted image containing the converted characters;
[0049] If not, then based on the traditional Chinese character, search for the corresponding simplified Chinese character from the preset character database to determine the simplified Chinese character.
[0050] By adopting the above technical solution, after the ancient text is converted, the original characters and the converted characters are stored in the text database in the form of paragraphs. When it is necessary to convert ancient text with a similarity higher than the preset similarity value, the corresponding converted characters can be retrieved directly from the text database, which can save the step of converting between simplified and traditional characters.
[0051] Preferred options also include:
[0052] Based on text images, identify text colors and generate color labels;
[0053] Based on the color identifier and the original and converted characters corresponding to the color identifier, the original characters in the first converted image and the converted characters in the second converted image are color-coded.
[0054] By adopting the above technical solution, the colors of the main text and the seal text in ancient texts are significantly different. The color markings can distinguish between the main text and the seal text, and can minimize confusion between the main text and the seal text due to their similar colors.
[0055] Secondly, this application provides an ancient script recognition system, which adopts the following technical solution:
[0056] An ancient script recognition system, comprising:
[0057] The text image acquisition module is used to acquire text images;
[0058] The character generation module is used to generate raw characters based on text images;
[0059] The character conversion module is used to find the corresponding simplified Chinese character for the traditional Chinese character from a preset text database, determine the simplified Chinese character, replace the traditional Chinese character with the simplified Chinese character, and generate the converted character.
[0060] The image output module is used to output a first converted image displaying the original characters, and a second converted image displaying the converted characters, based on the converted characters.
[0061] By adopting the above technical solution, the text image acquisition module can acquire text images, the character generation module can generate original characters from the text images, the character conversion module can convert traditional Chinese characters in the original characters into simplified Chinese characters, and the image output module can output the first converted image and the second converted image for readers to read, which can facilitate readers' understanding of classical Chinese texts.
[0062] Thirdly, this application provides a smart terminal, which adopts the following technical solution:
[0063] A smart terminal includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed using the ancient script recognition method described above.
[0064] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:
[0065] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed by any of the above-mentioned ancient script recognition methods.
[0066] In summary, this application includes at least one of the following beneficial technical effects:
[0067] 1. By generating original characters with the same simplified and traditional characters as the characters in the ancient text image, the simplified characters corresponding to the traditional characters in the original characters are searched in the text database. The traditional characters in the original characters can be replaced with simplified characters. The original characters are thus transformed into converted characters. Then, a second converted image containing the original characters is output based on the converted characters. A first converted image containing the original characters is also output based on the original characters. By comparing the first and second converted images, the simplified characters corresponding to the traditional characters can be observed. After the traditional characters in the ancient text are converted into simplified characters, it is easier for readers to understand the ancient text.
[0068] 2. By counting the number of times the original character is generated and modified, the accuracy rate of the original character can be determined. For original characters with a low accuracy rate, readers can pay special attention to whether the original character is misidentified so as to correct it in time.
[0069] 3. After the ancient text is converted, the original characters and the converted characters are saved in the text database in the form of paragraphs. When it is necessary to convert ancient text with a similarity higher than the preset similarity value, the corresponding converted characters can be retrieved directly from the text database, which can save the step of converting between simplified and traditional characters. Attached Figure Description
[0070] Figure 1 This is a flowchart of an ancient character recognition method according to an embodiment of this application.
[0071] Figure 2 This is a partial flowchart of an ancient character recognition method according to an embodiment of this application, mainly showing S200-S240 and S400-S430.
[0072] Figure 3 This is a partial flowchart of an ancient character recognition method according to an embodiment of this application, mainly showing S201-S205, S202a-S202c and S1400-S1500.
[0073] Figure 4 This is a partial schematic diagram of an ancient character recognition method according to an embodiment of this application, used to illustrate the boundary points and character extraction frame.
[0074] Figure 5 This is a system module diagram of an ancient character recognition system according to an embodiment of this application.
[0075] Explanation of reference numerals in the attached figures:
[0076] 100. Boundary point;
[0077] 200. Character capture box. Detailed Implementation
[0078] The present application will be further described in detail below with reference to all the accompanying drawings.
[0079] This application discloses a method for ancient character recognition. (Refer to...) Figure 1 and Figure 2 Ancient character recognition methods include:
[0080] S100: Acquire text image;
[0081] Specifically, text images can be in formats such as jpg, png, and pdf. Image files can be obtained by scanning ancient books, photographing ancient books, or downloading images of ancient books from the internet.
[0082] S200: Generate raw characters based on text images, including traditional Chinese characters;
[0083] Specifically, text recognition is performed on text images using OCR technology. The original characters are converted from the characters in the text image, including both simplified and traditional Chinese characters, and the font of the original characters is a preset font, such as SimSun.
[0084] S300: Based on the traditional Chinese character, search for the corresponding simplified Chinese character from the preset text database and determine the simplified Chinese character;
[0085] Specifically, the text database stores simplified and traditional Chinese characters, with each traditional character corresponding to a corresponding simplified character. Therefore, the simplified character corresponding to a traditional character can be found in the text database, and the simplified character corresponding to a traditional character can be identified.
[0086] S400: Replaces traditional Chinese characters with simplified Chinese characters to generate converted characters;
[0087] Specifically, after determining the simplified Chinese characters corresponding to the traditional Chinese characters, the traditional Chinese characters in the original characters are replaced with simplified Chinese characters, thus converting the original characters into converted characters, all of which are simplified Chinese characters.
[0088] S410: Generate the converted text paragraph based on the converted characters;
[0089] Specifically, the text paragraphs are converted into text paragraphs generated based on the converted characters, and then the text paragraphs are converted into text format.
[0090] S420: Generate the original text paragraph based on the original characters;
[0091] Specifically, the original text paragraph is a text paragraph generated based on the original characters, and the original text paragraph is in text form.
[0092] S430: Save the converted text paragraph and the corresponding original text paragraph to the preset text database.
[0093] Specifically, after the converted text paragraphs and the original text paragraphs are stored in the text database in text form, it is convenient to find the corresponding converted text paragraphs in the text database through the original text paragraphs. Moreover, when converting the same original characters in the future, the corresponding converted text paragraphs can be found directly through the original text paragraphs.
[0094] S510: Based on the original characters, output and display the first converted image containing the original characters;
[0095] S520: Based on the conversion characters, output and display the second conversion image containing the conversion characters;
[0096] Specifically, the first transformed image shows the original characters, and the second transformed image shows the transformed characters. That is, the characters on the first transformed image are the same as those on the text image, and the characters on the second transformed image are all simplified Chinese characters. Through the second transformed image, it is convenient for readers to understand ancient texts, and by comparing the first transformed image and the second transformed image, it is convenient for readers to learn traditional Chinese characters.
[0097] S600: Obtain an error correction instruction corresponding to the original character;
[0098] Specifically, when there is a deviation between the original character and the character in the text image, the original character is modified. The error correction instruction can be an instruction triggered by the reader selecting a certain character in the original characters through a mechanical button, such as clicking on a certain character in the original characters with a mouse; it can also be obtained by the method of triggering through virtual buttons, such as pressing relevant virtual trigger buttons in the interface of the corresponding software to achieve acquisition.
[0099] S700: Generate and display a modification control on the human-computer interaction interface based on the error correction instruction;
[0100] S800: Search for multiple similar characters similar to the original character from a preset character database based on the original character, and the modification control is used for selecting and inputting similar characters;
[0101] Specifically, the similar characters are multiple characters with similar shapes to the original character. For example, if the original character is "祗", the similar characters are "祇", "衹", "袛", etc. If the character on the text image is "祇", then "祇" can be selected through the modification control and "祇" can be input.
[0102] S900: Obtain a confirmation instruction for the similar character and the modification control;
[0103] S1110: Modify the original character based on the confirmation instruction of the modification control;
[0104] S1210: Update the original character based on the modified original character.
[0105] Specifically, after inputting the correct similar character into the modification control and triggering the confirmation instruction, the incorrect character in the original character is modified to the correct similar character. That is, "祗" is modified to "祇". After the original character is updated, the steps of S300, S400, S510, and S520 are continued to be executed, and the update of the first transformed image and the second transformed image can be completed, so as to correct the characters displayed on the first transformed image and the second transformed image.
[0106] S1120: Calculate the generation times corresponding to the original character;
[0107] Specifically, the original character corresponding to the original character to be modified, and the generation times are the times of generating the original character based on the text image, such as the times of generating "祗" as described above.
[0108] S1220: Calculate the modification times of the corresponding original character based on the confirmation instruction of the modification control;
[0109] Specifically, after the user inputs the correct similar character through the modification control and triggers the confirmation instruction of the modification control to confirm the modification, this confirmation instruction is the number of modification times.
[0110] S1300: Generate and display the correct rate of the corresponding original character according to the generation times of the corresponding original character and the modification times of the corresponding original character.
[0111] Specifically, the generation times is n, the modification times is m, and the correct rate is (n - m) / n; through the correct rate, the error probability of the corresponding original character transformed from the text image can be known.
[0112] Refer to Figure 2 , after the step of S200: Generate the original character based on the text image, and the original character includes traditional Chinese characters, it further includes:
[0113] S210: Generate the original text paragraph according to the original character;
[0114] S220: Determine whether there is an original text paragraph in the preset text database whose similarity is higher than the preset similarity value according to the original text paragraph;
[0115] Specifically, the original text paragraph searches in the text database for whether there is a similar original text paragraph, and the preset similarity value can be numerical values such as 90% and 95%.
[0116] S231: If so, retrieve and display the converted text paragraph corresponding to the original text paragraph;
[0117] S240: Generate the converted character according to the converted text paragraph;
[0118] S520: Output and display the second converted image with the converted character according to the converted character.
[0119] Specifically, when the similarity value of the original text paragraph and the original text paragraph in the text database is the preset similarity value, the converted text paragraph corresponding to the original text paragraph in the text database is retrieved, and then the converted character is generated according to the retrieved converted text paragraph. In this way, the second converted image can be output according to the converted character, that is, the text image can be converted into the second converted image by searching in the text database, and the steps of S300 and S400 can be omitted.
[0120] If not, then execute S300: According to the traditional Chinese characters, search for the corresponding simplified Chinese characters in the preset character database to determine the simplified Chinese character characters.
[0121] Specifically, when there is no original text paragraph in the text database with a similarity value to the original text paragraph equal to the preset similarity value, then execute the steps of S300.
[0122] Refer to Figure 1 And Figure 3 , S200: The steps of generating the original characters based on the text image are specifically the following steps.
[0123] S201: Based on the text image, identify the text size and text shape;
[0124] S202: According to the text size and text shape, generate a character extraction frame 200, and one character extraction frame 200 encloses one character;
[0125] Specifically, the text size and text shape can determine the outline of the text. According to the outline of the text, generate a character extraction frame 200. The character extraction frame 200 can enclose the text, and in this way, each text can be separated. After separating the text, it is convenient to identify the text and reduce the recognition error rate.
[0126] S203: Identify the text style of the text within the character extraction frame 200;
[0127] S204: Based on the text style, determine the text font;
[0128] Specifically, there are various styles of characters recorded in ancient books. The corresponding style of characters can be searched in the text database according to the text style, and in this way, the text font in the text image can be confirmed according to the text style, such as fonts like official script, regular script, running script, etc.
[0129] S205: According to the text shape and text font, generate the original characters.
[0130] Specifically, after determining the glyph and font, it can be determined what character it is. For example, the regular script "正", and in this way, the original characters can be generated, which can improve the accuracy of generating the original characters.
[0131] S202: The steps of generating the character extraction frame 200 according to the text size and text shape are specifically:
[0132] S202a: According to the text size and text shape, generate multiple boundary points 100, and each boundary point 100 is distributed on the periphery of the text;
[0133] Refer to Figure 3 And Figure 4, specifically, the boundary point 100 is a point on the periphery of the text, and the boundary points 100 can be evenly spaced along the periphery of the text. For example, the boundary point 100 on the top of the character "正" is distributed along the periphery of the character "一".
[0134] S202b: Generate a model box based on each boundary point 100. The model box is formed by sequentially connecting each boundary point 100;
[0135] Specifically, by sequentially connecting adjacent boundary points 100 along the periphery of the text, a model box can be formed. The model box can enclose the text. For a scattered text shape, such as the character "二", there are two upper and lower model boxes.
[0136] S202c: Generate a character extraction box 200 based on the model box. The character extraction box 200 is a rectangular box, and there are at least four minimum gaps of a preset value between the model box and the character extraction box 200.
[0137] Specifically, one character extraction box 200 encloses one character. The model box is located inside the character extraction box 200, and the minimum interval between the model box and the character extraction box 200 is a preset value. The character extraction box 200 can thus enclose the characters and further separate multiple characters.
[0138] S1400: Identify the text color based on the text image and generate a color identifier;
[0139] Specifically, generally in ancient books, the colors of the main text characters and the seal characters are different and the color difference is obvious. For example, most of the main text characters are black, and most of the seal characters are red. Generate color identifiers of black and red to distinguish the main text characters from the seal characters.
[0140] S1500: Perform color marking on the original characters in the first transformed image and the transformed characters in the second transformed image according to the color identifier and the original characters and transformed characters corresponding to the color identifier.
[0141] Specifically, in the first transformed image and the second transformed image, distinguish the colors of the main text characters and the seal characters according to the color identifier, which can尽量避免 the problem of difficult distinction due to the same color of the main text characters and the seal characters. At the same time, the transformation with the maximum similarity of the text image can be performed, that is, the colors of the main text characters and the seal characters in the first transformed image and the second transformed image are the same as the colors of the main text characters and the seal characters in the text image.
[0142] Refer to Figure 5 , this application embodiment also discloses an ancient character recognition system, including: a text image acquisition module, a character generation module, a character transformation module, and an image output module.
[0143] The text image acquisition module is used to acquire a text image;
[0144] The character generation module is used to generate raw characters based on text images;
[0145] The character conversion module is used to find the corresponding simplified Chinese character for the traditional Chinese character from a preset text database, determine the simplified Chinese character, replace the traditional Chinese character with the simplified Chinese character, and generate the converted character.
[0146] The image output module is used to output a first converted image displaying the original characters, and a second converted image displaying the converted characters, based on the converted characters.
[0147] This application also discloses an embodiment of a smart terminal, including a memory and a processor. The processor can be a central processing unit such as a CPU or MPU, or a host system built around a CPU or MPU. The memory can be a storage device such as RAM, ROM, EPROM, EEPROM, FLASH, disk, or optical disk. The memory stores a computer program that can be loaded by the processor and executed using the above-described ancient character recognition method.
[0148] This embodiment also provides a computer-readable storage medium, which can be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. The computer-readable storage medium stores a computer program that can be loaded by a processor and executed using the aforementioned ancient character recognition method.
[0149] The implementation principle of the ancient Chinese character recognition method, system, terminal and medium in this application embodiment is as follows: after generating original characters that are the same as the characters on the text image, the traditional characters in the original characters are replaced with simplified characters to generate converted characters. Then, a first converted image is generated from the original characters, and a second converted image is generated from the converted characters. Readers can read the original ancient text based on the first converted image, and can read the simplified Chinese version of the ancient text through the second converted image, which can facilitate readers' understanding of the ancient text.
[0150] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for recognizing ancient characters, characterized in that: include: Get text images; Based on the text image, generate the original characters, which include traditional Chinese characters; Based on the traditional Chinese characters, the corresponding simplified Chinese characters are searched from the preset text database to determine the simplified Chinese characters; Replace traditional Chinese characters with simplified Chinese characters to generate converted characters; Based on the original characters, output the first transformed image displaying the original characters; Based on the converted characters, output and display the second converted image containing the converted characters; After the step of replacing traditional Chinese characters with simplified Chinese characters to generate converted characters, the following steps are included: Based on the converted characters, generate converted text paragraphs; Generate the original text paragraph based on the original characters; Save the converted text paragraphs and their corresponding original text paragraphs to a preset text database; After the step of generating the original characters based on the text image, the following steps are included: Generate the original text paragraph based on the original characters; Based on the original text paragraph, determine whether there is an original text paragraph in the preset text database with a similarity higher than the preset similarity value; If so, retrieve and display the converted text paragraph corresponding to the original text paragraph; Generate converted characters based on the text paragraphs to be converted; Based on the converted characters, output and display the second converted image containing the converted characters; If not, then based on the traditional Chinese character, search for the corresponding simplified Chinese character from the preset character database to determine the simplified Chinese character.
2. The ancient script recognition method according to claim 1, characterized in that: The step of generating original characters based on text images specifically includes: Based on text images, recognize text size and shape; Based on the size and shape of the text, generate a text selection box (200), and each text selection box (200) selects one character; Identify the text style of the text within the character capture box (200); Determine the font based on the text style; Generate original characters based on the shape and font of the text.
3. The ancient script recognition method according to claim 1, characterized in that: After the step of outputting a first converted image displaying the original characters based on the original characters, the method further includes: Obtain the error correction instruction corresponding to the original character; Based on error correction instructions, generate and display modification controls on the human-computer interaction interface; Based on the original character, multiple similar characters are searched from a preset text database. A modification control is then used to allow selection and input of similar characters. Get similar characters and confirm the control modification command; Modify the original characters based on the confirmation command of the control; Update the original characters based on the modified original characters.
4. The ancient script recognition method according to claim 3, characterized in that: After the step of obtaining the confirmation instruction for the similar characters and the modification control, the method further includes: Calculate the number of times the corresponding original character is generated; Based on the confirmation command of the modification control, calculate the number of times the corresponding original character has been modified; Based on the number of times the corresponding original character was generated and the number of times the corresponding original character was modified, the accuracy rate of the corresponding original character is generated and displayed.
5. The ancient script recognition method according to claim 2, characterized in that: The step of generating the character extraction frame (200) based on the character size and shape specifically includes: Based on the size and shape of the text, multiple boundary points (100) are generated, and each boundary point (100) is distributed around the text. Based on each boundary point (100), a model frame is generated, which is formed by connecting each boundary point (100) in sequence; Based on the model frame, a character capture frame (200) is generated. The character capture frame (200) is a rectangular frame. There are at least four minimum gaps of a preset size between the model frame and the character capture frame (200).
6. The ancient script recognition method according to claim 1, characterized in that: Also includes: Based on text images, identify text colors and generate color labels; Based on the color identifier and the original and converted characters corresponding to the color identifier, the original characters in the first converted image and the converted characters in the second converted image are color-coded.
7. An ancient script recognition system, which implements the ancient script recognition method as described in any one of claims 1 to 6, characterized in that: include: The text image acquisition module is used to acquire text images; The character generation module is used to generate raw characters based on text images; The character conversion module is used to find the corresponding simplified Chinese character for the traditional Chinese character from a preset text database, determine the simplified Chinese character, replace the traditional Chinese character with the simplified Chinese character, and generate the converted character. The image output module is used to output a first converted image displaying the original characters, and a second converted image displaying the converted characters, based on the converted characters.
8. A smart terminal, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described in any one of claims 1 to 6 for the ancient script recognition method.
9. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 6 for the ancient script recognition method.
Citation Information
Patent Citations
Ancient book traditional and simplified Chinese character conversion method and device
CN112270201A
Character definition conversion method and system based on OCR (Optical Character Recognition) technology, terminal and medium
CN114220109A