Method and system for audio conversion of Zhuang ancient books and literatures

By using large-scale datasets and deep learning models, combined with OCR and TTS technologies, we have achieved automated recognition and speech synthesis of ancient Zhuang characters, solved the digitization problem of ancient Zhuang characters, improved the recognition rate and pronunciation accuracy, and promoted the dissemination of ancient Zhuang culture.

CN120977284APending Publication Date: 2025-11-18GUANGXI UNIV FOR NATITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511121682.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and digitize ancient Zhuang characters, especially in handwritten documents or blurry images, failing to effectively convey the cultural connotations of ancient Zhuang characters.

Method used

By employing large-scale datasets and deep learning models, combined with OCR technology and TTS systems, we achieve automated recognition and speech synthesis of ancient Zhuang characters. We identify characters through convolutional neural networks and long short-term memory networks, construct an ancient Zhuang character mapping lexicon, and generate natural and fluent speech using the Transformer self-attention mechanism.

Benefits of technology

It improves the recognition rate and digitization efficiency of ancient Zhuang characters, generates natural speech that conforms to the characteristics of the Zhuang language, realizes the automated audio digitization of ancient Zhuang texts, and supports the accurate pronunciation of polyphonic characters and the dissemination of culture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977284A_ABST
    Figure CN120977284A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for audio processing of Zhuang and ancient books, and the method achieves the automatic processing through five core modules: firstly, an ancient book image processing module carries out the denoising, binaryzation, character segmentation and quality optimization of an input image; secondly, establishing a corresponding relation between fonts and phonetic symbols by an ancient-Zhuang character mapping lexicon construction module; thirdly, the optical character recognition module adopts a CNN and LSTM combined learning model to output a standard text and a corresponding phonetic symbol; then, a text-to-speech module optimizes a speech synthesis model according to the six-tone characteristics of Zhuang language; and finally, the system integration module realizes an end-to-end automatic process from image input to audio output. According to the method, the technical blank of ancient Zhuang character digital processing is filled, the ancient book recognition efficiency and the speech synthesis naturalness are improved, the Zhuang nationality cultural heritage can be stored and spread conveniently, the Zhuang nationality cultural heritage can be popularized to groups such as visually impaired people conveniently, and meanwhile the potential of being expanded to other minority ancient book processing is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to digital technology, optical character recognition (OCR) and speech synthesis (TTS) technology, especially suitable for the digital conversion of ancient literature, which can realize the automatic conversion from ancient book images to audio, and is particularly aimed at processing minority languages such as ancient Zhuang characters, providing audio format ancient literature for the visually impaired and ancient literature lovers, and helping cultural preservation and dissemination. BACKGROUND

[0002] Ancient Zhuang characters are derived from Chinese characters and have both Chinese characters and Zhuang language characteristics, and are an important carrier of Zhuang culture, recording traditional culture, classical literature, genealogy, etc., but have problems such as many variant characters, complex structure, various writing forms, and non-unified system, increasing the difficulty of recognition.

[0003] Ancient Zhuang character texts carry rich Zhuang cultural heritage, but due to aging of writing materials, text damage, and lack of recognition by the younger generation, they face the risk of being lost, and efficient and reliable ancient Zhuang character digitization technology combined with speech output is urgently needed to realize the intuitive dissemination and preservation of culture.

[0004] Existing minority language digitization technology is mostly focused on modern pinyin and standardized characters, and there is little research on ancient Zhuang characters. Traditional OCR technology has low recognition rate for ancient Zhuang characters, especially on handwritten literature or blurred images, and cannot effectively handle complex characters and handwritten bodies. SUMMARY

[0005] The technical problem to be solved by the present application is to overcome the above technical defects and provide a Zhuang text ancient literature audio method and system.

[0006] To solve the above problems, the technical scheme of the present application is: Zhuang text ancient literature audio method and system,

[0007] Further, (from the right)

[0008] The present application has the following advantages compared with the existing technology:

[0009] (1) Automatic processing of multi-sound character problem: with the support of large-scale data set and the self-adaptive ability of deep learning model, the model can automatically select the correct pronunciation of multi-sound character according to the context, without the need for a separate disambiguation step.

[0010] (2) Efficient image recognition and text conversion: advanced OCR technology and handwriting recognition model can handle complex characters in ancient books, greatly improving the digital efficiency of ancient literature.

[0011] (3) Natural and fluent speech synthesis: Through the optimized TTS model, the generated speech conforms to the language characteristics of Zhuang language, ensuring accurate pronunciation and natural tone, and effectively conveying the cultural connotation of ancient literature.

[0012] (4) High automation and scalability: The system not only supports the digital transformation of existing ancient literature, but also has scalability, which can adapt to the digitalization needs of other minority languages and ancient literature. Through further accumulation and updating of data sets, the system can continuously optimize pronunciation effect and recognition accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is the ancient book digitization preprocessing diagram of the present application.

[0014] Figure 2 is the ancient Zhuang character mapping vocabulary construction module diagram of the present application.

[0015] Figure 3 is the ancient book image to speech synthesis automation system flowchart of the present application.

[0016] Figure 4 is the optical character recognition module diagram of the present application.

[0017] Figure 5 is the text-to-speech module diagram of the present application.

[0018] Figure 6 is the character shape similarity matching flowchart diagram of the present application. DETAILED DESCRIPTION

[0019] The specific embodiments of the present application will be further described below in conjunction with the accompanying drawings. Wherein the same parts are denoted by the same reference numerals.

[0020] It should be noted that the words "front", "back", "left", "right", "up" and "down" used in the following description refer to the directions in the drawings, and the words "in" and "out" refer to the directions towards or away from the geometric center of a particular part.

[0021] In order to make the content of the present application more easily understood, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present application.

[0022] As shown in Figures 1 to 6 The ancient Zhuang literature audio method of the present application realizes the conversion from ancient book image to audio through the end-to-end automation process of ancient book image processing, optical character recognition (OCR), ancient Zhuang character mapping vocabulary matching, text-to-speech (TTS) synthesis and audio output, without human intervention throughout. The user only needs to upload the ancient book image, and the system can automatically complete all processing steps:

[0023] Image processing: preprocessing of input Zhuang ancient book images for denoising, binarization, layout analysis and character segmentation, and image quality optimization, using Gaussian filtering or median filtering algorithm for image denoising, detecting text regions through convolutional neural network, and segmenting out text and non-text regions;

[0024] OCR recognition and phonetic symbol mapping: using a model combining CNN and LSTM to recognize preprocessed images of ancient Zhuang characters, and outputting standardized text and corresponding phonetic symbols by matching the similarity of character shapes in the ancient Zhuang character mapping vocabulary, wherein for multi-pronunciation characters, the correct pronunciation is selected in combination with the context, the ancient Zhuang character mapping vocabulary is constructed by collecting standard printed ancient Zhuang character images, labeling phonetic symbols for each character, and annotating different pronunciations in different contexts for multi-pronunciation characters, and supporting dynamic updates through manual or "crowdsourcing" methods;

[0025] Speech synthesis: based on the output phonetic symbols, modeling and processing are performed according to the intonation characteristics of Zhuang language to generate speech, multi-pronunciation characters are automatically disambiguated through Transformer self-attention mechanism, and Tacotron2, WaveNet and other deep neural networks are used to convert phonetic symbols into audio waveforms;

[0026] Audio post-processing: denoising and volume balancing optimization of generated speech to output the final audio.

[0027] A Zhuang ancient book document audio system, characterized by comprising:

[0028] Ancient book image processing module: for preprocessing of input Zhuang ancient book images for denoising, binarization, layout analysis and character segmentation, and image quality optimization, the ancient book image processing module uses Gaussian filtering or median filtering algorithm for image denoising, and detects text regions through convolutional neural network;

[0029] Gaussian filtering algorithm is used to process printed ancient book images, smooth the image and preserve the character edges, median filtering algorithm is used to remove impulse noise for handwritten images with stains and blurred handwriting, adaptive threshold method is used to convert color or grayscale images into black and white binary images: text area is set to black with pixel value 0, background is set to white with pixel value 255, enhancing the contrast between text and background, for images with uneven lighting, local threshold adjustment is used to ensure that the text in different areas is clear and visible, a target detection model based on convolutional neural network (CNN) is used to identify the text regions in the image, the text regions are segmented into lines and characters, each character is extracted as a sub-image, and the segmented character sub-images are subjected to contrast stretching and edge enhancement to improve the outline clarity of blurred characters.

[0030] Ancient Zhuang character mapping dictionary construction module: used to store the correspondence between ancient Zhuang character images and phonetic symbols, to mark the pronunciation of polyphonic characters in different contexts, and to support dynamic updates;

[0031] We collected over 30,000 images of standard printed ancient Zhuang characters, which were then annotated with the International Phonetic Alphabet by Zhuang language experts. For polyphonic characters, such as “” which is pronounced “fai” in place names… 5 It is pronounced "fai" in a person's name. 2 The system labels contextual tags and their corresponding pronunciations to form an initial mapping table. Based on the ancient Zhuang language corpus, it trains the BERT model to learn the semantics of character context and establishes a character, context, and pronunciation association model. When the OCR recognizes a polyphonic character, the model automatically extracts the five characters before and after it as contextual features and matches the most likely pronunciation in the dictionary. It supports manual updates, allowing users to upload images and phonetic symbols of newly recognized rare characters. After review by language experts, these images and phonetic symbols are included in the dictionary. It also supports crowdsourcing updates, collecting pronunciation suggestions for unrecognized characters through user feedback and using a voting mechanism to filter high-credibility content. The dictionary is updated in batches regularly.

[0032] Optical Character Recognition (OCR) Module: The OCR module uses a model based on a combination of CNN and LSTM to recognize ancient Zhuang characters in the preprocessed image. By matching the similarity of the characters with the characters in the ancient Zhuang character mapping dictionary, it outputs standardized text and corresponding phonetic symbols. The OCR module uses algorithms such as cosine similarity and Euclidean distance to perform character similarity matching.

[0033] This system employs a deep learning model to recognize character sub-images, outputting standard text and corresponding phonetic symbols. It addresses the challenges of recognizing complex characters, variant characters, and handwritten characters. A hybrid model of "CNN+LSTM+CTC" is constructed. The CNN layer extracts low-level features such as edges, textures, and structures from the character sub-images. The LSTM layer handles the temporal dependencies of character sequences and outputs character probability distributions. The CTC layer solves the alignment problem between character length and the output sequence, improving the accuracy of handwritten character recognition. The training dataset contains over 500,000 images of ancient Zhuang characters, expanded to over 2 million samples through data augmentation. The model achieves an accuracy rate of over 95%. For the OCR model's recognition results, the similarity between the OCR model and the characters in the dictionary is calculated. If the candidate character with the highest similarity has a confidence level ≥90%, its corresponding phonetic symbol is automatically matched; otherwise, it is marked as pending verification, requiring manual correction by the user through the interface. The recognized character sequences are converted into standard ancient Zhuang text, and a mapping dictionary is called to match phonetic symbols for each character, generating a text-phonetic symbol lookup table.

[0034] Text-to-Speech (TTS) module: Based on the phonetic symbols output by the OCR module, the module models and processes the intonation characteristics of Zhuang language to generate speech. The TTS module analyzes the context through the Transformer self-attention mechanism to achieve automatic disambiguation of polyphonic characters. It uses deep neural networks such as Tacotron2 and WaveNet to convert phonetic symbols into audio waveforms.

[0035] Based on phonetic symbols and context information, generate speech that conforms to Zhuang language intonation, natural and fluent, solve tone accuracy and disambiguation of polyphones, receive text and phonetic symbol table output by OCR module, check phonetic symbol format, complete incorrect phonetic symbols, tone modeling: use Transformer-based tone prediction model to learn Zhuang language "high-low tone" and "long-short tone" rules, generate tone curve, prosody processing: predict pause duration based on sentence structure and mark stress position, polyphone disambiguation: analyze sentence semantics through self-attention mechanism (Transformer encoder) to verify the phonetic symbols output by the OCR module, use Tacotron2 model to convert phonetic symbols, tone curve and prosody features into mel spectrum, and then use WaveNet model to generate original audio waveform, audio post-processing: use Wiener filter to remove synthesis noise, balance volume through loudness normalization, and add fade-out effect;

[0036] System integration and control module: used to control the cooperative work of the above modules, realize the automatic processing from the input of ancient book images to the output of final audio, and optimize the generated audio by denoising and volume balancing, user interface and application layer, support batch uploading of ancient book images, audio playback, audio download, keyword search, Zhuang language and Mandarin or English translation and bookmark marking functions;

[0037] Users upload ancient book images, the system automatically triggers the processing flow, the image is preprocessed, and the output character sub-image sequence is obtained, the OCR module recognizes characters and matches phonetic symbols, generates text and phonetic symbol data, the TTS module generates speech waveform based on phonetic symbols, and outputs audio files after post-processing, the system stores text, phonetic symbols and audio in association, and users can view text, play audio or download files through the interface;

[0038] Support batch uploading, the upload interface displays a progress bar and preprocessing status, automatically detect image quality, prompt users to re-upload if the blurriness is too high, or automatically enable enhancement mode, text area is typeset according to ancient book paragraphs, each paragraph displays a "play" button below, click to play the corresponding audio, when playing, the characters currently pronounced are highlighted in yellow, support progress bar dragging, support single audio download or whole book package download, the download interface displays file size and estimated download time, provide "recognition error" and "pronunciation error" feedback entry, users can mark error positions and input correct content, the system records feedback data for model optimization.

[0039] The above describes the present application and its embodiments, which are not limited, and the drawings only show one of the embodiments of the present application, and the actual structure is not limited thereto. In general, if a person skilled in the art is inspired thereby, without departing from the purpose of the present application, without creative design, similar structure and embodiments of the technical solution are not creative, and should belong to the protection scope of the present application.

Claims

1. A method for digitizing Zhuang ancient texts, characterized in that, Includes the following steps: Image processing: Preprocessing of the input Zhuang ancient book image, including denoising, binarization, layout analysis, character segmentation, and image quality optimization; OCR Recognition and Phonetic Symbol Mapping: A model based on a combination of CNN and LSTM is used to recognize ancient Zhuang characters in preprocessed images. By matching the similarity of the characters with the characters in the ancient Zhuang character mapping lexicon, standardized text and corresponding phonetic symbols are output. For polyphonic characters, the correct pronunciation is selected in combination with the context. Speech synthesis: Based on the output phonetic symbols, the speech is generated by modeling and processing the intonation characteristics of Zhuang language. Audio post-processing: Denoise and volume balance optimization are performed on the generated speech to output the final audio.

2. The method and system for digitizing Zhuang ancient texts according to claim 1, characterized in that: In the image processing steps, Gaussian filtering or median filtering algorithms are used for image denoising, and text regions are detected by convolutional neural networks to segment text and non-text regions.

3. The method and system for digitizing Zhuang ancient texts according to claim 1, characterized in that: The ancient Zhuang character mapping dictionary is constructed by collecting standard printed ancient Zhuang character images, marking each character with phonetic symbols, and marking the pronunciation of polyphonic characters in different contexts. It also supports dynamic updates through manual or crowdsourcing methods.

4. The method and system for digitizing Zhuang ancient texts according to claim 1, characterized in that: In the speech synthesis step, the Transformer self-attention mechanism is used to analyze the context to achieve automatic disambiguation of polyphonic characters, and deep neural networks such as Tacotron2 and WaveNet are used to convert phonetic symbols into audio waveforms.

5. A system for digitizing ancient Zhuang texts into audio format, characterized in that, include: Ancient Book Image Processing Module: Used for preprocessing input Zhuang ancient book images, including denoising, binarization, layout analysis, character segmentation, and image quality optimization; Ancient Zhuang character mapping dictionary construction module: used to store the correspondence between ancient Zhuang character images and phonetic symbols, to mark the pronunciation of polyphonic characters in different contexts, and to support dynamic updates; Optical Character Recognition (OCR) module: It uses a model based on a combination of CNN and LSTM to recognize ancient Zhuang characters in the preprocessed image. By matching the similarity of the characters with the characters in the ancient Zhuang character mapping dictionary, it outputs standardized text and corresponding phonetic symbols. Text-to-Speech (TTS) module: Based on the phonetic symbols output by the OCR module, the module models and processes the intonation characteristics of Zhuang language to generate speech; System integration and control module: Used to control the coordinated work of the above modules, realize the automated processing from ancient book image input to final audio output, and optimize the generated audio by noise reduction and volume balance.

6. The method and system for digitizing Zhuang ancient texts according to claim 5, characterized in that: The ancient book image processing module uses Gaussian filtering or median filtering algorithms for image denoising and detects text regions through convolutional neural networks.

7. The method and system for digitizing Zhuang ancient texts according to claim 5, characterized in that: The OCR module uses algorithms such as cosine similarity and Euclidean distance to perform character similarity matching.

8. The method and system for digitizing Zhuang ancient texts according to claim 5, characterized in that: The TTS module analyzes the context through the Transformer self-attention mechanism to achieve automatic disambiguation of polyphonic characters, and uses deep neural networks such as Tacotron2 and WaveNet to convert phonetic symbols into audio waveforms.

9. The method and system for digitizing Zhuang ancient texts according to claim 5, characterized in that: It also includes a user interface and application layer, supporting batch uploading of ancient book images, audio playback, audio download, keyword search, Zhuang language to Mandarin or English translation, and bookmarking functions.