Dataset generation method and system for tibetan text

By statistically analyzing the frequency of Tibetan characters and generating Tibetan text images, the problem of uneven data quality caused by the limited availability of Tibetan data was solved. A highly available general Tibetan text dataset was established, which improved the model training effect and promoted the development of the Tibetan language.

CN118096940BActive Publication Date: 2026-05-08HEFEI HIGH DIMENSIONAL DATA TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI HIGH DIMENSIONAL DATA TECH CO LTD
Filing Date
2024-01-04
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Due to the limited availability of Tibetan data, existing methods struggle to establish high-quality Tibetan text databases, resulting in inconsistent data quality and impacting the training performance of Tibetan models.

Method used

By statistically analyzing the frequency of Tibetan characters, high-frequency Tibetan main characters and auxiliary characters are selected. Combined with Tibetan grammar rules and information such as background images, text colors, and font sizes, high-quality and diverse Tibetan text images are generated, and a general Tibetan text dataset is established.

Benefits of technology

Without relying on external Tibetan text data, high-quality, diverse, and abundant Tibetan text data were generated, which improved the training effect of the Tibetan object detection model, promoted the development and dissemination of the Tibetan language, and avoided information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118096940B_ABST
    Figure CN118096940B_ABST
Patent Text Reader

Abstract

The application relates to a Tibetan text data set generation method and system, applied to the technical field of data generation, which comprises the following steps: acquiring high-frequency Tibetan main characters and Tibetan auxiliary characters based on preset Tibetan data statistics of Tibetan character appearance frequency; preprocessing the Tibetan data to acquire Tibetan processing information, wherein the Tibetan processing information at least comprises a Tibetan background picture, text color and text font size; and generating a Tibetan text picture according to a preset Tibetan distribution mode, the Tibetan auxiliary characters, the high-frequency Tibetan main characters and the Tibetan processing information. The application guarantees that high-quality, variant diversified and sufficient Tibetan text data can be generated without external Tibetan language data, thereby establishing a high-availability general Tibetan text data set, and further improving the training effect of a Tibetan target detection model, so as to meet the needs of various Tibetan application fields and promote the development and popularization of the Tibetan language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data generation technology, and in particular to a method and system for generating Tibetan text datasets. Background Technology

[0002] Tibetan is a language belonging to the Sino-Tibetan language family, possessing a unique character set and grammatical structure that distinguishes it significantly from many other languages. Existing methods for generating Tibetan text are primarily based on Natural Language Processing (NLP) to understand and generate Tibetan text.

[0003] Qinghai Nationalities University, in its patent application "A Method and System for Automatically Generating Tibetan Webpage Summaries" (Application No. CN202011433753.3, Publication No. CN112328946A), provides a method for automatically generating Tibetan webpage summaries. The implementation steps of this method include: First, using a Tibetan webpage crawler tool to crawl and obtain training and test samples for the Tibetan webpage summarization system; second, determining the length of the Tibetan webpage and whether its hyperlinks exist in the database; third, removing noise from the crawled Tibetan webpage to generate Tibetan webpage text, and then automatically segmenting the text; fourth, after sorting the Tibetan webpage text sentences by weight, setting a threshold for extracting the Tibetan webpage summary, and extracting the initial summary of the Tibetan webpage based on the threshold. This invention, by combining web crawling and natural language processing technology, can effectively output Tibetan webpage summaries.

[0004] Regarding the aforementioned technologies, it is believed that due to the limited availability of Tibetan data, establishing a high-quality Tibetan text database is impractical. Therefore, the methods described above are difficult to use to create a general Tibetan text dataset, resulting in inconsistent data quality and affecting the training of Tibetan models. Summary of the Invention

[0005] To address the issue that establishing a high-quality Tibetan text database is impractical due to the limited availability of Tibetan data, the aforementioned methods struggle to create a universal Tibetan text dataset, resulting in inconsistent data quality and hindering the training of Tibetan target detection models. This application provides a method and system for generating Tibetan text datasets.

[0006] Firstly, this application provides a method for generating a dataset of Tibetan text, employing the following technical solution: including:

[0007] Based on the preset Tibetan data, the frequency of Tibetan characters is statistically analyzed to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters;

[0008] The Tibetan data is preprocessed to obtain Tibetan processing information, which includes at least: Tibetan background image, text color, and text font size;

[0009] A Tibetan text image is generated based on the preset Tibetan distribution pattern, the Tibetan auxiliary characters, the high-frequency Tibetan main characters, and the Tibetan processing information.

[0010] Optionally, the step of statistically analyzing the frequency of Tibetan characters based on preset Tibetan data to obtain high-frequency Tibetan main characters includes:

[0011] All Tibetan text data were statistically analyzed, and the main Tibetan characters and auxiliary Tibetan characters were extracted to generate a Tibetan corpus.

[0012] Generate a frequency table of Tibetan main characters based on the Tibetan main characters in the Tibetan corpus;

[0013] According to the frequency table of Tibetan main characters, a preset number of Tibetan main characters are selected in descending order of frequency as the high-frequency Tibetan main characters.

[0014] Optionally, the preprocessing of the Tibetan data to obtain Tibetan processing information includes at least: a Tibetan background image, text color, and text font size, including:

[0015] Background images are extracted from a preset Tibetan background image library, and the background images are cropped and resized to obtain the Tibetan background image.

[0016] Set the character colors and font sizes of Tibetan text in a preset Tibetan color library to obtain the text color and the text size;

[0017] By mapping the Tibetan background image to the text color and the text font size, a text generation scheme is obtained.

[0018] Optionally, after mapping the Tibetan background image to the text color and the text font size to obtain the text generation scheme, the method further includes:

[0019] By associating multiple text font sizes and multiple text colors with the same Tibetan background image, a text enhancement scheme is generated.

[0020] Optionally, generating a Tibetan text image based on a preset Tibetan distribution pattern, the Tibetan auxiliary characters, the high-frequency Tibetan main characters, and the Tibetan processing information includes:

[0021] The first coordinate in the Tibetan background image is randomly selected and set as the position of the first high-frequency Tibetan main character generated.

[0022] Based on Tibetan grammar rules, the auxiliary Tibetan characters and the high-frequency Tibetan main characters are randomly selected from the Tibetan corpus to generate complete Tibetan characters;

[0023] The embedding position of the complete Tibetan character is determined according to the preset matrix and the first coordinate.

[0024] The complete Tibetan characters are embedded into the Tibetan background image according to the embedding position and the text generation scheme to generate the Tibetan text image.

[0025] Optionally, before embedding the complete Tibetan character into the Tibetan background image according to the embedding position, the method further includes:

[0026] The Tibetan text format is controlled according to the preset row and column spacing;

[0027] The edge spacing of Tibetan characters is set according to a preset standard spacing, and the number of Tibetan characters embedded on the Tibetan background image is controlled.

[0028] Optionally, generating a Tibetan text image based on a preset Tibetan distribution pattern, the Tibetan auxiliary characters, the high-frequency Tibetan main characters, and the Tibetan processing information includes:

[0029] Based on the Tibetan background image, several non-overlapping regions are randomly selected as the embedding coordinates of the high-frequency Tibetan main characters;

[0030] Based on Tibetan grammar rules, the auxiliary Tibetan characters and the high-frequency Tibetan main characters are randomly selected from the Tibetan corpus to generate complete Tibetan characters;

[0031] The complete Tibetan characters are embedded into the Tibetan background image based on the embedding coordinates and the text generation scheme to generate a Tibetan text image.

[0032] Secondly, this application provides a Tibetan text dataset generation device, which employs the following technical solution, including:

[0033] The frequency statistics module is used to count the frequency of Tibetan characters based on preset Tibetan data, and to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters;

[0034] The Tibetan processing module is used to preprocess the Tibetan data and obtain Tibetan processing information, which includes at least: Tibetan background image, text color and text font size;

[0035] The image generation module is used to generate Tibetan text images based on a preset Tibetan distribution pattern, the Tibetan auxiliary characters, the high-frequency Tibetan main characters, and the Tibetan processing information.

[0036] Thirdly, this application also provides a control device, the device comprising:

[0037] It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described above for generating a dataset of Tibetan text.

[0038] Fourthly, this application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in the method for generating a dataset of Tibetan text.

[0039] In summary, the system in this application performs frequency statistics on Tibetan characters, filters out the most frequent high-frequency Tibetan main characters, and then randomly extracts Tibetan background images, text colors, and text font sizes to form text generation and text enhancement schemes. The system then confirms the order and coordinates of the Tibetan characters in the background image using a matrix distribution or random distribution method. Next, the system randomly combines auxiliary Tibetan characters and high-frequency Tibetan main characters to form complete Tibetan characters. Finally, the complete Tibetan characters are embedded into the Tibetan background image according to the text generation or text enhancement scheme to generate Tibetan text images. This ensures the generation of high-quality, diverse, and abundant Tibetan text data without requiring external Tibetan language data, thereby establishing a highly usable general-purpose Tibetan text dataset. This improves the training effect of Tibetan object detection models, meets the needs of various Tibetan application fields, and promotes the development and dissemination of the Tibetan language. Furthermore, since the generation of the Tibetan text dataset is completely independent of external Tibetan data, potential information leakage problems are avoided. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a method for generating a dataset of Tibetan text.

[0041] Figure 2 These are images of Tibetan text arranged in a random distribution.

[0042] Figure 3 This is a structural block diagram of a device for generating datasets of Tibetan text.

[0043] Figure labeling: 210, Frequency statistics module; 220, Tibetan language processing module; 230, Image generation module. Detailed Implementation

[0044] The following combination Figure 1 - Figure 3 This application will be described in further detail.

[0045] Tibetan script can be composed of base characters, vowels, palatalization symbols, superscript characters, subscript characters, prefix characters, and suffix characters. Considering the generation of Tibetan text in real-world scenarios, the Tibetan character variant library includes base characters, base characters + vowels, base characters + subscript characters, base characters + subscript characters + vowels, base characters + palatalization symbols, base characters + palatalization symbols + vowels, and base characters + subscript characters + palatalization symbols + vowels, which is different from mainstream languages.

[0046] Current natural language processing models, such as the GPT series, perform well on some mainstream languages, but may perform poorly on a few languages ​​like Tibetan because they are typically pre-trained on large-scale languages. However, due to the limited availability of Tibetan data, building a high-quality Tibetan text database is impractical. Furthermore, existing methods cannot adequately support the diverse dialects and variations of Tibetan, as Tibetan fonts require strict grammatical or cultural norms to ensure text quality.

[0047] Based on this, this application ensures the generation of high-quality, diverse, and abundant Tibetan text data without the need for external Tibetan language data, thereby establishing a highly usable general Tibetan text dataset, improving the training effect of Tibetan object detection models, meeting the needs of various Tibetan application fields, and promoting the development and dissemination of the Tibetan language.

[0048] Reference Figure 1 The embodiments of this application include at least steps S10 to S30.

[0049] S10: Based on preset Tibetan data, statistically analyze the frequency of Tibetan characters to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters.

[0050] S20, preprocess the Tibetan data to obtain Tibetan processing information.

[0051] The Tibetan processing information includes at least the following: Tibetan background image, text color, and text font size.

[0052] S30 generates Tibetan text images based on preset Tibetan distribution patterns, Tibetan auxiliary characters, high-frequency Tibetan main characters, and Tibetan processing information.

[0053] Among them, the Tibetan script distribution pattern refers to the distribution format of Tibetan characters in the background image, which can be understood as the text format of Tibetan. Tibetan grammar rules refer to the composition methods of the main Tibetan characters and auxiliary Tibetan characters in the Tibetan font.

[0054] Specifically, the system performs character frequency statistics on existing Tibetan data, selecting the most frequently occurring Tibetan main characters, i.e., high-frequency Tibetan main characters. Then, based on Tibetan grammar rules, the system integrates Tibetan auxiliary characters with high-frequency Tibetan main characters to generate complete Tibetan characters. The system then preprocesses the Tibetan data to obtain information such as Tibetan background images, text colors, and text font sizes. Finally, based on Tibetan distribution patterns, complete Tibetan characters, and Tibetan processing information, the system generates Tibetan text images, thereby establishing a highly usable general Tibetan text dataset, improving the training effect of Tibetan object detection models, meeting the needs of various Tibetan application fields, and promoting the development and dissemination of the Tibetan language.

[0055] In some embodiments, step S10 specifically includes the following steps: statistically analyzing all Tibetan text data, extracting Tibetan main characters and Tibetan auxiliary characters, and generating a Tibetan corpus; generating a Tibetan main character frequency table based on the Tibetan main characters in the Tibetan corpus; and sequentially selecting a preset number of Tibetan main characters as high-frequency Tibetan main characters according to the order of frequency from large to small in the Tibetan main character frequency table.

[0056] Specifically, the system extracts the main Tibetan characters and auxiliary Tibetan characters that appear in Tibetan text, and then filters the most frequently occurring main Tibetan characters based on their frequency of occurrence, i.e., high-frequency main Tibetan characters. This makes it easier to collect commonly used Tibetan characters, and thus makes the generated Tibetan dataset more realistic, so that the subsequent model training will be more realistic.

[0057] In some embodiments, step S20 specifically includes the following steps: extracting a background image from a preset Tibetan background library, cropping and resizing the background image to obtain a Tibetan background image; setting the character color and font size of the Tibetan text in a preset Tibetan color library to obtain the text color and text size; and mapping the Tibetan background image to the text color and text size to obtain a text generation scheme.

[0058] Specifically, the system randomly extracts background images from the Tibetan background library, processes the background images to obtain Tibetan background images, sets the character colors and font sizes of Tibetan text according to the Tibetan color library, and finally selects a Tibetan background image, a text color, and a text size to generate a text generation scheme, so that the Tibetan text can be embedded into the Tibetan background image according to the text color and text size in the subsequent process.

[0059] Furthermore, the system maps multiple text font sizes and colors to the same Tibetan background image to generate a text enhancement scheme. This facilitates embedding Tibetan fonts with multiple font sizes and colors into a single Tibetan background image, thereby enhancing the noise reduction of Tibetan text and making the Tibetan object detection model more robust during training.

[0060] It should be understood that when using text enhancement schemes to embed Tibetan characters into Tibetan background images, the text size and color can be determined randomly or preset.

[0061] In some embodiments, step S30 specifically includes the following steps: randomly selecting the first coordinate in the Tibetan background image as the position of the first high-frequency Tibetan main character generated; randomly selecting Tibetan auxiliary characters and high-frequency Tibetan main characters from the Tibetan corpus based on Tibetan grammar rules to generate complete Tibetan characters; determining the embedding position of the complete Tibetan characters according to the preset matrix and the first coordinate; embedding the complete Tibetan characters into the Tibetan background image according to the embedding position and the text generation scheme to generate a Tibetan text image.

[0062] Specifically, the system randomly selects a coordinate in the Tibetan background image as the position of the first high-frequency Tibetan main character to be embedded. Then, based on Tibetan grammar rules, the system randomly selects Tibetan auxiliary characters from the Tibetan corpus and randomly selects high-frequency Tibetan main characters to generate complete Tibetan characters. The system then embeds the complete Tibetan characters into the Tibetan background image according to the row-deterministic arrangement order and the text generation scheme or text enhancement scheme, and finally obtains a Tibetan text image. In this way, high-quality, diverse and sufficient Tibetan text data can be generated without the need for external Tibetan language data, thereby establishing a highly available general Tibetan text dataset.

[0063] Furthermore, the system can control the Tibetan text format according to the preset row and column spacing; and set the edge spacing of Tibetan characters according to the preset standard spacing and control the number of Tibetan characters embedded in the Tibetan background image, so as to reduce the occurrence of blurry Tibetan characters due to too many embedded Tibetan characters in the Tibetan background image.

[0064] In one embodiment, step S30 specifically includes the following steps: randomly selecting several non-overlapping regions based on the Tibetan background image as the embedding coordinates of high-frequency Tibetan main characters; randomly selecting Tibetan auxiliary characters and high-frequency Tibetan main characters from the Tibetan corpus based on Tibetan grammar rules to generate complete Tibetan characters; embedding the complete Tibetan characters into the Tibetan background image according to the embedding coordinates and the text generation scheme to generate a Tibetan text image.

[0065] Reference Figure 2 , Figure 2To generate Tibetan text images according to random assignment, specifically, the system randomly selects several non-overlapping regions in the Tibetan background image as the embedding coordinates for high-frequency Tibetan main characters. The system then randomly selects Tibetan auxiliary characters from the Tibetan corpus based on Tibetan grammar rules and randomly selects high-frequency Tibetan main characters to generate complete Tibetan characters. The system then randomly embeds the complete Tibetan characters into the Tibetan background image according to the embedding coordinates and the text generation or text enhancement scheme, finally obtaining a Tibetan text image. This generates high-quality, diverse, and abundant Tibetan text data without the need for external Tibetan language data, thus establishing a highly usable general Tibetan text dataset.

[0066] It should be understood that this application uses two Tibetan distribution patterns to sort the Tibetan text embedded in the Tibetan background image, but it does not only represent these two sorting methods. Other sorting methods similar to those in this application should also be within the scope of protection of this application.

[0067] The implementation principle of the Tibetan text dataset generation method in this application is as follows: The system performs frequency statistics on Tibetan characters, filters out the most frequent high-frequency Tibetan main characters, and then randomly extracts Tibetan background images, text colors, text font sizes, etc., to form a text generation scheme and a text enhancement scheme. The system then confirms the order and coordinates of Tibetan characters in the Tibetan background image according to a row-determinant distribution or a random distribution method. Next, the system randomly combines Tibetan auxiliary characters and high-frequency Tibetan main characters to form complete Tibetan characters. Finally, the complete Tibetan characters are embedded into the Tibetan background image according to the text generation scheme or text enhancement scheme to generate Tibetan text images. This method can guarantee the generation of high-quality, diverse, and sufficient Tibetan text data without the need for external Tibetan language data, thereby establishing a highly usable general Tibetan text dataset. This improves the training effect of Tibetan object detection models, meets the needs of various Tibetan application fields, and promotes the development and dissemination of the Tibetan language. Furthermore, since the generation of the Tibetan text dataset does not depend on external Tibetan data, it also avoids potential information leakage problems.

[0068] Figure 1 This is a flowchart illustrating a method for generating a dataset of Tibetan text in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows; unless explicitly stated otherwise, there is no strict order requirement for the execution of these steps, and they can be executed in other orders; and Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0069] Based on the same technical concept, referring to Figure 3 This application also provides a Tibetan text dataset generation apparatus, which adopts the following technical solution: the apparatus includes:

[0070] The frequency statistics module 210 is used to count the frequency of Tibetan characters based on preset Tibetan data, and to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters;

[0071] The Tibetan processing module 220 is used to preprocess Tibetan data and obtain Tibetan processing information, which includes at least: Tibetan background image, text color and text font size.

[0072] Image generation module 230 is used to generate Tibetan text images based on preset Tibetan distribution patterns, Tibetan auxiliary characters, high-frequency Tibetan main characters, and Tibetan processing information.

[0073] In some embodiments, the frequency statistics module 210 is specifically used to count all Tibetan text data, extract Tibetan main characters and Tibetan auxiliary characters, and generate a Tibetan corpus;

[0074] Generate a frequency table of Tibetan main characters based on the Tibetan main characters in the Tibetan corpus;

[0075] Based on the frequency table of Tibetan main characters, a preset number of Tibetan main characters are selected in descending order of frequency as high-frequency Tibetan main characters.

[0076] In some embodiments, the Tibetan processing module 220 is specifically used to extract background images from a preset Tibetan background library, and to crop and resize the background images to obtain Tibetan background images.

[0077] Set the character colors and font sizes of Tibetan text in the preset Tibetan color library to obtain the text color and text size;

[0078] By mapping the Tibetan background image to the text color and font size, a text generation scheme is obtained.

[0079] In some embodiments, the Tibetan processing module 220 is further configured to correspond multiple text font sizes and multiple text colors to the same Tibetan background image to generate a text enhancement scheme.

[0080] In some embodiments, the image generation module 230 is specifically used to randomly select the first coordinate in the Tibetan background image and set it as the position of the first high-frequency Tibetan main character generated.

[0081] Complete Tibetan characters are generated by randomly selecting Tibetan auxiliary characters and high-frequency Tibetan main characters from the Tibetan corpus based on Tibetan grammar rules.

[0082] The embedding position of the complete Tibetan character is determined according to the preset matrix and the first coordinate.

[0083] Based on the embedding location and text generation scheme, the complete Tibetan characters are embedded into the Tibetan background image to generate a Tibetan text image.

[0084] In some embodiments, the image generation module 230 is also used to control the Tibetan text format according to a preset row and column spacing;

[0085] Set the edge spacing of Tibetan characters according to the preset standard spacing and control the number of Tibetan characters embedded on the Tibetan background image.

[0086] In some embodiments, the image generation module 230 is specifically used to randomly select several non-overlapping regions based on the Tibetan background image as the embedding coordinates of high-frequency Tibetan main characters;

[0087] Complete Tibetan characters are generated by randomly selecting Tibetan auxiliary characters and high-frequency Tibetan main characters from the Tibetan corpus based on Tibetan grammar rules.

[0088] Based on the embedding coordinates and the text generation scheme, the complete Tibetan characters are embedded into the Tibetan background image to generate a Tibetan text image.

[0089] This application also discloses a control device.

[0090] Specifically, the control device includes a memory and a processor, the memory storing a computer program that can be loaded by the processor and executed to generate the dataset of the Tibetan text described above.

[0091] This application also discloses a computer-readable storage medium.

[0092] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed as described above in the method for generating a dataset of Tibetan text. The computer-readable storage medium includes, for example, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0093] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for generating a dataset of Tibetan text, characterized in that, include: Based on preset Tibetan data, the frequency of Tibetan characters is statistically analyzed to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters, including: All Tibetan text data were statistically analyzed, and the main Tibetan characters and auxiliary Tibetan characters were extracted to generate a Tibetan corpus. Generate a frequency table of Tibetan main characters based on the Tibetan main characters in the Tibetan corpus; According to the frequency table of Tibetan main characters, a preset number of Tibetan main characters are selected in descending order of frequency as the high-frequency Tibetan main characters; The Tibetan data is preprocessed to obtain Tibetan processing information, which includes at least: Tibetan background image, text color, and text font size, including: Background images are extracted from a preset Tibetan background image library, and the background images are cropped and resized to obtain the Tibetan background image. Set the character colors and font sizes of Tibetan text in a preset Tibetan color library to obtain the text color and the text size; By mapping the Tibetan background image to the text color and the text font size, a text generation scheme is obtained; By associating multiple text font sizes and multiple text colors with the same Tibetan background image, a text enhancement scheme is generated. By combining the Tibetan grammar rules, the Tibetan auxiliary characters and the high-frequency Tibetan main characters are combined into complete Tibetan characters; Based on a preset Tibetan distribution pattern, the Tibetan processing information and the complete Tibetan characters are used to generate Tibetan text images to construct a general Tibetan text dataset.

2. The method for generating a dataset of Tibetan text according to claim 1, characterized in that, The step of generating a Tibetan text image based on a preset Tibetan distribution pattern, the Tibetan processing information, and the complete Tibetan characters includes: The first coordinate in the Tibetan background image is randomly selected and set as the position of the first high-frequency Tibetan main character generated. Based on Tibetan grammar rules, the auxiliary Tibetan characters and the high-frequency Tibetan main characters are randomly selected from the Tibetan corpus to generate complete Tibetan characters; The embedding position of the complete Tibetan character is determined according to the preset matrix and the first coordinate. The complete Tibetan characters are embedded into the Tibetan background image according to the embedding position and the text generation scheme to generate the Tibetan text image.

3. The method for generating a dataset of Tibetan text according to claim 2, characterized in that, Before embedding the complete Tibetan character into the Tibetan background image according to the embedding position, the method further includes: The Tibetan text format is controlled according to the preset row and column spacing; The edge spacing of Tibetan characters is set according to a preset standard spacing, and the number of Tibetan characters embedded on the Tibetan background image is controlled.

4. The method for generating a dataset of Tibetan text according to claim 1, characterized in that, The step of generating a Tibetan text image based on a preset Tibetan distribution pattern, the Tibetan processing information, and the complete Tibetan characters includes: Based on the Tibetan background image, several non-overlapping regions are randomly selected as the embedding coordinates of the high-frequency Tibetan main characters; Based on Tibetan grammar rules, the auxiliary Tibetan characters and the high-frequency Tibetan main characters are randomly selected from the Tibetan corpus to generate complete Tibetan characters; The complete Tibetan characters are embedded into the Tibetan background image based on the embedding coordinates and the text generation scheme to generate a Tibetan text image.

5. A dataset generation device for Tibetan text, characterized in that, The device includes: The frequency statistics module is used to count the frequency of Tibetan characters based on preset Tibetan data, and to obtain high-frequency Tibetan main characters and Tibetan auxiliary characters, including: All Tibetan text data were statistically analyzed, and the main Tibetan characters and auxiliary Tibetan characters were extracted to generate a Tibetan corpus. Generate a frequency table of Tibetan main characters based on the Tibetan main characters in the Tibetan corpus; According to the frequency table of Tibetan main characters, a preset number of Tibetan main characters are selected in descending order of frequency as the high-frequency Tibetan main characters; The Tibetan processing module is used to preprocess the Tibetan data to obtain Tibetan processing information, which includes at least: a Tibetan background image, text color, and text font size, including: Background images are extracted from a preset Tibetan background image library, and the background images are cropped and resized to obtain the Tibetan background image. Set the character colors and font sizes of Tibetan text in a preset Tibetan color library to obtain the text color and the text size; By mapping the Tibetan background image to the text color and the text font size, a text generation scheme is obtained; By associating multiple text font sizes and multiple text colors with the same Tibetan background image, a text enhancement scheme is generated. The image generation module, in conjunction with Tibetan grammar rules, combines the Tibetan auxiliary characters and the high-frequency Tibetan main characters into complete Tibetan characters; according to a preset Tibetan distribution pattern, it generates Tibetan text images from the Tibetan processing information and the complete Tibetan characters to construct a general Tibetan text dataset.

6. A control device, characterized in that, The device includes: It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Automatic Tibetan webpage abstract generation method and system

    CN112328946A

  • Tibetan font diversity expression method based on weighted random distribution model

    CN110728273A

  • Method and system for generating enhanced text composite image

    CN114119949A

  • Dunhuang Tibetan ancient book identification method, system and device and medium

    CN116740726A