Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9 results about "Tibetan language" patented technology

Standard Tibetan is the most widely spoken form of the Tibetic languages. It is based on the speech of Lhasa, an Ü-Tsang (Central Tibetan) dialect. For this reason, Standard Tibetan is often called Lhasa Tibetan. Tibetan is an official language of the Tibet Autonomous Region of the People's Republic of China.

Dataset generation method and system for tibetan text

The application relates to a Tibetan text data set generation method and system, applied to the technical field of data generation, which comprises the following steps: acquiring high-frequency Tibetan main characters and Tibetan auxiliary characters based on preset Tibetan data statistics of Tibetan character appearance frequency; preprocessing the Tibetan data to acquire Tibetan processing information, wherein the Tibetan processing information at least comprises a Tibetan background picture, text color and text font size; and generating a Tibetan text picture according to a preset Tibetan distribution mode, the Tibetan auxiliary characters, the high-frequency Tibetan main characters and the Tibetan processing information. The application guarantees that high-quality, variant diversified and sufficient Tibetan text data can be generated without external Tibetan language data, thereby establishing a high-availability general Tibetan text data set, and further improving the training effect of a Tibetan target detection model, so as to meet the needs of various Tibetan application fields and promote the development and popularization of the Tibetan language.
Owner:HEFEI HIGH DIMENSIONAL DATA TECH CO LTD +1

Boundary perception and pre-training enhanced low-resource Tibetan language recognition method and system

The invention discloses a boundary perception and pre-training enhanced low-resource Tibetan language recognition method and system, and the method comprises the following steps: 1, carrying out the standardization processing of data and texts, and enabling the language knowledge and pre-training model enhanced low-resource Tibetan language voice to be converted into text. Unified modeling of Tibetan language boundaries, normal characters and morphemes is realized under a low-resource condition; remarkable and stable alignment is realized through a tsheg-perception adapter and three-head joint learning; language knowledge is fused into search in a learnable mode through the micro WFST soft constraint, and illegal font and language / case marking errors are greatly reduced. Training data are effectively expanded through teacher semi-supervised and coverage rate active learning, and cross-domain and accent robustness is improved; the end side maintenance cost is remarkably reduced through'adapter + atlas' hot updating; iCR and a boundary F1 are introduced as closed-loop indexes, so that the quality of an optimization target is consistent with that of a Tibetan Chinese character method, and the whole method has the advantages in engineering landing, interpretability and maintainability.
Owner:INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES

Efficient automatic construction method for Tibetan-Chinese bilingual parallel corpora

The invention discloses an efficient and automatic construction method for Tibetan and Chinese bilingual parallel corpora, and relates to the technical field of Tibetan and Chinese bilingual parallel corpora construction.The Tibetan and Chinese bilingual parallel corpora construction comprises a base model unit, a localization processing unit and a Tibetan stability strengthening module, the base model unit is a Qwen2.5 series pre-training model as a technical base, and the localization processing unit is a local processing unit; the training model comprises a 70B parameter main generation model and a 7B parameter lightweight intention judgment model, and the 70B model is responsible for receiving an input text and generating Tibetan-Chinese bilingual parallel sentence pairs. A base model unit, a localization processing unit and a Tibetan stability strengthening module communicate with one another through shared memory communication, the Tibetan stability strengthening module further comprises a double-model cooperation mechanism, and the working process of the double-model cooperation mechanism comprises the steps that input Chinese sentences are preprocessed through a word segmentation module, 7B model execution intention triad analysis is carried out, and the Chinese sentences are classified into Chinese sentences; and semantic control codes are generated, so that the output stability of the Tibetan and Chinese bilingual parallel corpora is improved through a double-model collaborative architecture.
Owner:TIBET JUELUO DIGITAL IND MANAGEMENT CO LTD

A semantic enhancement-based ando language generation method and system

PendingCN122337173AMedicineSemantic feature
This application relates to a semantically enhanced Amdo Tibetan language generation method and system. The method extracts semantic features containing lexical boundaries through a pre-trained encoder with full-word masking. After dimensional transformation and phoneme-level repetition expansion, these features are added element-wise to the phoneme features, so that each phoneme directly carries the complete semantics of its corresponding word, fundamentally eliminating semantic fragmentation in prosodic modeling. Simultaneously, the sequence termination vector is independently input into the end-energy predictor and jointly trained with a variational inference adversarial training framework to obtain the generation model and the sentence-end energy decay coefficient. During inference, after synthesizing audio segments according to sentence-end punctuation, the end of the segments is faded out based on the decay coefficient and a silence buffer block is spliced. This ensures the naturalness and accuracy of the speech flow rhythm and logical stress while completely eliminating sentence-end elision and acoustic truncation noise, significantly improving the prosodic naturalness and speech purity of the synthesized Amdo Tibetan speech.
Owner:HAINAN TIBETAN AUTONOMOUS PREFECTURE TIBETAN INFORMATION TECH RES CENT

Tibetan language speech recognition method based on mixed weight and dynamic block training

PendingCN122416991AEncoder decoderTibetan language
This invention relates to the field of speech recognition technology, and particularly to a streaming Tibetan speech recognition method based on hybrid weights and dynamic chunk training. The core architecture is a shared encoder based on an improved Zipformer, with a backend connected to both RNN-T and CTC dual decoding branches. By fusing a pre-trained Tibetan language model within the shared encoder architecture, the method compensates for missing corpora, helping the model learn Tibetan language knowledge and thus optimizing its parameters. Dynamic chunk training achieves a unified streaming and non-streaming capability. The Hybrid RNN-T / CTC speech recognition model architecture consists of four parts: a shared encoder, a CTC decoder, an RNN-T prediction network / joint network, and a language model. The shared encoder is composed of an improved multi-layer Zipformer. This invention fuses a pre-trained Tibetan Transformer language model with the joint network in the Hybrid RNN-T / CTC model and randomly determines the size of audio chunks during training.
Owner:NORTHWEST UNIVERSITY FOR NATIONALITIES

Tibetan language Andor dialect speech synthesis method

The invention provides a Tibetan Anduodialect speech synthesis method and system, and belongs to the technical field of speech synthesis, and the method comprises the steps: obtaining a multi-language text-audio pairing data set, and constructing a text-to-speech model; training a text-to-speech model by using data in the multilingual text-audio pairing data set to obtain a multilingual universal model and initial model parameters; acquiring text-audio pairing data of Tibetan Andole dialects, inputting the audio data and the text data into the multi-language general model, and updating initial model parameters of the multi-language general model by using a full-amount fine tuning algorithm; and obtaining a to-be-recognized Tibetan language Andor dialect text, and inputting the to-be-recognized Tibetan language Andor dialect text into the multi-language general model with the updated parameters to obtain Tibetan language Andor dialect synthetic speech. The synthesized high-quality dialect voice is closer to real voice, and a good voice synthesis effect is provided for a low-resource language scene at low cost.
Owner:QINGHAI NORMAL UNIV

Weather warning language release graphical user interface for electronic devices

1. The name of the design product: weather warning Tibetan language release graphical user interface of electronic equipment. 2. The use of the design product: for displaying graphical user interface. 3. The design points of the design product: the graphical user interface in the display screen panel. 4. The picture or photo that best indicates the design points: front view. 5. The use of the graphical user interface: the interface is used to operate the weather warning Tibetan language release graphical user interface of electronic equipment according to the user. The front view is the interface after login and opening; click "weather service Chinese-Tibetan parallel corpus" in the front view to enter the weather service Chinese-Tibetan parallel corpus interface, as shown in the change state diagram.
Owner:青海省气象服务中心

Tibetan dialect protection and real-time cross-dialect translation system based on voiceprint recognition

The invention discloses a Tibetan dialect protection and real-time cross-dialect translation system based on voiceprint recognition, and relates to the technical field of voice processing, the Tibetan dialect protection and real-time cross-dialect translation system comprises a voiceprint recognition module, a voice recognition module, a cross-dialect translation module, a voice synthesis module and a safety and audit module, the speech recognition module writes an input dialect speech stream into a standard Tibetan text according to a judged dialect category, the cross-dialect translation module receives the Tibetan text and drives a Tibetan large language model to complete semantic conversion and text generation, and the speech synthesis module converts the dialect text into target Tibetan dialect speech. According to the method, voiceprint recognition and dialect classification are combined, so that the function of personalized accurate cross-dialect translation is realized, the problem of difficult communication caused by large difference of Tibetan dialects is effectively solved, users of different dialects can perform smooth and natural real-time dialogues, and language barriers are broken through.
Owner:TIBET JUELUO DIGITAL IND MANAGEMENT CO LTD

A system and method for Tibetan-Chinese bilingual corpus collaborative annotation and versioned release

The application belongs to the technical field of natural language processing, and relates to a Tibetan-Chinese bilingual corpus collaborative labeling and versioned publishing system and method. Through the system architecture formed by the front-end interaction module, the business processing module, the data storage module and the basic management and control module, relying on the Tibetan-Chinese bilingual labeling, intelligent collaborative management and control, corpus version management and standardized publishing core units integrated by the business processing module, cooperating with the distributed correlation index storage mechanism of the data storage module, the fine-grained permission and operation log management and control function of the basic management and control module, the core technical problems of the Tibetan language characteristics adaptation deficiency, the low efficiency and frequent conflicts of multi-person collaborative labeling, the lack of version tracing and quality grading control of the corpus, the non-standard publishing process and the disconnection of the corpus and the downstream model in the prior art are solved. The accuracy of Tibetan-Chinese bilingual corpus labeling and storage is ensured through exclusive Tibetan adaptation processing, and the orderly promotion of multi-person labeling is realized through modular collaborative management and control.
Owner:SICHUAN TIANFU GAOCHI INFORMATION TECHNOLOGY CO LTD