Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Transliteration" patented technology

Transliteration is a type of conversion of a text from one script to another that involves swapping letters (thus trans- + liter-) in predictable ways (such as α → a, д → d, χ → ch, ն → n or æ → ae). For instance, for the Modern Greek term "Ελληνική Δημοκρατία", which is usually translated as "Hellenic Republic", the usual transliteration to Latin script is "Ellēnikḗ Dēmokratía", and the name for Russia in Cyrillic script, "Россия", is usually transliterated as "Rossiya".

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

Internet of Things multi-source heterogeneous data management method based on artificial intelligence

InactiveCN121189325ASemantic analysisBiological modelsTransliterationEngineering
The invention discloses an Internet of Things multi-source heterogeneous data management method based on artificial intelligence, and the method comprises the following steps: collecting voice data and text data in an Internet of Things system, and constructing a voice and text pair; inputting the voice data into an improved Whisper model, and outputting an enhanced transliteration text; inputting the enhanced transliteration text and the text data into a semantic encoder based on dynamic weighting to generate a fused semantic embedding vector; calculating semantic similarity between the fused semantic embedding vectors, and screening semantic consistent sample pairs; performing event aggregation on the semantically consistent sample pairs; performing semantic level coding and confidence adaptive fusion processing to generate a unified semantic fusion vector; and constructing a backmarking training set based on expert feedback, and executing periodic increment optimization on the updatable parameters. The voice and text multi-modal data fusion quality and semantic consistency recognition precision are improved, and the method has good expandability and is suitable for intelligent data governance tasks in a complex Internet of Things environment.
Owner:CHONGQING PAILING INFORMATION TECHNOLOGY CO LTD

Voice transfer text error correction method and device, storage medium and computer equipment

PendingCN121706770ASemantic analysisSpeech recognitionAlgorithmTransliteration
According to the voice transliteration text error correction method and device, the storage medium and the computer equipment, after the voice transliteration text is received in real time, the text is subjected to primary error detection, and the first error confidence coefficient is obtained; when the first error confidence coefficient is smaller than a fast error correction threshold value, the voice transcription text is directly output, and unnecessary error correction is avoided; otherwise, performing word-level rapid error correction on the voice transcription text to obtain a rapid error correction text, and performing secondary error detection on the rapid error correction text to obtain a second error confidence coefficient. Judging whether the second error confidence is smaller than a depth error correction threshold value or not; if yes, the rapid error correction text is directly output, and if not, semantic-level deep error correction is conducted on the rapid error correction text through a deep error correction model, and a deep error correction text is obtained and then output. Through a hierarchical error correction processing strategy, the method can be adapted to different service scenes, and meanwhile, dynamic balance between quality and efficiency of different service scenes can be realized by adjusting two error correction thresholds.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Method and device for processing sensitive information in video, electronic equipment and storage medium

The invention provides a method and device for processing sensitive information in a video, electronic equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: extracting a voice segment in an audio signal, and determining a transliteration text of the voice segment and a phoneme timestamp of the transliteration text; performing semantic analysis on the transliteration text to obtain first sensitive information in the transliteration text, determining a sensitive voice segment corresponding to the first sensitive information based on the phoneme timestamp, and erasing the sensitive voice segment; and finally, sensitive area identification is carried out on each video frame to obtain second sensitive information in each video frame, and the second sensitive information is erased. According to the method, the sensitive information in the video can be automatically and accurately positioned and efficiently purified, the processed video is ensured to meet the requirements of national laws and regulations, the availability of the video in tasks such as action recognition and cross-modal learning can be remarkably improved, and a pure data basis is provided for high-quality visual model training.
Owner:ANHUI FEISHU INFORMATION TECHNOLOGY CO LTD

Structured electronic medical record generation method based on speech recognition and large language model

The invention belongs to the technical field of artificial intelligence and medical informatization, particularly relates to a structured electronic medical record generation method based on voice recognition and a large language model, and is particularly suitable for voice-driven electronic medical record automatic generation and structured extraction tasks. The method comprises the following steps: acquiring doctor-patient dialogue audio, and performing multilevel acoustic modeling and semantic coding by using a voice recognition model to obtain a standardized transliteration text; and then, combining the constructed prompt template with historical context information, guiding a large language model to automatically generate a plurality of structured electronic medical record fields, and assisting in outputting diagnosis suggestions. And finally, completing the optimization and confirmation of the structured result through a doctor interaction confirmation mechanism, and generating standardized output meeting the requirements of the electronic medical record system. According to the method, the medical record generation efficiency and the structural quality can be effectively improved, the document burden of doctors is reduced, and good clinical practicability and popularization value are achieved.
Owner:SHANXI ZHIJIE COMPUTER SOFTWARE ENG CO LTD

Conba Tibetan speech synthesis front-end modeling method, system, device, medium and program product

PendingCN120877709ASpeech synthesisTransliterationAcoustics
The invention provides a Kangba Tibetan speech synthesis front-end modeling method, which can be applied to the technical field of artificial intelligence. The front-end modeling method for voice synthesis of the Conba Tibetan language comprises the following steps: acquiring a Conba Tibetan language text and a corresponding real Tibetan language rhythm tag; based on a preset conbar international phonetic symbol transliteration rule, transliteration is conducted on the conbar Tibetan text to determine a conbar Tibetan phoneme vector; based on the three-layer structure modeling rhythm information of the Kangba Tibetan phoneme vector and the hierarchy-sensitive local self-attention mechanism, generating a prediction rhythm tag; and optimizing a rhythm modeling module according to the real Tibetan rhythm tag and the predicted rhythm tag to obtain a target rhythm encoder. The invention also provides a front-end modeling system and equipment for voice synthesis of the conoba Tibetan language, a storage medium and a program product.
Owner:TIANJIN UNIV

Voice transcription processing method, device, equipment and program product

PendingCN121884821ASemantic analysisSpeech recognitionLexical itemTransliteration
The invention discloses a voice transcription processing method and device, equipment and a program product. The method comprises the steps of obtaining audio data; and performing voice transcription processing on the audio data to obtain candidate transcription texts of at least two versions. Comparing the at least two versions of candidate transcription texts, identifying suspicious text segments with transcription ambiguity in the at least two versions of candidate transcription texts, and determining a text position corresponding to each suspicious text segment as a position to be filled. And for each position to be filled with the blank, extracting corresponding different text contents from all the candidate transcription texts to serve as candidate lexical items of the position to be filled with the blank. And calling a language model, and selecting matched candidate lexical items for each to-be-filled position based on context semantics of all the candidate transliteration texts to perform text completion so as to obtain a target transliteration text corresponding to the audio data.
Owner:PICC INFORMATION TECH CO LTD +1

Systems and methods for interpreting and transliterating the lyrics of a song in foreign languages

The present disclosure provides a non-transitory computer readable medium having instructions stored thereon that, when executed by a processing device, cause the processing device to carry out an operation comprising: ingesting metadata associated with a first song including a plurality of first song lyrics; dividing the plurality of first song lyrics into corresponding lyric blocks; encoding, via a phonetic color-number pairing, each lyric block by determining a first consonantal sound of each lyric, associating the first consonantal sound with a predetermined color-number pairing, and color-coding a plurality of corresponding lyric blocks based on the predetermined color-number pairing; generating structured grids comprised of the corresponding color-coded lyric blocks and displaying the structured grids via client devices.
Owner:MERKUR MICHAEL

Conference speaker conversion point recognition method based on multi-level judgment of text and voice representation fusion

The invention discloses a multi-level judgment conference speaker conversion point recognition method based on text and voice representation fusion, and relates to the technical field of voice signal processing and artificial intelligence. The method comprises the following steps: firstly, preprocessing an original audio stream to extract an effective voice segment, then obtaining a transliteration text and a character timestamp through low-delay stream type voice recognition, generating a word-level timestamp, and segmenting a word-level audio segment; converting points in the fragments are identified, cross-fragment multi-dimensional features are extracted, and affiliation of adjacent fragments is discriminated through fusion of a depth discrimination model; according to the method, text and voice representation are combined, the problems of inaccurate boundary segmentation, voice mixing, sensitivity to acoustic characteristic fluctuation and the like in the prior art are effectively solved through a multi-level judgment mechanism, the accuracy and robustness of speaker conversion point recognition are improved while low delay is guaranteed, and support is provided for high-quality conference record generation.
Owner:田进太

Machine translation method and related device for place names of class characters

The invention discloses a machine translation method and a related device for a class and graphic geographical name, and relates to the field of language translations. The method comprises the following steps: carrying out letter combination division on a class and graphic geographical name phrase to be translated by adopting a forward maximum matching algorithm and a reverse maximum matching algorithm; respectively calculating the mutual information value of the obtained corresponding forward letter combination division result and the mutual information value of the obtained corresponding reverse letter combination division result; and according to a comparison result of the mutual information values, determining an intermediate division result, and performing transliteration on the intermediate division result by using the class and graph Chinese translation table to obtain a transliteration result. According to the method, the to-be-translated phrases are segmented to obtain the intermediate division result which can directly perform accurate proper-name transliteration according to a class and graph Chinese translation table and has high association degree between adjacent phrases, so that the proper-name transliteration problem in class and graph geographical name translation is solved, a machine translation model can independently perform class and graph geographical name translation, and the translation efficiency is improved. The manpower consumption during place name generation is reduced, and the machine translation efficiency is improved.
Owner:CHINESE ACAD OF SURVEYING & MAPPING

Speech recognition method, device and system

The invention discloses a speech recognition method, device and system. Receiving a voice slice sent by the CCU through the websocket connection, and caching the voice slice; in response to the arrival time of a preset duration period, splicing all the cached voice slices into a complete long voice according to a caching sequence; sending the long voice to an online voice recognition server, enabling the online voice recognition server to perform voice recognition on the long voice to obtain a transcription text, and returning the transcription text; receiving a transliteration text returned by the online voice recognition server, sending the transliteration text to a large language model server, enabling the large language model server to judge the transliteration text through a preset large language model, obtaining a corresponding judgment result, and returning the judgment result; and in response to determining that the long voice is abnormal voice according to the judgment result, stopping voice slice splicing, and returning the judgment result to the CCU, so that the CCU performs processing according to a preset abnormality processing strategy. According to the method, voice recognition can be accurately realized in real time.
Owner:SHANGHAI XINFANG INTELLIGENT SYST CO LTD +1

Speech recognition method and device, electronic equipment and storage medium

PendingCN121393448ASpeech recognitionTransliterationEngineering
The invention relates to a voice recognition method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a transliteration text which is obtained through the transliteration of a to-be-processed voice based on a language prediction label; calling a large language model to perform semantic judgment on the transcription text to obtain a first judgment result; and outputting a transcription result of the to-be-processed voice based on a confirmation result of the first judgment result on the language prediction tag. According to the embodiment of the invention, the reliability and stability of speech recognition can be improved.
Owner:MOORE THREADS TECH CO LTD

Script Transliteration Input Pad

ActiveGB6524414S
Owner:BIKASH KUMAR PAL +2

Electronic device and method for creating customized language model

ActiveUS12718798B2Programming languageTransliteration
An example electronic device may include a memory configured to store instructions and a processor electrically connected to the memory and configured to execute the instructions. When the instructions are executed by the processor, the processor may be configured to create an automatic speech recognition (ASR) language model including information about a plurality of candidate transliterations for a variously utterable text, based on a context of a user indicating a situation of the user, a basic language model, or a customized language model and update the customized language model in response to an utterance of the user matching one of the plurality of candidate transliterations.
Owner:SAMSUNG ELECTRONICS CO LTD

Transliterated search method and electronic device supporting same

An electronic device includes a processor and a memory to store a dictionary and index information. The memory can store instructions to cause the processor to: extract a first text in a first language included in data; determine whether or not a transliteration pair is included in the dictionary; if the transliteration pair is identified, index the first text and at least one second text to the data and store same in the index information; and, if the transliteration pair is not identified, extract at least one third text in the second language by inputting the extracted first text to a machine learning model for transliterating an input in the first language into the second language and extracting same, and index the first text and the at least one third text which has been extracted to the data and store same in the index information.
Owner:SAMSUNG ELECTRONICS CO LTD

System and a method for phonetic-based transliteration

A system and a method for converting text in one of a plurality of input languages into a text in a second language using phonetic based transliteration are disclosed. The method includes receiving (802) an input text in a first script from a user; phonetically mapping (804) each character of the input text with a second script corresponding to the second language; validating (806) permutations of mapping of each input character with each character of second script and transliterating (808) input text in first script into an output text in second script. A transliteration engine (106) is configured to transliterate input text of first language into the output text of second language. The transliteration engine (106) includes a data reception module (108), a data transformation module (110), a training module (112), an inference module (114), and a database (116).
Owner:TALENT UNLIMITED ONLINE SERVICES PTE LTD

Higher-order function with reducing function parameter to transliterate characters

PendingUS20260141190A1Natural language translationProgramming languageTransliteration
Systems and methods of computational transliteration of an input sequence of characters in a source language such as Thai to a target language such as Latin. The output may include Romanization of Thai names. The output sequence may be used for machine transliteration understanding of words such as proper nouns. A system may execute a higher-order function that calls a reducing function that iterates through sliding windows of an input sequence and updates, with each iteration, an accumulator map that includes a vector of characters and an indication of a number of characters that can be skipped when processing the next window. Each window includes multiple characters in the input sequence for context-based transliteration using contextual transcription rules. Characters can be skipped when they have already been processed in a previous sliding window. Transliteration of certain source languages may include transposition, deletion, insertion, and transcription.
Owner:MASTERCARD INT INC

Speech recognition illusion suppression method and system based on double-decoding decision network, and medium

The invention provides a voice recognition illusion suppression method and system based on a double-decoding decision network, and a medium. The method comprises the steps of obtaining to-be-recognized voice data; decoding the voice data by adopting a biased decoding model to obtain a customized transliteration result A; decoding the voice data by adopting an unbiased decoding model to obtain a preliminary transcription result B; inputting representation tensors of the customized transliteration result A and the preliminary transliteration result B into a pre-training semantic encoder to obtain advanced representation, and inputting the advanced representation into a trained predictive projection network to obtain illusion probability; comparing the illusion probability with a set discrimination threshold value, and if the illusion probability is smaller than or equal to the discrimination threshold value, outputting a preliminary transcription result B; if so, outputting a customized transfer result A; according to the method, two paths of decoding are performed on the same section of voice in parallel, so that illusion insertion errors caused by prompt words can be effectively eliminated, and the overall recognition accuracy is obviously improved in a scene containing user-customized hot words.
Owner:SHENZHEN TIMEKETTLE TECH CO LTD

Video communication voice transcription method and device adopting artificial intelligence, and electronic equipment

The invention discloses a video communication voice transliteration method and device adopting artificial intelligence and electronic equipment, and relates to the technical field of voice transliteration, and the method comprises the following steps: obtaining a voice signal in a video communication process, and segmenting the voice signal according to a preset frame length to obtain a plurality of voice segments; respectively carrying out feature extraction on the plurality of voice segments, constructing a fuzzy feature vector and identifying a fuzzy pronunciation section; constructing a voice membership curve by adopting a triangular membership function, and evaluating the fuzzy feature vector change before and after the inflection point of the voice membership curve; correcting the speech membership curve slope corresponding to the fuzzy pronunciation section based on the evaluation result to obtain a corrected fuzzy interval; according to the method, the dynamically generated phoneme candidate paths are screened, and the character transliteration result is generated based on the screened optimal phoneme path, so that the problem that fuzzy sounds are forcibly distributed to non-main semantic paths due to overlarge curve inflection point offset when the semantic variation is excessively amplified is solved.
Owner:INFORMATION CENT OF YELLOW RIVER WATER RESOURCES COMMISSION +2

Language-agnostic multilingual modeling using effective script normalization

A method includes obtaining a plurality of training data sets each associated with a respective native language and includes a plurality of respective training data samples. For each respective training data sample of each training data set in the respective native language, the method includes transliterating the corresponding transcription in the respective native script into corresponding transliterated text representing the respective native language of the corresponding audio in a target script and associating the corresponding transliterated text in the target script with the corresponding audio in the respective native language to generate a respective normalized training data sample. The method also includes training, using the normalized training data samples, a multilingual end-to-end speech recognition model to predict speech recognition results in the target script for corresponding speech utterances spoken in any of the different native languages associated with the plurality of training data sets.
Owner:GOOGLE LLC

A method and related apparatus for machine translation of bantu place names

ActiveCN121615663Bimprove relevanceSolve the problem of proper name transliterationNatural language translationMatch algorithmsTheoretical computer science
The application discloses a Bantu language geographical name machine translation method and related device, and relates to the field of language translation. The method comprises the following steps: adopting a forward maximum matching algorithm and a reverse maximum matching algorithm to divide a Bantu language geographical name word group to be translated into letter combinations, and respectively calculating mutual information values of the obtained corresponding forward letter combination division result and reverse letter combination division result; determining an intermediate division result according to the comparison result of the mutual information values, and using a Bantu language Chinese translation and writing table to transliterate the intermediate division result to obtain a transliteration result. The application cuts the word group to be translated, obtains the intermediate division result which can be directly compared with the Bantu language Chinese translation and writing table for accurate proper name transliteration, and has high correlation between adjacent word groups, thereby solving the problem of proper name transliteration in Bantu language geographical name translation, enabling the machine translation model to independently perform Bantu language geographical name translation, reducing the human consumption during geographical name generation, and improving the machine translation efficiency.
Owner:CHINESE ACAD OF SURVEYING & MAPPING

Speech recognition model training method and device based on transliteration translation preference alignment

ActiveCN121148372BImproving proper name fidelityimprove consistencySpeech recognitionTransliterationSpeech sound
The embodiment of the specification provides a speech recognition model training method and device based on phonetic translation and meaning translation preference alignment, wherein the speech recognition model training method based on phonetic translation and meaning translation preference alignment comprises: processing sample audio according to a phonetic translation strategy and a meaning translation strategy to obtain phonetic translation text and meaning translation text; constructing a target sample according to the sample audio, the phonetic translation text, the meaning translation text, and preference annotation information corresponding to the sample audio; processing the target sample by using a speech recognition model to obtain predicted text; performing multi-dimensional evaluation on the predicted text based on a reward model of the speech recognition model to obtain evaluation information; determining a reward loss value and a preference loss value according to the evaluation information; and fine-tuning the speech recognition model based on the reward loss value and the preference loss value to obtain a target speech recognition model, wherein the target speech recognition model is used for phonetic translation recognition and / or meaning translation recognition on audio.
Owner:SHENZHEN WEIAIZHIYUN TECH CO LTD