Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Language identification" patented technology

In natural language processing, language identification or language guessing is the problem of determining which natural language given content is in. Computational approaches to this problem view it as a special case of text categorization, solved with various statistical methods.

Artificial intelligence speech recognition system

The invention discloses an artificial intelligence speech recognition system, and the system comprises a multi-modal feature extraction module which employs an improved Conformer architecture to synchronously extract the time-frequency features and text embedding vectors of speech signals; the joint training module is used for performing joint optimization on ASR and NMT loss functions through an adversarial training strategy, learning voice recognition and machine translation tasks at the same time through joint training, and completing direct mapping from voice features to a target language; the context perception translation engine is used for integrating an attention mechanism of a pre-training language model, carrying out deep coding on the extracted speech features and generating cross-language semantic representation; the self-adaptive post-processing module is used for dynamically optimizing an output result by adopting a reinforcement learning framework, dynamically adjusting the output result according to a reward function, and optimizing translation quality and a speech synthesis effect; the dynamic language recognition module is a real-time language classifier based on a Wave2Vec 2.0 framework and is used for recognizing the language of the input voice in real time; and the incremental field adaptation module is used for quickly updating a field term library by using a LoRA fine tuning technology.
Owner:ANKANG UNIV

Code conversion method and device based on domain-specific language and medium

The embodiment of the invention discloses a code conversion method and device based on a domain-specific language and a medium, and relates to the technical field of code conversion, the method comprises the steps that a DSL file and a target conversion language input by a user are received, the DSL file comprises at least one class definition, and the class definition comprises a class name, a field declaration and a method signature declaration; the DSL file is analyzed to extract a plurality of class annotations marked in the class definition, and the class annotations comprise target language identification information, unique class identification information, construction mode identification information and array access mode identification information; and according to the target language identification information in the target conversion language and the class annotations, a matched specified class annotation set is extracted from the class annotations, a code file conforming to the grammar of the target conversion language is generated based on the specified class annotation set, and the specified class annotation set comprises the class annotations, field annotations and method annotations.
Owner:INSPUR GENERSOFT CO LTD

Method and system for automatically arranging process engine of Internet of Things based on large model

The invention discloses an Internet of Things process engine automatic arrangement method and system based on a large model. The method comprises the steps of obtaining a process demand described by a natural language of a user; the method comprises the following steps: analyzing a natural language of a user through a first-layer large model, and identifying related equipment, variables, conditions and action information; information identified by the first-layer large model serves as input of a second-layer large model, and the second-layer large model matches actual corresponding equipment and variable information under the user permission based on the user token, the equipment and the variable information; transmitting actual corresponding equipment, variables, conditions and action information under the user permission to a third-layer large model; and the third-layer large model matches the corresponding flow nodes according to the received information, the matched flow nodes are connected in series through a connecting line, and a structured flow JSON is generated. According to the method, the fuzzy natural language intention is gradually converted into an accurate and executable structured process through the three-layer large model, and the accuracy of an output result of each link is improved.
Owner:SHANDONG YOU INTERNET OF THINGS CO LTD

Language therapy with multilingual ai-agent

A method for treating a language disorder in a patient includes receiving therapist input specifying a speech target and engagement indicator priority, capturing audio of a speech response to a therapy prompt, and identifying the language of the response using a multilingual language identification model. The method further comprises analyzing the speech response with a language-specific recognition model to extract speech features and classify errors across multiple linguistic and acoustic dimensions. Engagement indicators are extracted and used to compute an engagement score, which, along with the error classifications and speech target, informs a decision model that selects a therapy task. The selected task is presented to the patient, and a subsequent speech response is captured to update error classifications. The decision model is iteratively refined based on therapist input and revised error data, enabling adaptive, personalized therapy progression.
Owner:UNIV OF SOUTH FLORIDA

Large model-based secret language collection method, apparatus and device, and storage medium

The invention provides a large-model-based secret language collection method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical fields of artificial intelligence, large models, public opinion control, secret language recognition, risk content management and control and the like. According to the specific implementation scheme, a candidate content set is obtained based on seed expansion related to a secret language; screening candidate contents in the candidate content set by using a large model to obtain candidate secret words; performing retrieval enhancement on the basis of the candidate secret words to obtain a retrieval result; and screening the retrieval results by using a large model to obtain the to-be-collected secret words and explanations thereof.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Humanoid robot language identification model training method and language identification method

The application discloses a kind of humanoid robot language identification model training method and language identification method, wherein, the training method obtains multi-modal training sample, the multi-modal training sample includes audio mode sample, oral type mode sample and environment mode sample;The multi-modal training sample is encoded to multi-layer cross-modal, obtains several cross-modal encoding features, each cross-modal encoding feature includes acoustic encoding feature, oral type encoding feature and environment encoding feature;Cross-modal fusion is carried out to all the cross-modal encoding features, obtains first fusion feature;According to the environment mode sample and the first fusion feature, the parameter of initialization humanoid robot language identification model is updated, obtains well-trained humanoid robot language identification model.The training method provides a kind of humanoid robot language identification model, which is beneficial to improve the accuracy and robustness of language identification.The application relates to the technical field of intelligent robot.
Owner:广州里工实业有限公司

Pre-training method, system and electronic device of multilingual self-supervised model

Embodiments of the present application provide a multilingual self-supervised model pre-training method, system and electronic equipment. The method comprises: inputting unpaired unsupervised speech data selected from a multilingual data set to a language identification network to build a language classifier; extracting target language speech embedding and speech sentence embedding from a target language speech and speech sentence set; determining a training difficulty standard of extended dynamic curriculum learning of the multilingual self-supervised model based on an initial running loss of the multilingual self-supervised model, the target language speech embedding and the speech sentence embedding; and the language classifier performing pre-training of the multilingual self-supervised model in extended dynamic curriculum learning based on the training difficulty standard and a training set of different data amounts dynamically determined from the multilingual data set. The method of the embodiments of the present application makes the multilingual data set more efficient, eliminates potential harmful data on low-resource target language speech, and improves the performance of multilingual self-supervised learning downstream tasks.
Owner:AISPEECH CO LTD

Real-time dynamic display transcription method and system, electronic equipment and storage medium

The invention provides a real-time dynamic display transcription method and system, electronic equipment and a storage medium. The method comprises the following steps: segmenting a single-channel voice stream into continuous multi-frame voice streams; performing human voice detection on the voice stream to obtain a detection result; if a preset number of continuous target voice streams exist, streaming transcription is carried out on each frame of target voice stream and the acquired human voice stream in real time by using a current language through a streaming transcription model; in the streaming transcription process, accumulated audios are collected every preset time period, and the streaming text in the preset time period is corrected by utilizing an accumulated text obtained by transcription of the accumulated audios until a human voice ending signal is detected; correcting the streaming text in the target time period by using a target accumulated text obtained by transferring the target accumulated audio in the target time period to obtain a target corrected text; decoding the target correction text to obtain and display a final text; and updating the current language by using a language obtained by performing language identification on the target accumulated audio.
Owner:CTRIP TRAVEL INFORMATION TECH (SHANGHAI) CO LTD

Streaming end-to-end multilingual speech recognition using joint language identification

ActiveJP7808208B2Speech recognitionNeural learning methodsSpeech soundLanguage identification
The method (400) includes receiving a sequence of acoustic frames (110) as an input to an automatic speech recognition (ASR) model (200). The method also includes generating, by a first encoder (210), a first high-level feature representation (212) for the corresponding acoustic frame. The method also includes generating, by a second encoder (220), a second high-level feature representation (222) for the corresponding first high-level feature representation. The method also includes generating, by a language identification (ID) predictor (230), a language prediction representation (232) based on a combination (231) of the first high-level feature representation and the second high-level feature representation. The method also includes generating, by a first decoder (240a), a first probability distribution (120a) over possible speech recognition hypotheses based on a combination of the second high-level feature representation and the language prediction representation.
Owner:GOOGLE LLC

Multi-language verification system and method

The invention discloses a multi-language verification system and method.The system comprises a data collection module, a data cleaning module, a data model module and a comparison presentation module, and the data collection module is used for periodically recognizing source language identifiers and multiple languages in a logistics system and generating a source language dictionary table; the data cleaning module is used for preprocessing data, the data model module is used for constructing a data dictionary storing a source language and multiple languages corresponding to the source language and maintaining the data dictionary, and the comparison presentation module is used for comparing a latest collected text with the data dictionary. The comparison result is displayed, the specific position and content of the translation error are marked, and correction suggestions are provided according to the comparison result. The method comprises the steps of collecting data, preprocessing the data, constructing a data dictionary, and comparing and displaying translation errors. According to the method, the specific position and content of the translation error can be quickly displayed, the multi-language verification efficiency in the global logistics system is improved, and the multi-language verification time is saved.
Owner:上海捷晓信息技术有限公司

A cross-cultural dynamic portrait-driven multi-agent content generation method and system

This invention discloses a multi-agent content generation method and system driven by cross-cultural dynamic profiling, belonging to the fields of artificial intelligence and cross-cultural information processing technology. The method includes: S1, collecting multilingual text data through social media interfaces and performing structural verification, language identification, text normalization, deduplication, and quality filtering to construct a cross-cultural corpus knowledge base; S2, dividing the corpus into groups based on clustering algorithms, extracting group features, and constructing structured user profiles; S3, scheduling multiple agents to collaboratively perform content generation, trend analysis, and cross-cultural interaction tasks according to user profile tags and rule-based routing; S4, constructing a virtual audience cluster and pre-evaluating the generation strategy using a simulation mechanism that decouples the generation model and the evaluation model; S5, optimizing and iterating the generation strategy based on the evaluation results to form a closed-loop update mechanism. This invention effectively improves the stability and computational accuracy of the cross-language data processing link.
Owner:BEIJING TECH & BUSINESS UNIV

Information handling system with automatic detection and configuration of language model and layout of a keyboard

An information handling system includes an embedded controller that detects the presence of a keyboard within the information handling system. The embedded controller also determines a language identification associated with the keyboard. Based on the language identification being different than a current language identification, a processor receives the language identification from the embedded controller. Based on the language identification, the processor configures a localized language for an operating system of the information handling system.
Owner:DELL PROD LP

Subtitle language recognition methods, devices, computer equipment and computer-readable media

This disclosure provides a method for identifying the language of closed subtitles. The method involves acquiring a video stream with a preset encoding format from the bitstream, obtaining Network Abstraction Layer (NET) data of the supplementary enhancement information type from the video stream, and determining the character encoding of the closed subtitles if they are found in the NET data. The language of the closed subtitles is then determined based on the character encoding and a preset mapping relationship between character encodings and languages. This disclosure allows for the rapid and accurate identification of the language of closed subtitles even when the original bitstream lacks subtitle language information. This disclosure also provides a subtitle language identification device, a computer device, and a computer-readable medium.
Owner:ZTE TECH & SERVICE CO LTD

An input method, device, electronic equipment and computer storage medium

Embodiments of the present application provide an input method and device, electronic equipment and computer storage medium. According to the input scheme provided by the embodiments of the present application, the input text of the user is obtained, the preferred language distribution of the user is determined according to the identity of the user; a pre-trained language identification model is used to determine the first prediction distribution of the language of the input text; according to the first prediction distribution and the preferred language distribution, the second prediction distribution of the language of the input text is generated; and the predicted language of the input text is determined according to the second prediction distribution. By using the preferred language distribution of the user to correct the first prediction distribution of the language identification model, the individualization and accuracy of language identification are enhanced at the user granularity.
Owner:ALIBABA INNOVATION PRIVATE LIMITED

Vehicle-based speech unit with hybrid language detection

PendingCN121708926ASound input/outputSpeech recognitionMore languageLoudspeaker
A hybrid language identification (HLI) system includes one or more microphones configured to detect an acoustic utterance within an interior of a host system; a speaker operable to broadcast prompts or responses within the interior of the host system; a processor and a memory. A processor performs a method that classifies a sound utterance as a mixed utterance having two or more languages using mixed language detection logic stored in a memory, determines a relative language contribution of the languages, and commands a speaker to broadcast a prompt or response within the interior. The prompts and responses have relative language contributions. The HLI system may be used as part of a vehicle having a vehicle body defining a vehicle interior within which hybrid speech recognition occurs.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Image processing apparatus, image processing method, and computer-readable storage medium

An image processing apparatus, an image processing method, and a computer-readable storage medium are disclosed. The image processing apparatus includes a feature extraction unit configured to extract features of an input image; a text detection unit configured to detect text in the input image based on the features extracted by the feature extraction unit; a language identification unit configured to identify a language of the text detected by the text detection unit; a text recognition unit configured to recognize the detected text based on a recognition result of the language identification unit to obtain at least one character string set; and a first classification unit configured to classify the input image by matching the at least one character string set with a predetermined character string set to obtain a first classification result representing a category of an object involved in the input image for obtaining a final classification result of the category of the object.
Owner:FUJITSU LTD

A patent document intelligent classification method and system based on semantic understanding

The application discloses a patent literature intelligent classification method and system based on semantic understanding, acquires patent literature data and classification system configuration parameters, recognizes high-density semantic areas through text segmentation and information density analysis, and constructs a classification index library; adopts a pre-training model to perform vectorization coding to form a semantic vector space, identifies ambiguous feature points through bidirectional semantic detection to form a theme clustering space; implements label matching analysis on the theme clustering space, and performs priority sorting and disambiguation processing to establish a semantic classification rule library; decomposes a classification matching strategy into core feature and auxiliary feature matching sequences, extracts field attributes of a candidate classification set, and implements weight proportioning to determine classification attribution parameters; extracts semantic mapping rules in combination with source language identification, and realizes cross-language retrieval through semantic alignment, thereby realizing accurate understanding of patent technology semantics and supporting multilingual retrieval.
Owner:HUNAN KUANGCHU TECH CO LTD

Code multilingualization method, device, equipment, storage medium and program product

Embodiments of the present disclosure disclose a code multilingualization method, device, equipment, storage medium and program product, wherein the code multilingualization method comprises: determining a code variable capable of being recognized by a current code running environment; the code variable comprises original language text in source code and translated language text of the original language text in multiple languages; based on target original language text to be processed, the code variable and language identification, determining a to-be-replaced variable corresponding to the target original language text; replacing the target original language text in the source code with the to-be-replaced variable to obtain multilingual source code; the to-be-replaced variable is used to adjust the translated language text corresponding to the target original language text based on the language identification. In this way, the multilingual switching of the original language text in the source code can be intelligently realized, the operation is simple, and the applicability is strong.
Owner:PATEO CONNECT (NANJING) CO LTD

Long-text-oriented passage-level knowledge extraction method, system, device and medium

This application relates to a method, system, device, and medium for text-level knowledge extraction from long texts, belonging to the field of knowledge extraction technology. The method includes: inputting diverse heterogeneous data, parsing it into text according to a format, and then performing language identification and translation; connecting to a knowledge platform to obtain and structure entity information, and generating large language model prompt words based on prompt word templates; designing example data format conversion sample data, inputting it into the model to generate extraction results, parsing it into RDF triples, and performing entity alignment and disambiguation. This application can accurately extract knowledge from diverse heterogeneous long texts, solve the problem of text fragmentation, and also achieve entity alignment and disambiguation, thereby improving the consistency and quality of knowledge graphs.
Owner:TUPU INTELLIGENT TECH (BEIJING) CO LTD

Semantic meaning recognition method and device, equipment, medium and program product

The invention provides a semantic recognition method and device, equipment, a medium and a program product. Relates to the financial science and technology field. The method comprises the steps of obtaining a to-be-processed question of a user, and determining a first candidate semantic meaning corresponding to the to-be-processed question; when the semantic similarity probability of the first candidate semanteme meets a preset matching condition, determining the first candidate semanteme as a target semanteme of the to-be-processed problem; when the semantic similarity probability of the first candidate semanteme does not meet a preset matching condition, determining a second candidate semanteme corresponding to the to-be-processed problem from the candidate service semanteme with the highest semantic similarity probability in the candidate service semanteme; determining a first integration probability corresponding to the first candidate semantic meaning and a second integration probability corresponding to the second candidate semantic meaning; and based on the first integration probability and the second integration probability, determining a target semantic meaning corresponding to the to-be-processed problem. According to the method, the technical problem of low semantic recognition accuracy in the prior art is solved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Multilinguistic Query Response Agent With Agent Core

PendingUS20260187057A1Natural language processingCommand language
Techniques for responding to multilinguistic queries using an agent core are disclosed herein. An agent core is trained and / or fine-tuned in a first language to generate instructions (i.e., commands) for answering a query. A language identification and / or translation model receives queries, identifies languages associated with the queries, and translates the queries to the first language. The agent core generates instructions in the first language based on the translated queries. The instructions include instructions, in the first language, to perform actions such as retrieval, generation, contextual understanding, or calculation, in both the first language and the second language. The results of executing the instructions, including an action performed in the first language and an action performed in one or more second languages, are combined, ranked and / or reranked to generate an answer to the query.
Owner:ORACLE INT CORP

System and method for speech language identification

A method, computer program product, and computing system for speech language identification. An input voice signal of a specific language is received. The input speech signal is processed by a plurality of speech recognition processing paths, each speech recognition processing path identifying an associated subset of languages. Using each of the speech recognition processing paths, an input speech signal is processed using machine learning to identify a language that most closely matches a particular language of the input speech signal to produce a plurality of identified languages. An indication of each of the input speech signal and the plurality of identified languages is received in a further speech recognition processing path. The input speech signal is processed using machine learning to identify one of the identified languages as closest matching to the particular language.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A method, apparatus and storage medium for identifying a font language

Disclosed are a font language identification method, device and storage medium. The font language identification method comprises the following steps: obtaining font attributes of a font to be identified, wherein the font attributes comprise a font name language; identifying a language according to the font name language in the font attributes, and determining the identified language as a language supported by the font to be identified. According to the method provided herein, the same font can be determined to belong to the same language on different operating systems, thereby maintaining the same user experience.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

Multilingual text classification method, apparatus, device, and medium

The application discloses a multilingual text classification method, device and equipment and a medium. The method comprises the following steps: obtaining a target text and a pre-trained learning model, wherein the learning model comprises a shared feature extraction network and a plurality of sub-task identification networks; obtaining a sentence vector representation of the target text and a language identification prediction result through the shared feature extraction network; and calling a sub-task identification network corresponding to the language to process the sentence vector representation according to the language identification prediction result, so as to obtain a classification result of the target text, wherein a language self-learning module in the sub-task identification network learns the correlation between a plurality of languages corresponding to the language. The application can integrate the correlation knowledge between a plurality of languages into the model for learning, and classify the multilingual text through the model, so that the multilingual text can be better classified. Accordingly, the application also provides a multilingual text classification device, equipment and medium.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES