Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

330 results about "Language translation" patented technology

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

Animal language conversion methods, devices, electronic equipment and storage media

This disclosure provides a method, apparatus, electronic device, and storage medium for animal language conversion, relating to the field of artificial intelligence technology, specifically machine learning, deep learning, and natural language processing. The specific implementation involves: acquiring multimodal data related to the animal, including animal vocal data, animal behavioral data, and animal physical characteristics data; preprocessing the multimodal data to obtain fused multimodal data; identifying the animal's current emotion based on the fused multimodal data to obtain an emotion recognition result; and performing semantic mapping and language translation on the emotion recognition result to convert the animal language into human language, obtaining a language conversion result. This disclosure can accurately identify the animal's current emotional state and convert it into human language, thereby achieving deeper emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Sign language translation model, system and method based on full-modal alignment

The invention discloses a sign language translation model, a sign language translation system and a sign language translation method based on full-modal alignment. The sign language translation method comprises the steps of extracting multi-modal features of hand, face and body postures from an input video and performing preliminary fusion; deep secondary fusion and alignment are carried out through multi-scale time sequence coding and a cross-modal collaborative attention mechanism, and a spatial-temporal feature sequence of full-modal alignment is generated; performing boundary detection and dynamic segmentation on the feature sequence by using a CTC-based sequence prediction model, and outputting a discrete sign language word sequence with a timestamp; and finally, capturing a sign language grammar structure of the sequence through a graph structure enhanced Transform encoder, inputting the sign language grammar structure into a Transform decoder integrating grammar consistency loss, and generating a target text conforming to target natural language grammar and semantic rules. According to the method, the problems of continuous sign language action adhesion and grammar structure difference are effectively solved, and the sign language translation accuracy and the natural language generation fluency are greatly improved.
Owner:SURELY ACCESSIBLE TECH (SUZHOU) CO LTD

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Low-resource language translation model training method based on large language

The invention discloses a low-resource language translation model training method based on a big language, which comprises the following steps of: continuously pre-training a basic big language model by utilizing a multilingual text corpus, and storing a plurality of middle check points as candidate models; selecting a model with the optimal downstream translation task performance from the candidate models based on the performance of the verification set, and performing instruction supervision fine tuning by using the parallel instruction data set to obtain an intermediate model; and finally, training the intermediate model by using a preference optimization algorithm to obtain a target translation model by using a preference data set consisting of preferred translation and rejected translation. According to the method, through three-step progressive training, the problems of data sparsity, insufficient model capability, single training method and the like in low-resource language translation are effectively solved, and the translation quality is remarkably improved.
Owner:北京中科闻歌科技股份有限公司 +2

Asynchronous anthropomorphic communication method and system based on voice transfer

The invention provides an asynchronous anthropomorphic communication method and system based on voice transfer, which are suitable for various terminals such as wearable equipment, earphones, dolls, bolsters and the like. The method comprises the steps of user voice input, edge or cloud recognition, playback confirmation, content translation and voice synthesis, asynchronous transmission and broadcast and the like. The system does not depend on a specific hardware form, emphasizes a user confirmation mechanism and anthropomorphic voice broadcast and supports multilingual translation and personalized voice styles, the communication process is bound based on a device ID or nickname identity, social account login is not needed, and interaction privacy security is guaranteed. The method is widely applicable to various asynchronous social application scenes such as children, old people, lovers, autism rehabilitation and the like. The system supports nickname binding and friend relationship establishment, users can complete social connection through voice instructions or two-dimensional codes, and controllability and interestingness of communication interaction are enhanced.
Owner:GUANGDONG OPERATOR WIRE INTELLIGENT TECHNOLOGY CO LTD

Robot operation method, system and terminal based on visual language model

The invention relates to the field of robot intelligent control, and discloses a robot operation method, system and terminal based on a visual language model, and the method comprises the steps: obtaining a visual image and a natural language instruction, carrying out the combined analysis of the visual image and the natural language instruction through a visual language large model, and generating an initial planning strategy; the robot is controlled to execute operation actions according to the initial planning strategy, and original tactile signals in the operation process are collected; inputting the original tactile signal into a tactile language translation model for feature extraction and semantic mapping to obtain tactile semantic description; when the prompt operation of the tactile semantic description is abnormal or the physical attribute does not accord with the visual expectation, the tactile semantic description is fed back to the visual language large model as an enhanced prompt word; and according to the visual image and the enhanced prompt word, the current task scene is reasoned again, a corrected planning strategy is generated, and the robot is controlled to execute an operation action. The operation precision of the robot is improved.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

Guiding language translation with translation documents using machine learning

In accordance with the described techniques, a system receives a plurality of facets describing language-agnostic aspects of language translation, a translation document describing language-specific rules for translating from a source language to a target language, and a source text in the source language. Using one or more machine learning models, a plurality of guidelines are extracted from the translation document and assigned to respective facets of the plurality of facets. The system translates the source text to a translated text in the target language using one or more machine learning models conditioned on the plurality of guidelines assigned to the respective facets.
Owner:ADOBE INC

AUTOMATIC TRANSCRIPT-ASSISTED SPEECH LANGUAGE TRANSLATION USING LANGUAGE MODELS

Devices, systems, and techniques are disclosed that implement the training and deployment of automatic transcription-based translation systems using language models. The techniques include: processing, using a first speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language. The second input includes at least a representation of the spoken language and the transcription of the spoken language.
Owner:NVIDIA CORP

Computing technologies for evaluating linguistic content to predict impact on user engagement analytic parameters

Correlations between a set of linguistic features identified in an unstructured text recited in a source language and a set of user engagement analytic parameters may be measured by a machine learning model selected based on a set of performance metrics from a set of machine learning models trained by a set of supervised machine learning algorithms on (i) a set of unstructured texts recited in the source language and containing the set of linguistic features and (ii) the set of user engagement analytic parameters measured for the set of unstructured texts. The machine learning model grades the unstructured text recited in the source language to determine whether the unstructured text recited in the source language should be (1) edited in the source language and then translated into the target language or (2) translated from the source language to the target language as is.
Owner:WELOCALIZE INC

Portable multi-language intelligent acquisition and translation system

The invention relates to the technical field of language translation, and discloses a portable multi-language intelligent acquisition and translation system, which comprises a multi-mode acquisition module used for acquiring audio signals and video streams; the environment sensing and scheduling module is used for evaluating the complexity of the current environment so as to dynamically adjust computing resources; the target voice construction module is used for determining a current speaker and extracting a target audio stream of the current speaker, and generating a voice text based on the target audio stream; the analysis module is used for generating visual situation metadata; the translation module is used for generating a preliminary translation result and a corresponding translation confidence score; and the interactive output module is used for carrying out ambiguity clarification to generate a final translation result. According to the invention, audio-visual fusion is carried out on the audio signal and the video stream, and the current speaker is determined in a multi-person noisy environment in combination with the face position and lip movement information of the video stream, so that the interference of background noise and other non-target speakers is eliminated, and the accuracy of subsequent speech recognition and translation is improved.
Owner:CHANGJI UNIV

Visual sign language translation training device and method

Methods, devices and systems for training a pattern recognition system are described. In one example, a method for training a sign language translation system includes generating a three-dimensional (3D) scene that includes a 3D model simulating a gesture that represents a letter, a word, or a phrase in a sign language. The method includes obtaining a value indicative of a total number of training images to be generated, using the value indicative of the total number of training images to determine a plurality of variations of the 3D scene for generating of the training images, applying each of plurality of variations to the 3D scene to produce a plurality of modified 3D scenes, and capturing an image of each of the plurality of modified 3D scenes to form the training images for a neural network of the sign language translation system.
Owner:AVODAH INC

Federal multi-language machine translation method based on efficient fine tuning

The invention relates to a federal multi-language machine translation method based on efficient fine tuning, and belongs to the technical field of natural language processing. Aiming at the problems of high communication cost and long training time in a federated learning-based multi-language machine translation method, the invention provides a federated multi-language machine translation method based on efficient fine tuning, which comprises the following steps of: efficiently fine-tuning a multi-language translation model of a client; performing gradient similarity clustering on the fine-tuned multi-language translation model; carrying out average aggregation based on the clustered multi-language translation model; and deploying a federal multi-language machine translation device based on efficient fine tuning. According to the method, the calculation and communication overhead is greatly reduced while the translation performance is kept, and the method is suitable for distributed translation tasks in a multi-language scene.
Owner:KUNMING UNIV OF SCI & TECH

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

Biometric sensor of a tactical gear to monitor health and stress of a wearer

Disclosed are a method, system, and apparatus of a biometric sensor of a tactical gear to monitor health and stress of a wearer. In one embodiment, personal protective equipment includes a language translator module integrated in a tactical gear, a biometric sensor, and a responsive device integrated in the tactical gear. A wearer bi-directionally communicates with an individual in an ambient environment using any language other than a primary language spoken by the wearer when the language translator module is activated. The biometric sensor detects the health condition of the wearer. The responsive device vibrates to notify the wearer when the biometric sensor detects an elevated stress level of the wearer of the tactical gear. The biometric sensor is detachable from the tactical gear and placeable on an injured person nearby to the wearer through an armband extendable from the biometric sensor when it is removed from the tactical gear.
Owner:GOVERNMENTGPT INC

Intelligent speech translation mobile phone and system capable of realizing multi-language inter-translation

The invention belongs to the technical field of mobile terminal communication and language translation, and particularly relates to an intelligent speech translation mobile phone and system capable of realizing multilingual inter-translation, which are characterized in that a two-way speech separation module and unit, a language recognition related module and unit, an end-cloud collaborative translation module and unit and a system are compatible with related module and unit collaboration; two-way voice collection, separation, language recognition and end-cloud collaborative translation are completed in a call scene, the system is compatible with a mainstream mobile operating system and communication application, a resource isolation and process linkage mechanism is adopted to guarantee operation compatibility, multilingual communication obstacles in the call scene are effectively solved, complex environment interference is adapted, the real-time performance and accuracy of translation are considered, and the system is suitable for large-scale popularization and application. And the native function of the terminal does not need to be modified.
Owner:SHENZHEN GUO ELECTRONIC INFORMATION CO LTD

Method and system for packaging webpage embedded sign language translation JS library

The invention discloses a method and system for packaging a webpage embedded sign language translation JS library, and belongs to the field of electrical digital data processing.The method comprises the steps that by constructing an ECS core architecture with perception feedback, a bidirectional data flow between modules is established to achieve dynamic collaboration of rendering and grammar; predictive streaming resource management is adopted, and resources are intelligently preloaded based on gesture sequence contexts; multi-instance physical isolation is realized by utilizing a sandboxed container and a ShadowDOM (Document Object Model); implementing multi-level adaptive rendering, and dynamically switching a rendering path and a frame rate according to equipment performance; and finally, shielding the bottom layer difference through cross-end unified rendering of the abstraction layer. According to the method, the problems of animation and semantic dislocation, resource loading conflict and poor equipment compatibility in a multi-user concurrent scene are effectively solved, and high-fidelity, low-delay and cross-end consistent experience of sign language translation animation is realized.
Owner:SURELY ACCESSIBLE TECH (SUZHOU) CO LTD +1

Pet language translation method and system based on audio learning

The invention relates to a pet language translation method and system based on audio learning, and belongs to the technical field of animal training. The method comprises the following steps: constructing a pet standardized sound library; when the preset scene is triggered, playing the target sound signal; the target sound signal is a sound signal related to a preset scene in a pet standardized sound library; when it is detected that the first sound signal sent by the pet is matched with the sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matched sound signal; collecting pet sound in the current environment in real time, responding to the matching of the pet sound and the sound signal in the pet standardized sound library, and outputting a pet demand of the matched sound signal; the pet demand corresponds to a semantic translation result of the pet sound. By means of the mode, the training logic which can be stably recognized and can be copied and executed can be provided, accurate and reasonable pet language translation can be achieved, and the probability of mistranslation or wrong translation is reduced.
Owner:SHENZHEN KOLAMAMA TECH CO LTD

Biometric sensor of a tactical gear to monitor health and stress of a wearer

Disclosed are a method, system, and apparatus of a biometric sensor of a tactical gear to monitor health and stress of a wearer. In one embodiment, personal protective equipment includes a language translator module integrated in a tactical gear, a biometric sensor, and a responsive device integrated in the tactical gear. A wearer bi-directionally communicates with an individual in an ambient environment using any language other than a primary language spoken by the wearer when the language translator module is activated. The biometric sensor detects the health condition of the wearer. The responsive device vibrates to notify the wearer when the biometric sensor detects an elevated stress level of the wearer of the tactical gear. The biometric sensor is detachable from the tactical gear and placeable on an injured person nearby to the wearer through an armband extendable from the biometric sensor when it is removed from the tactical gear.
Owner:GOVERNMENTGPT INC

Large language models for creating a multi-lingual, low-resource code translation dataset

One or more unit-test cases are generated from a monolingual code corpus and the generated unit-test cases are filtered to generate a corpus of unit-test cases which have acceptability scores exceeding one or more predefined thresholds. One or more of the code samples of the monolingual code corpus are translated from a source language to a target language using a pretrained Large Language Model and the generated unit-test cases are translated from the source language to the target language. The LLM-translated code samples are validated using the translated unit-test cases and a parallel-data training corpus comprising the LLM-translated code samples that pass the validation is created. The pretrained large language model (LLM) is fine-tuned using the parallel-data training corpus, a given code segment is translated using the fine-tuned large language model (LLM), the translated given code segment is tested and the tested given code segment is deployed.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION +1

AI intelligent glasses

The invention provides AI intelligent glasses which comprise a glasses body, a hardware system and a software system, the hardware system comprises a main control chip, a camera module, a microphone array, a loudspeaker, a display module, a communication module and a power supply module, and the software system comprises an image processing module, a visual analysis module and a semantic generation and multi-language module which are integrated on the main control chip. The system integrates visual AI, natural language processing, AR interaction and other technologies in an interdisciplinary manner, has the functions of image recognition, attribute analysis, knowledge reasoning, language generation and multi-language translation, realizes intelligent perception of any object in the real world, and enhances the understanding ability of a user for unknown objects; edge reasoning and real-time voice feedback are supported, the interaction efficiency is improved, and the intelligence and adaptability of visual auxiliary equipment are improved; the modular design is convenient to adapt to different languages and application scenes, and can be widely applied to scenes such as tourism, education, medical treatment and industry.
Owner:ZHEJIANG NORMAL UNIV

Sign language translation method and device based on large language model

The invention belongs to the crossing field of artificial intelligence and barrier-free assistance technology, and particularly discloses a sign language translation method and device based on a large language model.The method is applied to a terminal and comprises the steps that key points are extracted from sign language video frames; based on the coordinates, types and global indexes of the key points in the sign language video frames, obtaining fusion features through feature fusion; based on the fusion features corresponding to the sign language video frames in the frame window, time sequence modeling and classification are carried out, and a word prediction result of the frame window is obtained; a sliding window voting and deduplication mode is adopted, the candidate word sequence is preprocessed, a preprocessed word sequence is obtained, and the candidate word sequence is formed by connecting word prediction results of all frame windows in series according to a time sequence; based on the preprocessed word sequence, word error correction and statement generation are carried out by calling a large language model, and a sign language translation result is generated. According to the invention, sign language video translation can be realized for the terminal, and the accuracy and fluency of the translation result can be guaranteed.
Owner:SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES

Generating Clinical Documentation Using Large Language Models and Artificial Intelligence

Systems and methods generate clinical documentation using large language models and artificial intelligence (AI). A template management module is provided to create customizable templates. A processing unit can receive input data from various sources and use AI to generate transcripts, summarize sessions, and produce clinical documentation such as clinical notes. The processing unit may also generate Current Procedural Terminology (CPT) and diagnosis codes, generate after-visit summaries, and generate referral letters. The AI may be trained on past clinical notes and can adapt to the clinician's style over time, with a feedback loop for continuous improvement. Additional features include cohort-based training, real-time language translation, predictive text, and analytics for documentation trends. The system supports customization of note length, style, and keywords, as well as integration with external medical databases and patient portals.
Owner:ORCHID EXCHANGE INC

Multi-agent collaborative cross-language system translation method and system based on loAs

The invention provides a multi-agent collaborative cross-language system translation method and system based on loAs, and relates to the technical field of intelligent translation.The method comprises the steps that to-be-translated text information and conference theme information are obtained, and first fusion semantic information is obtained; obtaining first target semantic information through the first translation agent; obtaining second target semantic information through a second translation agent; obtaining second fusion semantic information through the first target semantic information and the conference theme information, and determining a second target language translation text; and acquiring third fused semantic information through the second target semantic information and the conference theme information, and determining the first target language translation text. According to the method and the device, the translation accuracy of key texts can be improved by referring to conference theme information, and multi-language semantic information can be mutually corrected through a cross attention mechanism to improve the semantic consistency of multi-language translation, so that the overall translation accuracy is improved, and the communication cost is reduced.
Owner:AIYU (SHANGHAI) INFORMATION TECHNOLOGY CO LTD

Sign language real-time translation system based on multi-dimensional data fusion

The invention relates to the field of sign language translation. The invention discloses a sign language real-time translation system based on multi-dimensional data fusion. The sign language real-time translation system comprises an image information acquisition module, a data processing module and a result display module, the image information acquisition module captures sign language actions of a user in real time through a camera to generate sign language videos, and operates four core task architectures including an HTTP server task, a camera acquisition task, a streaming media transmission task and a video downloading task on the basis of FreeRTOS, the HTTP server task receives a client request, and the video downloading task receives the client request. Comprising an MJPEG video stream, a JPEG single frame capture and a video downloading access endpoint, a built-in CORS cross-domain access support, a camera acquisition task continuously captures JPEG frames with a fixed frame rate, each task is connected with a queue transfer client through semaphore protection frame data to realize efficient collaboration, multi-task parallel processing is realized, the transmission frame rate is accurately controlled to meet the algorithm requirement, and the algorithm requirement is met. And ESP32 dual-core resources are utilized to the maximum extent.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Oral language translation for interactivity in virtualized worlds

Methods, systems, and computer-readable storage media are disclosed for translating a user input to a virtual environment into a contextualized output. The input is converted into a first textual representation by a recognition model, and a first set of tokens based on the textual representation is generated. The first set of tokens is fused with a second set of tokens stored in a contextualized language database. The second set of tokens is based on a second textual representation of previously collected user interactivity metrics, a virtual environment engine configuration, or displayable attributes. A trained neural network uses the fused set of tokens and at least a portion of the second set of tokens to generate an assessment of user activity to adjust a first display attribute, change the current position of the user within the virtual environment, or generate a natural language audio or textual output from the virtual environment.
Owner:WOODARD JR KENNETH LA-VERNE

Language translation system and method based on big data

The invention relates to the technical field of big data translation systems, and discloses a language translation system and method based on big data. The method comprises the steps of collecting multi-language historical data, and obtaining a source language text data set, a target language text data set and a translation evaluation data set which are collected within a preset time range; then, executing stability evaluation on the three data sets to obtain a language use stability factor, a translation consistency stability factor and an evaluation reliability stability factor; then, the stability factors serve as key indexes, dense exploration is executed in a translation parameter configuration domain, and priority translation configuration is determined; and finally, translating the input source language text according to the priority translation configuration and a predetermined translation task set to generate a serialized translation data set, analyzing the set by adopting a translation result analyzer, and outputting a final translation result.
Owner:SANYA UNIVERSITY

A large language model multilingual enhancement method and system based on model combination

This application discloses a method and system for multilingual enhancement based on a large language model using model ensemble. The system includes: a pre-trained multilingual translation model, a semantic representation mapping module, and a large language model. The multilingual translation model is used for multilingual semantic modeling and language generation, including a multilingual encoder module and a multilingual decoder module. The semantic representation mapping module is used to transform the latent space representations of different models into an interactive unified semantic space based on a cross-model representation mapping mechanism. The output of the multilingual encoder is mapped to the unified semantic representation space of the large language model, and the mapped semantics are input into the large language model to perform language-independent instruction understanding. The intermediate semantic representation output by the large language model is mapped and transformed to a cross-attention representation space, generating the final output text under the target language distribution. The system of this application outperforms existing technologies in terms of efficiency, stability, and generation quality in multilingual capability extension.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Headphone components

ActiveCN309763371SHeadphonesTesting Methods
1. Name of the product in this design: Headphone Assembly. 2. Purpose of this design: Primarily used for communication, real-time language translation, and audio transmission. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design points: Component 1 3D view 1. 5. Other situations requiring explanation: Component 1 of this design is the right earphone, and component 2 is the left earphone.
Owner:IFLYTEK CO LTD

Earphone assembly

1. The name of the design product: earphone assembly. 2. The use of the design product: mainly used for communication, instant language translation, audio transmission. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: assembly 1 perspective view. 5. Other circumstances that need to be explained: the design product assembly 1 is the left earphone, and assembly 2 is the right earphone.
Owner:IFLYTEK CO LTD