Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

441 results about "Language translation" patented technology

Low-resource language translation method and system based on deconstruction distillation

The invention relates to the technical field of language translation, in particular to a low-resource language translation method and system based on deconstruction distillation, and the method comprises the steps: obtaining parallel corpus data of a low-resource language and a general language; the method comprises the following steps: constructing a teacher model by taking parallel corpus data as input, constructing a student model trunk based on a pre-trained BERT big language model and optimizing the student model, calculating cross-task attention alignment loss based on the teacher model and the student model, and outputting and executing logits distillation based on the teacher model and the student model. Deployment of low-resource language translation is completed based on the trained student model, a high-quality knowledge migration source is provided for the student model by constructing the BERT teacher model subjected to full-parameter fine adjustment, and meanwhile, the problem of translation precision caused by insufficient low-resource language data is effectively solved by means of dual supervision of cross-task attention alignment and logits distillation.
Owner:YANTAI UNIV

Multi-modal time sequence alignment AI video translation method and system

The invention relates to the technical field of subtitle translation, in particular to a multi-modal time sequence alignment AI video translation method and system, and the method comprises the steps: 1, carrying out the multi-modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint feature data and visual time sequence data; 2, performing cross-language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint feature data and the target language text; generating a mouth shape animation matched with the target language voice based on the lip key point data and the limb action time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the mouth shape animation and the limb action sequence through a cross-modal time sequence encoder, and dynamically adjusting the layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.
Owner:HANGZHOU BAOMIHUA TECH CO LTD

Sign-language translation

System and techniques to facilitate the translation of a sign language into another language are described herein. A modular architecture may be used in which the output of different classifiers may be used to produce intermediate representations, or final translations, of the sign language. These classifiers may be trained on different types of signs to enhance accuracy while reduce training time and complexity.
Owner:SORENSON IP HOLDINGS LLC

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

Real-time language translation systems embodied in a physical device

The present disclosure provides a real-time language translation system embodied in a physical device. Further, the real-time language translation system may include an input device which may be configured for receiving generating a user input data representing a linguistic input from a user. Further, the linguistic input corresponds to a user language. Further, the real-time language translation system may include a processing device which may be configured for generating a translation data based on the user input data. Further, the translation data represents a translation of the linguistic input. Further, the generating may be based on an AI module. Further, the processing device may be communicatively coupled to the input device. Further, the real-time language translation system may include a presentation device which may be configured for presenting the translation data. Further, the presentation device may be communicatively coupled to the processing device.
Owner:MYLANGUAGE INC

Animal language conversion methods, devices, electronic equipment and storage media

This disclosure provides a method, apparatus, electronic device, and storage medium for animal language conversion, relating to the field of artificial intelligence technology, specifically machine learning, deep learning, and natural language processing. The specific implementation involves: acquiring multimodal data related to the animal, including animal vocal data, animal behavioral data, and animal physical characteristics data; preprocessing the multimodal data to obtain fused multimodal data; identifying the animal's current emotion based on the fused multimodal data to obtain an emotion recognition result; and performing semantic mapping and language translation on the emotion recognition result to convert the animal language into human language, obtaining a language conversion result. This disclosure can accurately identify the animal's current emotional state and convert it into human language, thereby achieving deeper emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Sign language translation model, system and method based on full-modal alignment

The invention discloses a sign language translation model, a sign language translation system and a sign language translation method based on full-modal alignment. The sign language translation method comprises the steps of extracting multi-modal features of hand, face and body postures from an input video and performing preliminary fusion; deep secondary fusion and alignment are carried out through multi-scale time sequence coding and a cross-modal collaborative attention mechanism, and a spatial-temporal feature sequence of full-modal alignment is generated; performing boundary detection and dynamic segmentation on the feature sequence by using a CTC-based sequence prediction model, and outputting a discrete sign language word sequence with a timestamp; and finally, capturing a sign language grammar structure of the sequence through a graph structure enhanced Transform encoder, inputting the sign language grammar structure into a Transform decoder integrating grammar consistency loss, and generating a target text conforming to target natural language grammar and semantic rules. According to the method, the problems of continuous sign language action adhesion and grammar structure difference are effectively solved, and the sign language translation accuracy and the natural language generation fluency are greatly improved.
Owner:SURELY ACCESSIBLE TECH (SUZHOU) CO LTD

Bilingual medical mixture of experts large language model

A computer-implemented system and computer instructions stored on non-transitory computer readable medium for bilingual medical inquiry in both Arabic and English, including multiple-choice question answering, open-ended question answering, and multi-turn question answering. The system and instructions use a mixture of experts large language model (MOE LLM) having a router network connected to multiple expert networks. The MOE LLM is trained with medical domain data and is used to receive the input bilingual text in a format for a medical inquiry, and output text in a format of a response to the medical inquiry, in sequence. The system and instructions incorporate an English-to-Arabic translation pipeline having a language translation model to generate Arabic language medical instruction sets from English language medical instructions, for large scale use in Arabic and English medical inquiry.
Owner:MOHAMED BIN ZAYED UNIV OF ARTIFICIAL INTELLIGENCE

Performing artificial intelligence sign language translation services in a video relay service environment

Video relay services, communication systems, non-transitory machine-readable storage media, and methods are disclosed herein. A video relay service may include at least one server configured to receive a video stream including sign language content from a video communication device during a real-time communication session. The server may also be configured to automatically translate the sign language content into a verbal language translation during the real-time communication session without assistance of a human sign language interpreter. Further, the server may be configured to transmit the verbal language translation during the real-time communication session.
Owner:SORENSON IP HOLDINGS LLC

Multilingual work order processing method and device based on LoRA network, equipment and medium

The invention discloses a multilingual work order processing method and device based on a dynamic LoRA network, computer equipment and a storage medium, and the method comprises the steps: receiving multilingual voice input, and converting the voice input into a source language text in real time; identifying a feature vector of the source language text; calculating the correlation with a plurality of LoRA networks to be selected according to the feature vectors; determining an activation probability of each to-be-selected LoRA network based on the correlation, and dynamically activating a target LoRA network through a pre-configured gating network according to the activation probability; fusing the target LoRA network into a basic large language model to generate a fusion model; and performing cross-language translation and work order content generation on the source language text by using the fusion model, and outputting a standard work order of a target language so as to distribute the work order to a corresponding processing module according to the work order content. By dynamically activating the proper LoRA network, the efficiency and accuracy of multilingual work order processing are improved.
Owner:HANGJU DATA SERVICES (SHANGHAI) CO LTD

Multinational universal conference language translation method, device and equipment and storage medium

InactiveCN120725030ANatural language translationMathematical modelsTranslation languageTimestamp
The invention relates to a multi-national universal conference language translation method, device and equipment and a storage medium, and the translation method comprises the steps: continuously collecting all on-site voice information through a sensor, and carrying out the voice information preprocessing; identifying each sounding main body in the first voice information and corresponding first target voice information, temporarily storing the first voice information in a buffer queue, and marking a timestamp and the sounding main body of each target voice information in the first voice information, processing the second voice information and other newly-added voice information according to the same processing flow of the first voice information; and extracting the multi-dimensional acoustic features of the marked first target voice information, generating translation content through a pre-trained hybrid model according to a preset target translation language, and outputting the translation content to an application scene in a preset mode, thereby realizing the real-time performance and accuracy of multi-national conference language translation, improving the efficiency and effect of an international conference, and improving the user experience. And the communication and cooperation of the conference are promoted.
Owner:JI LI MI (SHEN ZHEN) JI SHU YOU XIAN GONG SI

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Fine-grained per-vector scaling for neural network quantization

Today neural networks are used to enable autonomous vehicles and improve the quality of speech recognition, real-time language translation, and online search optimizations. However, operation of the neural networks for these applications consumes energy. Quantization of parameters used by the neural networks reduces the amount of memory needed to store the parameters while also reducing the power consumed during operation of the neural network. Matrix operations performed by the neural networks require many multiplication calculations, so reducing the number of bits that are multiplied reduces the energy that is consumed. Quantizing smaller sets of the parameters using a shared scale factor improves accuracy compared with quantizing larger sets of the parameters. Accuracy of the calculations may be maintained by quantizing and scaling the parameters using fine-grained per-vector scale factors. A vector includes one or more elements within a single dimension of a multi-dimensional matrix.
Owner:NVIDIA CORP

Low-resource language translation model training method based on large language

The invention discloses a low-resource language translation model training method based on a big language, which comprises the following steps of: continuously pre-training a basic big language model by utilizing a multilingual text corpus, and storing a plurality of middle check points as candidate models; selecting a model with the optimal downstream translation task performance from the candidate models based on the performance of the verification set, and performing instruction supervision fine tuning by using the parallel instruction data set to obtain an intermediate model; and finally, training the intermediate model by using a preference optimization algorithm to obtain a target translation model by using a preference data set consisting of preferred translation and rejected translation. According to the method, through three-step progressive training, the problems of data sparsity, insufficient model capability, single training method and the like in low-resource language translation are effectively solved, and the translation quality is remarkably improved.
Owner:北京中科闻歌科技股份有限公司 +2

Sign language translation training system based on gesture tracking and sign language tracking device

The invention relates to the technical field of sign language translation, in particular to a sign language translation training system based on gesture tracking and a sign language tracking device, and the system comprises a gesture tracking module, a skeleton structure completion module, a gesture standard comparison module, a gesture dynamic correction module and a training effect analysis module. According to the method, through screening visible skeleton joint points, detecting a space displacement trend and calculating a motion rate and a direction change rate, key frame data meeting a stability standard can be effectively extracted, the accuracy of gesture track analysis is improved, missing key points are complemented by using skeleton structure constraints, interference on an identification result caused by loss of the key points is avoided, and the accuracy of gesture track analysis is improved. Dynamic posture correction is performed by calculating the joint angle change rate and analyzing the deviation range and combining the finger opening and closing angle and the wrist rotation angle, adjustment suggestions are provided according to the correction angle, sign language training optimization guidance is achieved, and the pertinence and the refinement level of sign language training are improved.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Identify and obfuscate sensitive data before ingesting to generative ai engines

An approach is provided that identifies a sensitive data in a request to an artificial intelligence (AI) engine. In one embodiment, the identifying further includes: translating the request from one natural language to other natural languages, thus creating translated requests, comparing fields in the first natural language to fields in the other natural languages with the comparing resulting in some untranslated fields that are treated as the sensitive data. The approach further includes creating an obfuscated request by obfuscating the sensitive data identified in the request before the obfuscated request is transmitted to the AI engine.
Owner:KYNDRYL INC

Deaf-mute sign language translation pronunciation system

The invention relates to the technical field of sign language intelligent communication, in particular to a deaf-mute sign language translation pronunciation system which comprises a data acquisition and preprocessing module, a sign language model recognition module, a dynamic gesture model module and a voice conversion module. According to the sign language translation pronunciation system for the deaf-mute, wide-angle camera intelligent glasses are adopted for visual collection, wearable equipment is not needed, video streams and thermodynamic diagrams are combined, visual information and key point probability distribution are complementary, the influence of single-mode noise is reduced, video space-time features are extracted through a 3D CNN, long-range dependence is captured in combination with a self-attention mechanism of Transform, and therefore the sign language translation pronunciation system for the deaf-mute is obtained. The sign language recognition precision is improved, an internal reference matrix and a distortion coefficient are calculated through a calibration board, wide-angle lens distortion is corrected, it is ensured that hand key points are accurately positioned, YOLOv8 is adopted to segment an interference object, background noise is prevented from affecting recognition, 21 key point thermodynamic diagrams are generated through MediaPipe, fault tolerance of low-confidence-coefficient key points is enhanced through Gaussian kernel diffusion, and the recognition accuracy is improved. And the model identification precision is improved.
Owner:马也顺

Asynchronous anthropomorphic communication method and system based on voice transfer

The invention provides an asynchronous anthropomorphic communication method and system based on voice transfer, which are suitable for various terminals such as wearable equipment, earphones, dolls, bolsters and the like. The method comprises the steps of user voice input, edge or cloud recognition, playback confirmation, content translation and voice synthesis, asynchronous transmission and broadcast and the like. The system does not depend on a specific hardware form, emphasizes a user confirmation mechanism and anthropomorphic voice broadcast and supports multilingual translation and personalized voice styles, the communication process is bound based on a device ID or nickname identity, social account login is not needed, and interaction privacy security is guaranteed. The method is widely applicable to various asynchronous social application scenes such as children, old people, lovers, autism rehabilitation and the like. The system supports nickname binding and friend relationship establishment, users can complete social connection through voice instructions or two-dimensional codes, and controllability and interestingness of communication interaction are enhanced.
Owner:GUANGDONG OPERATOR WIRE INTELLIGENT TECHNOLOGY CO LTD

Sign language translation method based on multi-modal guidance and language generator

The invention provides a sign language translation method based on multi-modal guidance and a language generator. The sign language translation method comprises the following steps: step 1, acquiring and preprocessing a multi-modal signal; 2, designing an independent modal encoder; (3) cross-modal fusion is achieved through Q-Former; step 4, high-order cross-modal semantic bridging; step 5, reasoning a perception language generator; step 6, a combined training mechanism; and step 7, generative output and sign language translation, signal input and cross-modal bridge vector generation, language generation initialization and decoder input, autoregression generation process, continuous sign language input and multi-round generation, and translation result output. Based on the technical scheme of the invention, the complementary characteristics of each mode are fully mined, the accuracy and stability of gesture expression are improved, accurate signal-to-language alignment is realized, the logicality and context coherence of the generated language are improved, the common ambiguity and omission problems in sign language-language conversion are solved, and the conversion efficiency is improved. The method is suitable for small-sample and zero-sample sign language translation tasks and has a wide application prospect.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Robot operation method, system and terminal based on visual language model

The invention relates to the field of robot intelligent control, and discloses a robot operation method, system and terminal based on a visual language model, and the method comprises the steps: obtaining a visual image and a natural language instruction, carrying out the combined analysis of the visual image and the natural language instruction through a visual language large model, and generating an initial planning strategy; the robot is controlled to execute operation actions according to the initial planning strategy, and original tactile signals in the operation process are collected; inputting the original tactile signal into a tactile language translation model for feature extraction and semantic mapping to obtain tactile semantic description; when the prompt operation of the tactile semantic description is abnormal or the physical attribute does not accord with the visual expectation, the tactile semantic description is fed back to the visual language large model as an enhanced prompt word; and according to the visual image and the enhanced prompt word, the current task scene is reasoned again, a corrected planning strategy is generated, and the robot is controlled to execute an operation action. The operation precision of the robot is improved.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

Guiding language translation with translation documents using machine learning

In accordance with the described techniques, a system receives a plurality of facets describing language-agnostic aspects of language translation, a translation document describing language-specific rules for translating from a source language to a target language, and a source text in the source language. Using one or more machine learning models, a plurality of guidelines are extracted from the translation document and assigned to respective facets of the plurality of facets. The system translates the source text to a translated text in the target language using one or more machine learning models conditioned on the plurality of guidelines assigned to the respective facets.
Owner:ADOBE INC

Deep learning-based sign language-to-multilingual text speech mutual conversion method and application thereof

The invention relates to a sign language-to-multilingual text speech mutual conversion method based on deep learning and application thereof, belongs to the field of artificial intelligence, man-machine interaction and barrier-free communication, and particularly relates to a sign language recognition, semantic understanding and multilingual speech synthesis method based on deep learning and an application system thereof. The sign language to voice text conversion method comprises the following steps: 1, an image acquisition and preprocessing module, 2, a hand key point detection module, 3, time sequence modeling and sign language recognition, 4, natural language understanding and semantic optimization, 5, a multi-language translation module, and 6, a text and voice synthesis output module. According to the sign language-to-multilingual text voice mutual conversion method based on deep learning and the application thereof, sign languages can be automatically, flexibly and accurately converted into voices and characters, and deaf-mutes are helped to efficiently communicate with common people; the system can be widely applied to hospitals, schools, government affair windows, public transportation and other scenes.
Owner:MINXI VOCATIONAL & TECHN COLLEGE

Language translation data processing method and system based on large model

The invention provides a language translation data processing method and system based on a large model, and the method comprises the steps: carrying out the coding segmentation of language translation data, and obtaining a plurality of text segments; performing dependency syntactic analysis on each text fragment, determining a dependency tree structure of each text fragment, and determining an explicit semantic dependency relationship between every two text fragments according to the semantic similarity between every two text fragments and the dependency tree structure of the corresponding text fragment; performing dependency analysis on an implicit semantic relationship between every two text segments according to the implicit semantic association degree between the text segments and the segmentation loss of the text segments to obtain an implicit semantic dependency relationship between every two text segments; and constructing a paragraph tag of the language translation data through a dependency relationship between explicit semantics and implicit semantics between every two text fragments, and translating the to-be-translated data based on the paragraph tags. By the adoption of the scheme, cross-segment semantic guidance translation of the long complex text can be achieved.
Owner:HUNAN COMM POLYTECHNIC

AUTOMATIC TRANSCRIPT-ASSISTED SPEECH LANGUAGE TRANSLATION USING LANGUAGE MODELS

Devices, systems, and techniques are disclosed that implement the training and deployment of automatic transcription-based translation systems using language models. The techniques include: processing, using a first speech-to-text (S2T) model, an initial input that includes spoken language in a first language to generate a transcription of the spoken language; and processing, using a second S2T model, a second input to generate a translation of the spoken language into a second language. The second input includes at least a representation of the spoken language and the transcription of the spoken language.
Owner:NVIDIA CORP

Sign language translation system and method based on lightweight mask enhancement

The invention discloses a sign language translation system and method based on lightweight mask enhancement, and the method comprises the following steps: 1, employing an inter-frame difference technology to recognize and reject redundant frames having no substantial contribution to motion expression, and further distinguishing the foreground and background information of sign language motion through a local dynamic mask module; step 2, inputting the sign language video optimized by the lightweight mask module to a visual module; accurate and comprehensive visual features are generated; 3, converting the visual features obtained from the visual module into language features suitable for language processing; and step 4, the language module receives the language features converted by the mapping module, and processes the language features in combination with an encoder-decoder model mBART to improve the quality of the output text. According to the method, the accuracy and efficiency of sign language-to-spoken language translation are remarkably improved; the method not only optimizes the use of computing resources, but also greatly enhances the readability and semantic coherence of translation results, and provides powerful technical support for realizing more popular and convenient sign language communication.
Owner:XIDIAN UNIV

Computing technologies for evaluating linguistic content to predict impact on user engagement analytic parameters

Correlations between a set of linguistic features identified in an unstructured text recited in a source language and a set of user engagement analytic parameters may be measured by a machine learning model selected based on a set of performance metrics from a set of machine learning models trained by a set of supervised machine learning algorithms on (i) a set of unstructured texts recited in the source language and containing the set of linguistic features and (ii) the set of user engagement analytic parameters measured for the set of unstructured texts. The machine learning model grades the unstructured text recited in the source language to determine whether the unstructured text recited in the source language should be (1) edited in the source language and then translated into the target language or (2) translated from the source language to the target language as is.
Owner:WELOCALIZE INC

Method and system for translating PDF (Portable Document Format) text containing complex features

The invention belongs to the technical field of text translation, and provides a method and a system for translating a PDF (Portable Document Format) text containing complex features. The method comprises the steps that a PDF analysis engine is initialized, a PDF file is read, and basic information of the PDF file is extracted; judging whether the PDF file is a complex file or not, and recording complex features; preprocessing an image in the complex document, and calling a corresponding text detection model according to the complex features to perform text region identification; and extracting layout information of the text in the text area, translating the text in the text area through a translation model, performing typesetting according to the layout information after a final translation result is obtained, and outputting a target translation file in a user-defined manner. According to the method, the complex content in the PDF document can be intelligently identified, and the original document format is completely reserved in the translation process; meanwhile, multi-language translation is supported, and the accuracy of PDF document translation is guaranteed.
Owner:AFIRSTSOFT CO LTD

Portable multi-language intelligent acquisition and translation system

The invention relates to the technical field of language translation, and discloses a portable multi-language intelligent acquisition and translation system, which comprises a multi-mode acquisition module used for acquiring audio signals and video streams; the environment sensing and scheduling module is used for evaluating the complexity of the current environment so as to dynamically adjust computing resources; the target voice construction module is used for determining a current speaker and extracting a target audio stream of the current speaker, and generating a voice text based on the target audio stream; the analysis module is used for generating visual situation metadata; the translation module is used for generating a preliminary translation result and a corresponding translation confidence score; and the interactive output module is used for carrying out ambiguity clarification to generate a final translation result. According to the invention, audio-visual fusion is carried out on the audio signal and the video stream, and the current speaker is determined in a multi-person noisy environment in combination with the face position and lip movement information of the video stream, so that the interference of background noise and other non-target speakers is eliminated, and the accuracy of subsequent speech recognition and translation is improved.
Owner:CHANGJI UNIV

Visual sign language translation training device and method

Methods, devices and systems for training a pattern recognition system are described. In one example, a method for training a sign language translation system includes generating a three-dimensional (3D) scene that includes a 3D model simulating a gesture that represents a letter, a word, or a phrase in a sign language. The method includes obtaining a value indicative of a total number of training images to be generated, using the value indicative of the total number of training images to determine a plurality of variations of the 3D scene for generating of the training images, applying each of plurality of variations to the 3D scene to produce a plurality of modified 3D scenes, and capturing an image of each of the plurality of modified 3D scenes to form the training images for a neural network of the sign language translation system.
Owner:AVODAH INC