Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

42 results about "Language recognition" patented technology

Language Recognition. Summary: The goal of the NIST Language Recognition Evaluation (LRE) series is to establish the baseline of current performance capability for language recognition of conversational telephone speech and to lay the groundwork for further research efforts in the field.

Display device and subtitle language identification method

The application provides a display device and a subtitle language recognition method. The method can acquire a subtitle corresponding to a media video in response to a play instruction of the media video, extract visual texture features of the subtitle in a case where the subtitle is an image subtitle, perform language recognition according to the visual texture features to obtain a language recognition result, extract syntax topology features of the subtitle in a case where the subtitle is a text subtitle, perform language recognition according to the syntax topology features to obtain a language recognition result, write the language recognition result into metadata corresponding to the media video, control a display to play the media video and the subtitle, and display a language identifier corresponding to the subtitle according to the language recognition result in the metadata corresponding to the media video. The method recognizes a language by analyzing features of different languages in physical forms and syntax structures, so that the language identifier of the subtitle can be correctly displayed when the media video is played.
Owner:HISENSE ELECTRONICS TECH SHENZHEN CO LTD

An AI-assisted large-scale lip reading dataset automatic construction method

The application relates to an AI-assisted large-scale lip language recognition data set automatic construction method, which comprises the following steps: S1: a distributed automatic crawler system is constructed to capture video materials; S2: the collected original video is preprocessed; S3: the video stream is subjected to shot boundary detection, different shot segments are segmented, and a pre-trained SyncNet model is used for audio and video synchronization alignment; S4: the audio stream is subjected to text recognition, and one-to-one correspondence between a sentence and an audio timestamp is realized to obtain a structured sentence-level subtitle data set; S5: an MTCNN algorithm is used for face detection, and the speaker in the video is subjected to identity distinction and clustering; S6: face detection data is obtained according to the MTCNN algorithm, lip key point data is obtained, and a corresponding region of interest (ROI) is extracted. The application effectively improves the construction efficiency and quality of the lip language data set.
Owner:FUZHOU UNIV

Sign language recognition method and device based on cross-stage focus distillation and multi-branch temporal learning

ActiveCN122049994BLearning unitAlgorithm
The application discloses a sign language recognition method and device based on cross-stage focus distillation and multi-branch time sequence learning, acquires a sign language video to be recognized, and inputs the sign language video to a trained sign language recognition model to obtain a recognition result; the method and the device realize complementary deep and shallow features through a cross-stage focus distillation unit, and simultaneously capture time sequence dynamics of different granularities with the help of a multi-branch time sequence learning unit, so that the overall modeling capability for key semantics and action evolution in a continuous flow of sign language is improved; the cross-stage focus distillation unit guides the sign language recognition model to focus on key discrimination areas in the sign language video through selective reinforcement of channel and spatial dimensions, effectively solving the contradiction between incomplete shallow feature semantics and lost deep feature fine-grained information; and the multi-branch time sequence learning unit models short-time actions at the word level and long-range dependencies at the sentence level synchronously through a parallel multi-branch convolution structure, overcoming the deficiency of a traditional time sequence network in capturing multi-granularity time dynamics.
Owner:ZHEJIANG UNIV OF TECH

Real-time gesture recognition system based on 16-way imu array and spatiotemporal feature fusion and application

The application discloses a kind of real-time gesture recognition systems based on 16-way IMU array and space-time feature fusion, including data acquisition module, data preprocessing module, space-time feature extraction module and real-time identification display module.The IMU array is composed of 16 six-axis inertial measurement units, and 96-dimensional original motion data is continuously sent in fixed format by UDP protocol;Data preprocessing module uses Euclidean alignment and sliding window segmentation technology, window slicing is carried out according to 30 frame length, 15 frame step, and 96-dimensional data is reconstructed into 6×16×30 three-dimensional tensor.The space-time feature extraction module supports ListenNet, DARNET and three kinds of deep network structures based on multi-channel time sequence patch visual Transformer (ViT), realizes the fusion modeling of local structure feature and global time sequence dependence.Has the advantages of strong data structure retention, excellent cross-subject generalization ability, simple deployment, high real-time performance, etc., applicable to wearable interaction, robot control, sign language recognition and other scenes.
Owner:TONGJI UNIV

A small language-oriented audio and video subtitle optimization generation method and system

PendingCN122313949ASoutheast asiaSpoken language
This invention relates to the field of artificial intelligence technology and discloses a method and system for optimizing and generating audio and video subtitles for less commonly spoken languages. The method includes the following steps: Step 1: Audio extraction and preprocessing; Step 2: Language recognition for the less commonly spoken language; Step 3: Language-aware punctuation restoration module; Step 4: Subtitle readability-driven segmentation module; Step 5: Subtitle format encapsulation and output. This application establishes a complete technical process encompassing audio preprocessing, language-aware speech recognition, structured text restoration, subtitle segmentation optimization, and time alignment. This method significantly improves the accuracy of speech-to-text conversion in less commonly spoken language videos and enhances the structural integrity and readability of subtitles. It is particularly suitable for practical application scenarios involving less commonly spoken languages ​​in regions such as Southeast Asia, where data is scarce and languages ​​are diverse. It has broad application value and practical significance in media dissemination, educational videos, and government services.
Owner:XINGZHOU DIGITAL TECH (ZHUHAI) CO LTD

A multi-modal real-time sign language recognition method and system in a medical scenario

PendingCN122336848AFeature extractionMedicine
This invention discloses a multimodal real-time sign language recognition method and system in medical scenarios. The recognition method includes the following steps: 1) acquiring a continuous multi-frame sign language image sequence; 2) annotating the hand skeletal joints of the sign language image sequence; 3) recognizing the annotated sign language image sequence through a spatiotemporal feature extraction module to obtain a gesture recognition sequence; 4) inputting the gesture recognition sequence into a language reasoning module for medical semantic reconstruction, outputting structured and highly readable Chinese medical terms or sentences. This invention achieves high accuracy and low latency recognition of medical terminology-level sign language by integrating visual spatiotemporal features with language-level medical semantic reasoning, realizing an effective connection from low-level visual perception to high-level medical semantic understanding, and providing reliable technical support for efficient communication for hearing-impaired patients in real medical scenarios.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Multilingual speech recognition model training and speech recognition method

The application discloses a kind of multilingual speech recognition model training and speech recognition method.The language recognition layer in the speech recognition model to be trained is based on a plurality of hidden feature calculations to obtain predicted language probability distribution, realize language classification discrimination, provide basis for subsequent inference stage routing specific language expert module.Each target language is configured with a specific language expert module, to achieve customized recognition, taking into account multilingual versatility and monolingual recognition accuracy.With the help of the standard language probability distribution generated by the pre-trained teacher model, the first loss is used to constrain the prediction output of the language recognition layer, so that the speech recognition model to be trained can indirectly learn the generalization ability of the teacher model by aligning the standard language probability distribution, reducing the recognition bias caused by insufficient sample coverage in the training process itself, and realizing the knowledge distillation of the language recognition ability of the teacher model.
Owner:SHANGHAI NORMAL UNIVERSITY +1

A sign language recognition method and related apparatus, device, and storage medium

The application discloses a sign language recognition method and related device, equipment and storage medium. The sign language recognition method comprises the following steps: acquiring a sentence video frame sequence, the sentence video frame sequence is obtained by collecting a sign language action sequence, and the content expressed by the sign language action in the sentence video frame sequence is a sentence; dividing the sentence video frame sequence according to word segmentation to obtain a plurality of word video frame sequences, the content expressed by the sign language action in the sentence video frame sequence is a word; performing action recognition on each word video frame sequence to obtain a word corresponding to each word video frame sequence; and obtaining a sentence corresponding to the sentence video frame sequence by using the word corresponding to each word video frame sequence. The above scheme can improve the communication efficiency of video calls.
Owner:IFLYTEK CO LTD +1

A dialect language recognition method based on multi-modal fusion

This invention provides a dialect language recognition method based on multimodal fusion. The method includes the following steps: S101, collecting first multimodal data of dialect language and Mandarin data, and preprocessing them; S102, extracting features from the first multimodal data and Mandarin data to obtain dialect language features and Mandarin features; S103, performing feature compatibility processing on the dialect language features and Mandarin features to obtain a feature association dataset; S104, constructing a dialect language recognition model and training the model using the feature association dataset; S105, collecting second multimodal data of the dialect language to be recognized, extracting multimodal features from the second multimodal data, inputting the extracted multimodal features into the trained dialect language recognition model, and outputting the corresponding dialect recognition result. This invention improves the model's learning and understanding of dialect knowledge, and enhances the accuracy and robustness of dialect recognition.
Owner:海南经贸职业技术学院

APPROACHES TO EXPANDING OUTPUT FROM LANGUAGE RECOGNITION

ActiveDE602023019467T2Extension languageStructural engineering
Owner:PALANTIR TECHNOLOGIES INC AVENTURA

A speech synthesis method, device, computer equipment and storage medium

Embodiments of the present disclosure relate to the technical field of speech processing, and specifically relate to a speech synthesis method and device, computer equipment and a storage medium. The main steps of the foregoing method include: obtaining text data to be synthesized, performing language recognition on the text data, and determining at least one target language category to which the text data belongs. Based on the at least one target language category, the text data is subjected to text normalization processing to obtain normalized text. The unified phoneme sequence corresponding to the normalized text is input into a pre-trained shared acoustic model to obtain target acoustic features; and based on the target acoustic features, synthesized speech is obtained. By using a unified phoneme sequence defined based on cross-language pronunciation features as an intermediate identifier, a single shared acoustic model is used for processing, which significantly reduces the overall parameter quantity, storage space requirement and computing resource consumption of the system, thereby directly reducing the deployment and operation costs.
Owner:BEIJING BAILONG MAYUN TECH CO LTD

Automatic switching between languages during virtual conferences

In some aspects, a computing device may access audio information comprising an audio stream from a client device, and a source-language. The computing device may provide an audio segment from the audio stream to a language identification process of the computing device comprising a machine learning model that is trained to identify a language of a plurality of languages within recorded speech. The computing device may identify an identified-language of the plurality of languages for the speech based at least in part on the audio segment. The computing device may update the source-language to the identified language. Numerous other aspects are described.
Owner:ZOOM VIDEO COMM INC

A sign language translation glove and a sign language recognition method

This invention discloses a sign language translation glove and a sign language recognition method, including a glove body, a curvature sensor disposed on the fingers, an inertial sensor disposed on the back of the hand, a data processing module, and an output module. The data processing module extracts hand shape features based on finger bending data collected by the curvature sensor, extracts motion features based on hand motion data collected by the inertial sensor, and extracts coupling features based on the relative temporal relationship between hand shape features and motion features. Then, based on the hand shape prototype, motion prototype, and coupling prototype, a joint decision is made to output the sign language recognition result. This invention can improve the ability to distinguish similar gestures and supports the rapid expansion of new sign language words.
Owner:任永富

Degradation-robust quality-aware adaptive fusion continuous sign language recognition method

The present application relates to the technical field of continuous sign language recognition, and particularly provides a quality-aware adaptive fusion continuous sign language recognition method for degradation robustness. The method comprises the following steps: extracting an RGB frame sequence and a skeleton key point sequence; obtaining a video frame sequence of training input; inputting the video frame sequence into an RGB encoding branch and inputting the skeleton key point sequence into a skeleton encoding branch to obtain frame-level features extracted respectively; performing time sequence convolution on the frame-level features to obtain time-dimensionally aligned RGB time sequence features and skeleton time sequence features; using a quality-aware adaptive fusion module QAAF to estimate learnable fusion weights according to the RGB time sequence features and obtaining fused time sequence features; inputting the fused time sequence features into a BiLSTM network and a CTC module to output a sign language vocabulary sequence. The method improves the robustness and accuracy of the continuous sign language recognition system in the actual deployment environment.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

system

We provide the system. [Solution] An audio acquisition means that acquires audio during a business negotiation and generates audio data, Language recognition means that analyzes the aforementioned audio data and converts it into text data, A generation mechanism that extracts key points based on the aforementioned text data and generates meeting minutes, A document input means for inputting the generated meeting minutes into an information processing system, A descriptive analysis means that analyzes memo data entered after a business negotiation, converts it into detailed information, and records it in the information processing system, A presentation means that processes data during a business negotiation in real time and displays the results on a visual device that possesses it, A system that includes this.
Owner:SOFTBANK GROUP CORP

A sign language translation method, system, electronic device and storage medium

ActiveCN122223786BHigh precisionAchieve collaborative translationSemantic alignmentFeature extraction
A sign language translation method, system, electronic device and storage medium relate to the field of video natural language generation, and alleviate the problems that existing sign language translation technology cannot consider gesture dynamic change and semantic information. A sign language translation method comprises a video processing step, uniformly sampling an RGB sign language video to be recognized to obtain a visual image sequence; a visual feature extraction step, obtaining visual features by passing the visual image sequence through a visual encoder; a language feature extraction step, encoding a plurality of predetermined sign language labels to obtain a plurality of corresponding text semantic features; a semantic alignment step, aligning the text semantic features and the visual features to obtain aligned text semantic features and visual features; and a translation step, completing the translation of the RGB sign language video to be recognized. The present application is suitable for the field of sign language recognition and translation, and is especially suitable for the field of semantic understanding of continuous and natural sign language videos.
Owner:CHANGCHUN UNIV

Continuous sign language recognition method and system based on hierarchical attention feature fusion

This invention proposes a continuous sign language recognition method and system based on hierarchical attention feature fusion. The method involves processing sign language videos, outputting multi-level, multi-scale feature maps through a visual encoder. These maps are then processed by a cross-scale aggregation module containing a hierarchical attention network before being fed into a classifier for word classification. The CTC decoder outputs a word sequence. Simultaneously, a two-layer fully connected network performs binary classification prediction to determine whether each frame is a boundary frame for a sign language word. An online self-distillation mechanism based on exponential moving average (EMA) is established, using an exponential moving average version of the model parameters as the EMA teacher model. The student model continuously learns to obtain a more stable word sequence output. A multi-level regularization framework is constructed, coordinating optimization at four levels: label space, temporal space, feature space, and parameter space. This invention effectively improves the model's generalization ability under small sample conditions, thereby reducing the word error rate (WER) in continuous sign language recognition.
Owner:TIANJIN POLYTECHNIC UNIV

A lightweight adaptive temporal hybrid continuous sign language recognition method and apparatus

The application discloses a kind of lightweight adaptive timing hybrid continuous sign language recognition method and device, using trained continuous sign language recognition network model to carry out sign language recognition, the continuous sign language recognition network model includes feature extraction module, timing construction module and classification module, the feature extraction module uses lightweight convolutional neural network, the lightweight convolutional neural network includes several stages for carrying out feature extraction, for the stage feature extracted in each stage, mixed spatial feature and mixed timing feature are extracted respectively, mixed spatial feature and mixed timing feature are weighted and fused, then residual connection is carried out with stage feature, to obtain the feature finally extracted in current stage, input to next stage.The application can better capture the motion information of sign language, so as to maintain higher recognition accuracy under different environments and conditions.
Owner:ZHEJIANG UNIV OF TECH

A sign language recognition method based on hand key point fusion and language model

This invention discloses a sign language recognition method based on hand keypoint fusion and a language model, comprising the following steps: A: acquiring preprocessed image frame sequence data containing hand movements; B: constructing a set of hand detection boxes and a set of hand keypoints based on a hand keypoint detection network; C: acquiring the corresponding hand keypoint feature vector in each frame; D: identifying the number of hands in the image to obtain sign language word markers; E: constructing a stable sign language word cache sequence to be output based on the acquired sign language word markers; F: dividing the stable sign language word cache sequence to be output based on the duration of no-hand movement in the image to obtain complete semantic units for output; G: generating natural language sentences based on a language model and combining the complete semantic units to complete the output of sign language sentences. This invention can achieve stable output of continuous sign language into natural language sentences and is suitable for deployment in embedded devices.
Owner:ZHENGZHOU UNIV

A sign language recognition interaction system based on end-cloud cooperation

This invention provides a sign language recognition and interaction system based on edge-cloud collaboration, belonging to the field of artificial intelligence technology. It includes an image acquisition module for acquiring images of gestures from deaf and mute individuals and extracting keyframe images from these images; a client processing module, wirelessly connected to the image acquisition module, for receiving keyframe images, extracting key hand points from the keyframe images, and transmitting the extracted key hand point sequence to the cloud; a cloud processing module for processing the transmitted key hand point sequence using a sign language recognition model, recognizing the gestures as text based on the key hand point sequence; the cloud processing module also transmits the text to the client processing module, which in turn performs voice playback of the text and simultaneously acquires the voice response from a hearing person, converting the voice response into text for display. This system improves the portability and accuracy of sign language translation, reduces production costs, and improves the quality of life for deaf and mute individuals.
Owner:SHANDONG WOMENS UNIV

Realtime ai sign language recognition

PendingUS20260188141A1AcousticsLanguage recognition
A real time sign language recognition method that allows Deaf and Hard of Hearing individuals to sign into any apparatus with a camera to extract target information (such as a translation in a target language) is proposed.
Owner:SIGN-SPEAK INC

DISAMBIGUING OF LANGUAGE RECOGNITION COMMANDS FOR A VEHICLE

ActiveDE102017109097B4Speech soundSubvocal recognition
A method for speech recognition in a vehicle (12), comprising the steps of: (a) receiving a voice command at the vehicle (12) via a microphone (32) in the vehicle (12); (b) obtaining a recognition result from speech recognition performed on the received voice command, wherein the recognition result represents the voice command and indicates each of two or more available vehicle commands; and (c) selecting one of the two or more available vehicle commands based on a secondary feature and an attribute of the selected vehicle command, characterized in that the secondary feature includes a time at which the voice command is received at the vehicle (12), and wherein selecting one of the two or more available vehicle commands includes comparing the time with an expected availability period associated with the selected vehicle command.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Sign language recognition methods, devices and equipment

This invention provides a sign language recognition method, apparatus, and device, relating to the field of image recognition technology. The method includes: acquiring a sequence of image frames to be recognized, and determining skeletal detection results based on the image frame sequence; inputting the image frame sequence and skeletal detection results into a pre-trained sign language recognition model to obtain the sign language recognition result corresponding to the image frame sequence; wherein the sign language recognition model includes a cross-modal attention fusion module; the cross-modal attention fusion module is used to calculate a first semantic enhancement result of the visual feature vector on the skeletal feature vector, and a second semantic enhancement result of the skeletal feature vector on the visual feature vector, and based on the first and second semantic enhancement results, determining a frame-level prediction probability sequence; the frame-level prediction probability sequence is used to determine the sign language recognition result. This invention can fully leverage the advantages of multimodal feature complementarity, effectively improving the accuracy of sign language recognition.
Owner:HEBEI UNIVERSITY OF ECONOMICS AND BUSINESS

A method for adaptive selection of embedding vector models based on document language identification

This application relates to an adaptive selection method for embedding vector models based on document language recognition. The method includes: preprocessing sample text from document data; performing language statistical processing to generate corresponding language proportion indices; classifying the document data according to the language proportion indices to determine the language category and selecting the corresponding embedding vector model; performing vector encoding on the document data based on the embedding vector model to generate corresponding document vector data and constructing a corresponding vector index library; recording the embedding vector model corresponding to the document vector data; performing vector encoding on the query text based on the embedding vector model to generate corresponding query vectors; routing the query vectors to the corresponding vector index library for vector similarity retrieval. This method reduces model complexity and computational overhead while ensuring retrieval accuracy, achieving flexible adaptation in multilingual scenarios and ultimately striking a balance between accuracy, efficiency, and system scalability.
Owner:SHENZHEN CHUANGZHICHENG TECH CO LTD

Sign language recognition method, apparatus, and device

PendingCN122157369ACharacter and pattern recognitionTeaching apparatusLanguage recognition
The application discloses a sign language recognition method and device and equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: in response to a first input of a user, a first sign language video is acquired; the first sign language video is input into a gesture understanding model, and gesture features of each video frame in the first sign language video are extracted; the gesture understanding model is trained based on a plurality of training samples, each training sample comprises a plurality of sign language video samples corresponding to a vocabulary sample; and a sign language recognition result is determined according to the gesture features.
Owner:VIVO MOBILE COMM HANGZHOU CO LTD

Sign Language Recognition Method Based on Multi-Level State Feature Optimization

This invention relates to the fields of computer vision and artificial intelligence, and discloses a sign language recognition method based on multi-level state feature optimization. The method includes acquiring a sign language video sequence and extracting original visual features; constructing a semantic potential energy field, calculating the potential energy value of the current frame features and their gradient vector relative to the feature space; constructing a local orthogonal basis based on the gradient vector, and orthogonally decoupling the features into normal transition components along the gradient direction and tangential steady-state components along the equipotential surface direction; performing differential evolution processing on the normal transition components based on multi-level dynamic states, using graph networks to repair transition features, and smoothing the tangential steady-state components; finally, reconstructing features and temporally decoding the processed components to obtain the sign language text. This invention utilizes the geometric properties of the potential energy field to achieve physical-level decoupling between core semantics and transitional actions, effectively solving the feature drift problem caused by action switching in continuous sign language, and significantly improving recognition accuracy.
Owner:GUANGDONG POLYTECHNIC COLLEGE

A navigation method for identifying and correcting place names, a device, an electronic device and a medium

The present application relates to a kind of navigation identification place name correction method, device, electronic equipment and medium, belong to speech recognition navigation technical field, wherein, the method includes: based on the complete intent detection model of training, the classification of the language recognition result obtained, obtain user intent type, user intent type includes no correction intent and correction intent, intent detection model uses large language model;When user intent type is correction intent, extract initial place name from language recognition result, based on initial place name, in the matching of pre-set geographic knowledge base, obtain target place name, the target place name is as the final navigation identification place.This application carries out intent speculation to user speech recognition data by intent detection model, only when user has correction intent, subsequent rewriting operation is carried out, user no correction intent does not carry out any rewriting operation, avoid unnecessary computing overhead and additional time-consuming.
Owner:AISPEECH CO LTD

A Multilingual Adaptation Method for Used Car Web Applications Targeting Overseas Users

This application discloses a multilingual adaptation method for a used car web application for overseas users, including: responding to a web application access request from an overseas user; obtaining a target language identifier using a progressive language recognition method based on the web application access request; detecting whether the user performs a language switching operation in the web application; if the user performs a language switching operation, synchronously writing the target language identifier after switching to local storage and the root domain cookie; determining an interface layout strategy based on the target language identifier; loading the corresponding language resource file based on the target language identifier, and loading page code and image resources according to the interface layout strategy. The method provided by this application enables users to automatically inherit language settings when switching between the main site and overseas sites, eliminating the need for repeated language settings, and ensuring that the same interactive component can correctly respond to user operations in both left-to-right and right-to-left layout environments, guaranteeing the consistency of interaction logic.
Owner:BEIJING HAOCHA SHUAISHUAI TECHNOLOGY CO LTD