Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8738 results about "Speech recognition" patented technology

Speech recognition is a interdisciplinary subfield of computational linguistics that develops methodologies and technologies that enables the recognition and translation of spoken language into text by computers. It is also known as automatic speech recognition (ASR), computer speech recognition or speech to text (STT). It incorporates knowledge and research in the linguistics, computer science, and electrical engineering fields.

Assistant System Using Multimodal Multitask Medical Machine-Learned Models to Perform Image Processing to Answer Natural Language Queries

An example assistant system can use a multimodal multitask medical machine-learned model to perform image processing to answer natural language queries. A device can process speech data or other natural language inputs to obtain a query. The query can be processed alongside image data that provides context for the query. The example system can receive a query associated with a particular task domain; generate, based on the query, a query input that comprises query instruction data from a first modality and query context data from a second modality; generate a combined input comprising the query input and an exemplar input, wherein the exemplar input comprises exemplar instruction data from the first modality and an exemplar context placeholder in lieu of exemplar context data from the second modality; process the combined input with a multimodal machine-learned model to generate output data; and output a query response based on the output data.
Owner:GOOGLE LLC

Structure fatigue damage identification method based on acoustic emission and deep learning

The invention relates to the technical field of structural health monitoring and intelligent diagnosis, in particular to a structural fatigue damage identification method based on acoustic emission and deep learning, and the method comprises the steps: collecting a structural response signal under a fatigue load through an acoustic emission sensor array, inputting the structural response signal to a CNN-BiLSTM-Attention mixed deep learning model, and carrying out the recognition of the structural fatigue damage through the CNN-BiLSTM-Attention mixed deep learning model; the model extracts local time domain features through a dynamic adaptive convolution kernel, captures long time sequence dependence by using a bidirectional long-short-term memory network, focuses key damage features through a bimodal space-time attention mechanism, divides damage stages based on a nonlinear dynamic threshold algorithm of fracture opening amount, constructs a training data set of physical-data fusion, and performs dynamic time domain feature extraction. The learning rate is optimized by adopting a gradient sensitive cosine annealing algorithm, and the robustness of the model is improved in combination with an anti-noise and anti-loss function. The method integrates physical characteristics and an intelligent algorithm, and has the advantages of adaptive noise suppression, strong cross-domain generalization ability, high real-time performance and the like.
Owner:FUJIAN UNIV OF TECH

Electrical audio signal processing systems and devices

According to an aspect of the present invention, there is provided an electrical audio signal processing system and device, comprising: a computer graphics processing and selective visual display system with a screen; an eye tracking device; a processor; one or more computer memory devices; wherein the processor is arranged for operations comprising: measuring the user's eye movements to ascertain the specific word on which the user is fixated, by the eye tracking device; modifying the display at the user's current fixation point; applying a delay between the presentation of successive graphic elements based on the user's calculated rate to accommodate the user's required time; and presenting elements to the user at a rate based upon the user's required time.
Owner:DECHARMS RICHARD CHRISTOPHER

Audio noise reduction method, device and system based on deep learning

The invention relates to an audio noise reduction method, device and system based on deep learning, and the method comprises the steps: obtaining an input audio signal with noise, and carrying out the multi-scale time-frequency decomposition, and obtaining a mixed time-frequency feature and a noise fingerprint spectrum; performing parameter parallel processing on the noise fingerprint spectrum through a preset dynamic kernel generation network, and performing preliminary noise reduction processing on the mixed time-frequency characteristics to obtain noise-reduced mixed data; performing dual-path processing structure construction on the noise reduction mixed data to obtain amplitude optimization data and phase optimization data; performing dynamic time-frequency domain cross fusion on the amplitude optimization data and the phase optimization data to obtain fused audio data; and carrying out differentiable acoustic equation constraint adversarial training on the fused audio data, and carrying out inverse time-frequency transformation processing to obtain a target noise-reduced audio signal. According to the invention, the overall efficiency and effect of audio signal processing can be effectively improved.
Owner:DONGGUAN HUAZE ELECTRONIC TECH CO LTD

Real-time video translation and audio and picture synchronization method and system based on multi-modal large model

The invention provides a real-time video translation and audio and picture synchronization method and system based on a multi-modal large model, and relates to the technical field of video translations, and the method comprises the steps: obtaining a source video; extracting the source video based on the multi-modal large model to obtain a multi-modal feature; fusing the multi-modal features through a cross-modal attention mechanism to generate a context semantic vector; translating into a target language text in real time based on the context semantic vector, and processing the translated language text based on the multi-modal features to obtain a translated language sound source; and performing mouth shape adjustment on the source video based on the translation language sound source, and merging the translation language sound source and the mouth shape animation video to obtain a real-time translation video with synchronous sound and picture. According to the method, the limitation of traditional single-modal translation is broken through, and the semantic accuracy of translation is remarkably improved by dynamically aligning the context information through the multi-modal features in combination with a cross-modal attention mechanism.
Owner:SHANGHAI YINGZHUO INFORMATION TECH CO LTD

Method for realizing synchronization of character expression and lip shape in video through emotion in sound and cloned digital human system

The invention discloses a method for realizing synchronization of a character expression and a lip shape in a video through an emotion in sound and a cloned digital human system, and the method for realizing synchronization of the character expression and the lip shape in the video through the emotion in the sound comprises the steps: collecting audio information, and extracting multi-dimensional features of the sound; applying a pre-trained expression and lip shape generation model to generate corresponding expression parameters and lip shape parameters according to the multi-dimensional features of the sound; performing fusion processing on the expression parameters and the lip shape parameters, and generating a continuous animation sequence according to the fused parameters; and rendering the continuous animation sequence to generate video information with expressions, lip shapes and sound emotion synchronization. The high-precision expression and lip shape synchronization is realized, the natural fidelity of the cloned digital human is improved, and the application range and value of the cloned digital human are expanded.
Owner:SHANGHAI YANTU TECHNOLOGY CO LTD

Dance video generation method based on multi-mode music driving and frequency domain-space double-flow decomposition

The invention provides a dance video generation method based on multi-mode music driving and frequency domain-space double-flow decomposition, and the method comprises the steps: extracting multi-granularity music features through a composite encoder (Librosa + Jukebox), employing a beat gating attention mechanism, enabling the key actions such as dancing hand raising, kicking and the like to be strictly aligned with a music re-beat point, and enabling the synchronization error to be reduced to 118 ms through the verification of a test data set; for the problem of visual detail loss, a frequency domain-space double-flow decomposition architecture is provided, a Butterworth filter bank is used to decouple a reference image into a low-frequency energy diagram and a high-frequency residual error, and a double-flow diffusion mechanism is used to optimize a global attitude and local details respectively; a joint confidence prediction module is introduced for the generation stability in a shielding scene, and the motion trail of an abnormal joint point is dynamically corrected through a time domain sliding window weighted fusion strategy, so that a reasonable action conforming to ergonomics can still be generated under a 50% limb shielding rate.
Owner:湖南马栏山视频先进技术研究院有限公司

Name-detection based attention handling in active noise control systems

PCT designated stage expiredWO2025128140A1Sound producing devicesSpeech recognitionControl systemNoise
Automated attention handling techniques are described herein for use with wearable audio components with active noise control (ANC) to suppress ambient sound. A name embedding model is trained automatically to convert name audio samples into acoustic segments based on a knowledge distillation model. The name embedding model is used to generate reference embeddings for each of a user-enrolled set of names, and a relation network and a false rejection network are also trained. In real-time operation, the name embedding model converts real-time audio samples to real-time embeddings, the relation network compared the real-time embeddings to the reference embeddings to look for candidate matches, and the false rejection network validates the candidate matches to detect when one of the user-enrolled names has been invoked. Detecting such an invocation automatically triggers the ANC to switch to a conversation mode.
Owner:GOOGLE LLC

Systems and methods for generating an equal-loudness contour response using an auricular device

A system may include a storage device, configured to store computer-executable instructions. A system may include an ear-bud configured to be positioned within an ear canal of a user, the ear-bud comprising: a speaker, a microphone; and one or more processors in communication with the storage device, wherein the computer-executable instructions, when executed by the one or more processors, cause the one or more processors to: obtain a user hearing profile, obtain an equal-loudness hearing profile, receive audio data from the microphone, and generate a second audio data based on a first sound-pressure level, a second sound-pressure level, a first frequency; and cause the speaker to emit the second audio data within the ear canal of the user, such that the user perceives the audio data as if the user has normal hearing.
Owner:MASIMO CORP

Automated attention handling in active noise control systems based on linguistic name embedding

PCT designated stage expiredWO2025128138A1Sound producing devicesSpeech recognitionControl systemNoise
Automated attention handling techniques are described herein for use with wearable audio components with active noise control (ANC) to suppress ambient sound. A name embedding model is trained automatically to convert name audio samples into linguistically distinct name classifications and / or unified audio samples. The name embedding model is used to generate reference embeddings for each of a user-enrolled set of names, and a relation network and a false rejection network are also trained. In real-time operation, the name embedding model converts real-time audio samples to real-time embeddings, the relation network compared the real-time embeddings to the reference embeddings to look for candidate matches, and the false rejection network validates the candidate matches to detect when one of the user-enrolled names has been invoked. Detecting such an invocation automatically triggers the ANC to switch to a conversation mode.
Owner:GOOGLE LLC

Audio watermark generation method and device and computer storage medium

The invention provides an audio watermark generation method, audio watermark generation equipment and a computer storage medium. The audio watermark generation method comprises the following steps: acquiring target audio data; extracting amplitude spectrum features and phase spectrum features of the target audio data; acquiring watermark adding frame information of the amplitude spectrum features; selecting one watermark information code from a watermark codebook library, and performing frame-level copying according to the frame number of the amplitude spectrum characteristics to obtain first watermark data; performing frame selection on the first watermark data according to the watermark adding frame information to obtain second watermark data; fusing the amplitude spectrum feature with the second watermark data to obtain a watermark amplitude spectrum feature; and fusing the watermark amplitude spectrum feature and the phase spectrum feature to obtain watermark audio data. According to the audio watermark generation method, the watermark adding selector is designed, the frame number and the frame number needing to be added with the watermark are selected through the frame selection logic, and it is guaranteed that after the audio watermark is generated, the audio listening feeling is not affected.
Owner:IFLYTEK CO LTD

Ontology brain wave audio auditory perception synchronous feedback method, device and system and electronic equipment

The invention relates to the technical field of electroencephalogram signal processing, in particular to an ontology brain wave audio auditory perception synchronous feedback method, device and system and electronic equipment. The method comprises the following steps: receiving electroencephalogram signals of a collected user on line from electroencephalogram collection equipment through an upper computer, obtaining electroencephalogram signals of a specified frequency band from the electroencephalogram signals, and extracting corresponding electroencephalogram characteristics; according to the electroencephalogram features or preset rhythm parameters, the electroencephalogram signals of the specified frequency band are segmented into a plurality of electroencephalogram segments, a plurality of audio expressions corresponding to the electroencephalogram segments are generated, and feature parameters of the audio expressions are determined according to the electroencephalogram features of the corresponding electroencephalogram segments; generating brain wave audio representation data according to the audio representation corresponding to the electroencephalogram signals of the one or more designated frequency bands, and obtaining brain wave audio according to the brain wave audio representation data. Therefore, the physiological suitability of nerve regulation and control and the artistic expressivity of audio generation are met at the same time, and organic unification of nerve regulation and control and audio generation is achieved.
Owner:WEIZHINAO DATA SERVICE (TIANJIN) CO LTD +1

Automated attention handling in active noise control systems based on universal sound conversion

PCT designated stage expiredWO2025128139A1Sound producing devicesSpeech recognitionControl systemNoise
Automated attention handling techniques are described herein for use with wearable audio components with active noise control (ANC) to suppress ambient sound. A name embedding model is trained automatically to convert name audio samples into linguistically distinct name classifications and / or unified audio samples. The name embedding model is used to generate reference embeddings for each of a user-enrolled set of names, and a relation network and a false rejection network are also trained. In real-time operation, the name embedding model converts real-time audio samples to real-time embeddings, the relation network compared the real-time embeddings to the reference embeddings to look for candidate matches, and the false rejection network validates the candidate matches to detect when one of the user-enrolled names has been invoked. Detecting such an invocation automatically triggers the ANC to switch to a conversation mode.
Owner:GOOGLE LLC

System and Methods for Upsampling of Decompressed Audio Data Using a Neural Network

A computer system for upsampling decompressed audio data after lossy compression using specialized neural network techniques. The system processes compressed audio channels through an audio pre-processor that extracts spectral information, detects speech activity, segments audio, and normalizes input levels. A trained deep learning algorithm with multi-channel transformers using channel-wise and self-attention mechanisms recovers information lost during compression. The system further enhances audio quality through a time-frequency domain transformer applying Fourier transforms and Mel-scale frequency processing, while a perceptual quality assessor employing psychoacoustic models evaluates the output. This specialized audio processing approach significantly improves reconstructed audio quality by leveraging correlations between audio channels, addressing both spectral and temporal features, and optimizing for human perception characteristics, resulting in higher fidelity audio reproduction from compressed formats.
Owner:ATOMBEAM TECH INC

Robot sign language communication method, related device and storage medium

The invention discloses a robot sign language communication method, a related device and a storage medium. The method comprises the steps of collecting each frame of gesture image of a current user and performing feature extraction; inputting the features of each frame of gesture image into a trained gesture recognition model, and recognizing a current gesture recognition result; based on the mapping relation between the gesture symbols and vocabularies, mapping the current gesture recognition result to obtain a vocabulary sequence; forming the vocabulary sequence into a current complete recognition text according to a preset grammar rule; correcting the current complete recognition text by combining the multi-modal large language model with the current complete recognition text, the current scene understanding information and the current action understanding information; inputting the corrected current complete recognition text, the current scene picture and the historical communication information into a third visual language model, and analyzing a current response text; converting the current response text into a current sign language sequence; and driving the robot to simulate sign language actions according to the current sign language word order.
Owner:DIGITAL HUAXIA (SHENZHEN) TECHNOLOGY CO LTD

Learning to compress prompt in natural language formats

Methods, systems, and apparatuses for performing natural language prompt compression, the method being performed by an electronic device and including: obtaining a prompt for an artificial intelligence (AI) inference model, wherein the prompt corresponds to a first plurality of tokens; providing the prompt as an input to an AI compression model; and obtaining a compressed prompt based on an output of the AI compression model, wherein the compressed prompt corresponds to a second plurality of tokens which is smaller than the first plurality of tokens, and wherein the prompt and the compressed prompt are expressed using natural language.
Owner:SAMSUNG ELECTRONICS CO LTD

Sign-language translation

System and techniques to facilitate the translation of a sign language into another language are described herein. A modular architecture may be used in which the output of different classifiers may be used to produce intermediate representations, or final translations, of the sign language. These classifiers may be trained on different types of signs to enhance accuracy while reduce training time and complexity.
Owner:SORENSON IP HOLDINGS LLC

Natural language semantic recognition model and system for linguistics

The invention discloses a natural language semantic recognition model and system for linguistics, and relates to the technical field of semantic recognition, the natural language semantic recognition model comprises a multi-level context perception semantic recognition module, the multi-level context perception semantic recognition module receives pre-processing data output by a text pre-processing module, and the text pre-processing module sends the pre-processing data to the text pre-processing module; the semantic weight is adjusted under the guidance of the disambiguation optimization module, and the multi-level context perception semantic recognition module is used for extracting semantic features from a plurality of context levels. According to the method, a multi-level context perception semantic recognition module is designed, semantic features are extracted to capture the logic relation between a complex context and a long text, the problem that a traditional model is insufficient in understanding of complex sentence patterns and multi-level contexts is solved, and the accuracy and expression richness of overall semantic recognition are improved.
Owner:GUANGDONG OCEAN UNIVERSITY

Casting industry abnormal sound detection and grading response method based on voiceprint recognition

The invention provides a casting industry abnormal sound detection and grading response method based on voiceprint recognition, and belongs to the technical field of casting industry detection. A high-temperature-resistant microphone array is arranged at an easy-to-leak part of cast aluminum equipment, and three-stage filtering noise reduction and amplitude normalization preprocessing are adopted, so that the problem of poor signal quality caused by noise interference in a complex environment is effectively solved; the characteristics of the molten aluminum leakage sound in different frequency bands and different stages can be captured through variable window long-short time Fourier transform and extended Mel frequency cepstrum coefficient in combination with extraction of an energy change rate and a frequency spectrum gravity center; a Transform-CNN hybrid deep learning model based on an attention mechanism is constructed, and feature screening is optimized through principal component analysis and recursive feature elimination, so that the recognition and generalization ability of the model to the abnormal sound in the casting industry is significantly improved; and meanwhile, graded response measures are made based on the detection result, so that the abnormal conditions of molten aluminum leakage with different severity degrees are processed.
Owner:SHENZHEN POLYTECHNIC

TF-IDF and cross entropy-based cue word compression method and system

The invention discloses a cue word compression method and system based on TF-IDF and cross entropy, belongs to the technical field of large model cue word compression, and aims to solve the problems that redundant information is introduced into long cue words, the model efficiency is reduced and the cost is increased. To-be-compressed content is divided into sentences at the sentence level and then converted into embedded vectors, and the Euclidean distance is calculated in combination with problem vectors so as to screen related sentences; calculating a TF-IDF value at the word level through a word frequency and an inverse document frequency to extract keywords and recombine sentences; and selecting a reference model and a basic model at the Token level, identifying the key Token based on a cross entropy loss difference value, and splicing the key Token in sequence to generate a compressed cue word. According to the method, a complex calculation structure is avoided, the inference efficiency is improved while the semantic integrity is maintained, and the resource consumption is reduced.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Sound effect adjusting method and training method and device of sound effect adjusting multi-mode large model

The invention relates to a sound effect adjustment method, a training method of a sound effect adjustment multi-modal large model, computer equipment and a storage medium. The method comprises the following steps: in response to a sound effect adjustment request for target music sent by a terminal, obtaining an original audio of the target music; inputting the original audio of the target music into a target encoder module of the trained sound effect adjustment multi-mode large model to obtain audio features of the target music, inputting the audio features into a projection module, and converting the audio features into music description semantic features of the target music; the music description semantic feature is a semantic feature of a description text of the target music; and inputting the music description semantic features of the target music into the large language model module to obtain sound effect adjustment parameters of the target music. By adopting the method, a user does not need to manually select a sound effect adjustment mode, and the adjusted parameters can be determined according to the target music, so that the sound effect adjustment effect can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Training encoder model and / or using trained encoder model to determine responsive action(s) for natural language input

Systems, methods, and computer readable media related to: training an encoder model that can be utilized to determine semantic similarity of a natural language textual string to each of one or more additional natural language textual strings (directly and / or indirectly); and / or using a trained encoder model to determine one or more responsive actions to perform in response to a natural language query. The encoder model is a machine learning model, such as a neural network model. In some implementations of training the encoder model, the encoder model is trained as part of a larger network architecture trained based on one or more tasks that are distinct from a “semantic textual similarity” task for which the encoder model can be used.
Owner:GOOGLE LLC

VEM-Token vocal music emotion multi-mode token song and accompaniment deep learning method

The invention discloses a VEM-Token vocal music emotion multi-mode token song sound and accompaniment deep learning method, which is different from an existing method that information is segmented into textual token lexical elements by artificial intelligence and then recognition is carried out on the textual token lexical elements. The method comprises the following steps: performing frequency spectrum processing on a vocal music file, detecting rhythm, segmenting the vocal music file subjected to frequency spectrum processing into a VEM-Token sequence according to the vocal music rhythm, establishing a VEM coordinate system, a VEM function and a VEM library according to multiple modes such as lyrics, singing sound, accompaniment, singer emotion, accompaniment emotion, video and image, performing VEM-Token identification, separating a singing sound stream and an accompaniment stream, and determining the vocal music according to a vocal music expert. And performing multi-modal emotion scoring on the vocal music sample, obtaining a VEM parameter by adopting supervised learning and deep learning algorithms, and learning to obtain the multi-modal emotion of the vocal music sample. For other vocal music works, vocal music multi-mode emotions can be recognized, and a lyric score, a VEM-Token song sound score, a VEM-Token accompaniment score and a VEM-Token music score are output. AI systems including a common large model and the like are accessed, and a vocal music agent Agent capable of listening to the singing and identifying the music score is developed.
Owner:GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD

Lightweight sound signal enhancement method based on adaptive time-frequency modeling, medium and equipment

The invention discloses a lightweight sound signal enhancement method based on adaptive time-frequency modeling, a medium and equipment, and relates to a sound signal processing technology, the lightweight sound signal enhancement method based on adaptive time-frequency modeling mainly comprises the following steps: constructing a sound signal noise reduction model; training a sound signal noise reduction model by using the sound signal data training set to obtain a trained sound signal noise reduction model; and performing noise reduction on the target sound signal by using the trained sound signal noise reduction model to obtain a noise-reduced sound signal. According to the lightweight sound signal enhancement method based on adaptive time-frequency modeling, the medium and the equipment provided by the invention, effective enhancement of the sound signal in a complex noise environment can be realized, the expression ability, the environmental adaptability and the signal restoration precision of the model are improved, and the calculation cost is reduced.
Owner:HAINACORD (HUBEI) TECH CO LTD

Music reactive animation of human characters

Example methods for generating an animated character in dance poses to music may include generating, by at least one processor, a music input signal based on an acoustic signal associated with the music, and receiving, by the at least one processor, a model output signal from an encoding neural network. A current generated pose data is generated using a decoding neural network, the current generated pose data being based on previous generated pose data of a previous generated pose, the music input signal, and the model output signal. An animated character is generated based on a current generated pose data; and the animated character caused to be displayed by a display device.
Owner:SNAP INC

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

Slang usage detection and mitigation for large language models

Certain aspects of the disclosure provide techniques for slang usage detection and mitigation. A method generally includes receiving an input sentence comprising a plurality of tokens; processing, with a first machine learning (ML) model trained for slang classification, the input sentence and thereby generating at least one of: a first classification output, for the input sentence, comprising a slang instance classification; or a second classification output for each of the plurality of tokens of the input sentence, at least one second classification output comprising a slang token classification; and based on the first classification output comprising the slang instance classification and / or the at least one second classification output comprising the slang token classification, processing with a second ML model trained for slang mitigation, the slang-containing input sentence to generate a slang-free output sentence, wherein an entailment score between the input sentence and the output sentence satisfies a threshold.
Owner:INTUIT INC

A method for conditionally transmitting a prompt as input to large language model

PCT designated stage expiredWO2025149444A1Programme controlComputer controlLinguistic modelUser input
A method for conditionally transmitting a prompt as input to a large language model, and for using an output of said large language model to control a lighting system. The method comprising the steps of: receiving a textual user input indicative of an intention to control one or more devices of the lighting system, analyzing said textual user input to determine at least one characteristic of said textual user input, determining a level of complexity of said textual user input based on said at least one characteristic, and only if said level of complexity is above a first complexity threshold, generating a prompt, said prompt comprising at least said textual user input, and transmitting said prompt to said large language model via an output interface, receiving the output of said large language model and controlling said lighting system according to said output.
Owner:SIGNIFY HOLDING BV