Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4438 results about "Speech recognition" patented technology

Speech recognition is a interdisciplinary subfield of computational linguistics that develops methodologies and technologies that enables the recognition and translation of spoken language into text by computers. It is also known as automatic speech recognition (ASR), computer speech recognition or speech to text (STT). It incorporates knowledge and research in the linguistics, computer science, and electrical engineering fields.

Systems and methods for generating an equal-loudness contour response using an auricular device

A system may include a storage device, configured to store computer-executable instructions. A system may include an ear-bud configured to be positioned within an ear canal of a user, the ear-bud comprising: a speaker, a microphone; and one or more processors in communication with the storage device, wherein the computer-executable instructions, when executed by the one or more processors, cause the one or more processors to: obtain a user hearing profile, obtain an equal-loudness hearing profile, receive audio data from the microphone, and generate a second audio data based on a first sound-pressure level, a second sound-pressure level, a first frequency; and cause the speaker to emit the second audio data within the ear canal of the user, such that the user perceives the audio data as if the user has normal hearing.
Owner:MASIMO CORP

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

Multimodal Machine-Learned Models for Unified Attention and Response Predictions for Visual Content

Aspects of the disclosed technology include computer-implemented systems and methods for machine-learned multimodal models. A machine-learned multimodal model includes one or more embedding layers configured to generate one or more image tokens and one or more text tokens in response to the imagery and the text, a transformer encoder configured to receive the one or more image tokens and the one or more text tokens and generate one or more fused image tokens and one or more fused text tokens, a heatmap predictor configured to obtain the one or more fused image tokens and generate at least one image heatmap, and a sequence predictor configured to obtain the one or more fused image tokens and the one or more fused text tokens and generate a predicted sequence associated with the image.
Owner:GOOGLE LLC

Natural language prompt generation

Techniques for determining a follow-up natural language prompt, to continue a user-system dialog, are described. The system determines ASR output data representing a user input and / or a system-generated responsive thereto. The system determines one or more entities represented in the ASR output data and / or the system-generated response, and identifies one or more natural language prompts associated with the one or more entities in storage. The system filters out prompts classified as likely to result in an unsatisfactory user experience, using dialog history data including a previous user input(s) and / or a previous system-generated response(s). The system determines context(s) associated with the instant user and / or device, and uses this context(s), the ASR output data, and / or the system-generated response to determine which of the follow-up prompts is to be presented to the user.
Owner:AMAZON TECH INC

Bluetooth hearing aid earphone optimization method and system for adaptive hearing compensation

The invention discloses a Bluetooth hearing-aid earphone optimization method and system for adaptive hearing compensation. The method comprises the following steps: S1, acquiring a personalized hearing curve of a user and background audio data of left and right ears in each typical scene, and respectively preprocessing the background audio data; s2, extracting time-frequency fusion features of each typical scene, generating scene feature vectors, and constructing a scene feature vector library; s3, calculating a scene perception conversion vector, an abrupt change risk vector and a smooth gain vector; s4, calculating a binaural gain change vector, a binaural synchronous entropy vector, a phase alignment factor vector and a left and right ear cooperative gain vector; and S5, according to the left and right ear cooperative gain vectors, processing the background audio data after left and right ear preprocessing in the current scene, and outputting the processed background audio data to the Bluetooth hearing-aid earphone. According to the invention, the problem of gain abrupt change generated during scene switching of the existing Bluetooth hearing-aid earphone and the problem of spatial perception distortion caused by binaural coordination imbalance can be solved.
Owner:HUNAN DINO INTELLIGENT TECHNOLOGY CO LTD

System and method for detecting deep fake audio

A system for analyzing audio includes a memory configured to store known digital audio representation containing known fraudulent audio streams and a processor operably coupled to the memory. The processor receives a portion of an audio stream from an external device and produces a transcript of the portion of the audio stream. The processor then determines a timing score, an emotional score, a background score, and a content score by analyzing the portion of an audio stream and the corresponding transcript and comparing them to the known digital audio representations and transcripts. The processor then determines if the audio stream is malicious by combining the timing score, emotional score, background score, and content score to produce a combined score and comparing the combined score to a threshold. The processor notifies a user that the call may be fraudulent when the combined score is greater than the threshold.
Owner:BANK OF AMERICA CORP

Atmosphere lamp control method, electronic device and program product

The invention discloses an atmosphere lamp control method, electronic equipment and a program product. The method comprises the following steps: acquiring PCM data of an audio signal in real time in the music playing process of a vehicle; performing frequency domain analysis on the PCM data corresponding to each audio frame, and extracting frequency information and loudness information; the method comprises the following steps: calculating frequency information and loudness information of a preset number of continuous audio frames, and respectively calculating a frequency change rate and a loudness change rate within a preset time; respectively comparing the frequency change rate and the loudness change rate with corresponding change rate thresholds; and if at least one of the frequency change rate and the loudness change rate exceeds the corresponding change rate threshold value, a corresponding light updating instruction is sent to the atmosphere lamp control module. The music rhythm function of the automotive interior atmosphere lamp can be enhanced, and the cooperative interaction ability of music and the atmosphere lamp is improved.
Owner:ANHUI KAIYANG TECHNOLOGY CO LTD +1

Methods and systems for generating and rendering object based audio with conditional rendering metadata

Methods and audio processing units for generating an object based audio program including conditional rendering metadata corresponding to at least one object channel of the program, where the conditional rendering metadata is indicative of at least one rendering constraint, based on playback speaker array configuration, which applies to each corresponding object channel, and methods for rendering audio content determined by such a program, including by rendering content of at least one audio channel of the program in a manner compliant with each applicable rendering constraint in response to at least some of the conditional rendering metadata. Rendering of a selected mix of content of the program may provide an immersive experience.
Owner:DOLBY LABORATORIES LICENSING CORP +1

Robust noise reduction processing method and system for sound wave signal self-supervised learning enhancement

ActiveCN121415799ASpeech analysisPhysical realisationTime domainProbability propagation
The invention provides a sound wave signal self-supervised learning enhanced robust noise reduction processing method and system, and relates to the technical field of signal processing, and the method comprises the steps: carrying out the feature enhancement of an initial time-frequency representation through a dynamic adaptive mask strategy, constructing a self-supervised reconstruction task based on the mask time-frequency representation, separating noise and signal subspaces in a semantic manifold space, and carrying out the self-supervised learning enhanced robust noise reduction. And establishing a probability propagation network in combination with time domain continuity characteristics to model a local dependency relationship, and finally generating a noise reduction weight and realizing semantic fidelity optimization. The method can effectively improve the noise reduction effect and semantic integrity of sound wave signals in a noise complex environment.
Owner:BEIJING GUANYU INFORMATION TECHNOLOGY CO LTD

Audio signal generation model and training method using generative adversarial network

A generative adversarial network-based audio signal generation model for generating a high quality audio signal may comprise: a generator generating an audio signal with an external input; a harmonic-percussive separation model separating the generated audio signal into a harmonic component signal and a percussive component signal; and at least one discriminator evaluating whether each of the harmonic component signal and the percussive component signal is real or fake.
Owner:ELECTRONICS & TELECOMM RES INST +1

Data analysis and processing system based on clinical hearing detection

The invention relates to the technical field of medical data processing, in particular to a data analysis and processing system based on clinical hearing detection, and aims to solve the problems that dynamic characteristics of auditory nerve response cannot be systematically described, high-dimensional multi-modal feature vectors cannot be constructed, and the data analysis and processing efficiency is low in the prior art. A stable and reliable input basis cannot be provided for intelligent classification, anomaly detection and individualized analysis, and the sensitivity and specificity of hearing state recognition are reduced; the multi-dimensional hearing feature extraction module fuses time domain, frequency domain and non-linear complexity analysis methods, systematically depicts dynamic features of auditory nerve response, constructs a high-dimensional multi-modal feature vector by integrating three-dimensional features, remarkably improves sensitivity and specificity of hearing state recognition, and improves the accuracy of hearing state recognition. The feature representation method provides a stable and reliable input basis for subsequent intelligent classification, anomaly detection and individualized analysis.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Adjusting audio output in a vehicle

In aspects of adjusting audio output in a vehicle, a vehicle audio system implements an audio playback manager that detects a location of a first person in the vehicle, the first person located in proximity of a first speaker device configured for audio output. The audio playback manager also detects an additional location of a second person in the vehicle, the second person located in proximity of a second speaker device configured for the audio output. The audio playback manager determines a preference for the first person related to the audio output and an additional preference for the second person related to the audio output. Based on the preference of the first person or the additional preference of the second person, the audio playback manager adjusts a volume of at least one of the first speaker device or the second speaker device.
Owner:MOTOROLA MOBILITY LLC

Large model cue word injection protection method and device, equipment, medium and product

The invention provides a large model cue word injection protection method and device, equipment, a medium and a product, and relates to the technical fields of computers, network security, large models and the like. The large model cue word injection protection method comprises the steps of obtaining a current keyword in a current cue word and a current semantic vector; the current cue word is a cue word to be input into a large model; matching the current keyword with candidate keywords of each candidate fragment to determine a target fragment in the plurality of candidate fragments; matching the current semantic vector with a candidate semantic vector of each candidate cue word of a target fragment so as to determine a target cue word in the plurality of candidate cue words; determining a risk level of the current cue word based on the current cue word and the target cue word; and processing the current prompt word based on the risk level. According to the invention, the large model cue word injection protection effect can be improved.
Owner:RIVER INFORMATION TECH SHANGHAI CO LTD

Construction method of vem-token vocal emotion multimodal magic modification model

The construction method of the VEM-Token vocal emotion multi-modal magic modification model is different from the natural language processing model NLP-Token which interprets music through text. Instead, the VEM-Token sound-to-text innovation model comes with vocal emotion multi-modal information. The model captures and aligns the music tempo of sample songs and user learning songs, identifies various modalities of vocal emotion, divides the file into word units based on tempo, obtains VEM parameters through supervised learning and reinforcement learning, and decomposes songs into vocals, accompaniment, and emotion. The magic modification model provides a multi-modal magic modification method for imitating sample songs, including vocals, accompaniment, emotion overtones, emotion fluctuations, learning to sing, voice cloning, lyrics, pitch calibration, ornaments, tempo length, rhythm speed, tempo strength, freedom, and multiple sample magic modification. It also provides member management, mobile and PC application systems, dedicated support hardware, and communication protocols including MIDI, making it easy to access popular AI large models, reducing model hallucinations, and forming AI vocal agents and AI karaoke.
Owner:GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD

Space-time multi-mode video understanding method based on Video LLeMA2

The invention discloses a time-space multi-mode video understanding method based on Video LLeMA2. The time-space multi-mode video understanding method comprises the following steps: S1, extracting an image frame sequence and an audio stream sequence of video data; s2, carrying out pretreatment; s3, inputting the image frame sequence into a visual encoder to generate a visual initial feature; s4, inputting the visual initial features into a space-time convolution connector, and generating visual modal features by using three-dimensional convolution and a RegStage module; s5, inputting the audio stream sequence into an audio encoder to obtain audio modal features; s6, aligning the audio modal features with the visual modal features; s7, performing feature fusion; and S8, obtaining a video content understanding result through a language decoder. The video content semantic understanding method is based on the Video LLeMA2 model, integrates multi-modal space-time modeling and language generation technologies, realizes video content semantic understanding, and has the advantages of natural expression and high precision.
Owner:WUHAN RUANBANG INTELLIGENT TECHNOLOGY CO LTD

Audio-lip movement correlation measurement for dubbed content

Methods and apparatus are described for evaluating dubbing of media content. Phonemes in dubbed audio are extracted and mapped to visemes. Lip poses in video frames of the media content corresponding to the phonemes of the dubbed audio are compared to the visemes determined from the dubbed audio. A notification may be generated based on the comparison that indicates synchronization of the dubbed audio to lip poses of the video.
Owner:AMAZON TECH INC

System and method for training open-vocabulary object detectors using generated region-text pairs

Disclosed herein is a method of generating region-text pairs for training open-vocabulary object detection. The method innovates text-to-region and region-to-text processes, along with the introduction of a Scene-Aware Inpainting Guider and a Localization-Aware Region-Text Contrastive Loss.
Owner:CARNEGIE MELLON UNIV

Transmission-agnostic presentation-based program loudness

This disclosure falls into the field of audio coding, in particular it is related to the field of providing a framework for providing loudness consistency among differing audio output signals. In particular, the disclosure relates to methods, computer program products and apparatus for encoding and decoding of audio data bitstreams in order to attain a desired loudness level of an output audio signal.
Owner:DOLBY LABORATORIES LICENSING CORP +1

Video synthesis method and system

The invention discloses a video synthesis method and system, and relates to the technical field of audio and video processing. A video synthesis system comprises a video source acquisition and processing module, a cross-modal semantic understanding module, an attention tensor generation module, a hierarchical progressive fusion module and a quality evaluation and optimization module. According to the method, the spatial-temporal joint features are extracted through the three-dimensional convolutional network, and the audio-visual cross-modal attention mechanism is constructed, so that the dynamic association strength of the audio event and the video content can be quantified, and the main body space mask can be generated, and therefore, the traditional isolated visual processing can be expanded into sound and picture semantic linkage understanding; in this way, deep guidance of multi-modal information on the synthesis process is achieved, and the synthesis effect of the video synthesis method and system is improved.
Owner:SUZHOU BROADCASTING SYST +1

Respiratory disease diagnosis device and method

The invention particularly relates to a respiratory disease diagnosis device and method, and the device comprises a respiratory sound signal collection unit which is used for obtaining a current respiratory sound signal; the breathing sound signal preprocessing unit is used for performing segmentation processing and denoising processing on the current breathing sound signal to obtain a preprocessed breathing sound signal; the feature extraction unit is used for extracting multi-modal features of the preprocessed breath sound signals; and the detection unit is used for inputting the multi-modal features into a preset deep learning network model to obtain a respiratory system disease detection result. Therefore, through multi-modal feature extraction, dynamic fusion and an efficient classification strategy, the problems of relatively low feature identification degree, relatively low disease classification accuracy and the like in the process of researching disease diagnosis based on the breath sound signals in related technologies are solved, and the detection performance of chronic respiratory system diseases based on the breath sound signals is remarkably improved.
Owner:GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)

Sleep stage identification method, device and equipment based on multi-mode signal and medium

The invention discloses a sleep stage identification method and device based on a multi-mode signal, equipment and a medium. The method comprises the following steps: extracting a target lead signal comprising an electroencephalogram signal and an electro-oculogram signal from a polysleep monitoring signal of a target object; performing data preprocessing on the target lead signal to obtain a reference lead signal; inputting the reference lead signal into a sleep stage recognition model to obtain a sleep stage prediction result of the target object; wherein the sleep stage recognition model comprises a preliminary feature extraction unit, two deep feature extraction units, a dynamic gating fusion unit, a bidirectional long-short-term memory network unit, an attention unit and a classification unit, the deep feature extraction units are constructed based on a focus adjustment mechanism, and multi-scale one-dimensional deep convolution is adopted; and the two deep feature extraction units are used for performing deep feature extraction on the electroencephalogram signal and the electro-oculogram signal respectively. According to the scheme, the sleep stage recognition precision and robustness can be improved by using the sleep stage recognition model.
Owner:YANGTZE RIVER DELTA GUOZHI (SHANGHAI) INTELLIGENT MEDICAL TECH CO LTD

Extracting audio signal from audio mixed signal using neural network

The present disclosure provides an audio system, method, and system for facilitating machine operation. The machine includes an actuator that assists the tool in performing the task. In an example, an audio system is configured to receive an audio mixed signal of a signal generated by an audio source that includes at least one of a tool or an actuator that is performing a task. An audio source forming an audio mixed signal is identified by a relative position to each microphone of a microphone array that measures the audio mixed signal. The audio system is configured to extract an audio signal from an audio mixed signal generated by the identified audio source based on a correlation of spectral features in a multi-channel spectrogram of the audio mixed signal and directional information indicative of a relative position of the identified audio source. The audio system outputs the extracted audio signal to facilitate operation of the machine.
Owner:MITSUBISHI ELECTRIC CORP

Bearing vibration diagnosis method and system based on large language model

The invention relates to the technical field of wind power technologies, in particular to a bearing vibration diagnosis method and system based on a large language model. The method comprises the following steps: acquiring bearing vibration signal data including a fault frequency spectrum signal x, a reference bearing vibration signal and a reference frequency spectrum signal; performing data preprocessing on the acquired bearing vibration signal data; constructing a fault classification model, wherein the step of constructing a feature recognition network and constructing an alignment network to train the constructed fault classification model; and performing fault prediction by using the trained fault classification model. According to the invention, the fault classification network, the feature alignment module and the large language model are utilized to combine the vibration time domain signal and the cue word together, so that the large language model can indirectly process the ultra-long data text.
Owner:NAT NUCLEAR INFORMATION TECH CO LTD

Lightweight acoustic feature extraction method for end-side audio and video quality inspection

The invention relates to a lightweight acoustic feature extraction method for end-side audio and video quality inspection. The method comprises the following steps: end-side equipment separates audio and video stream data through a double-time-sequence anchor point alignment method, and acquires independent audio data; performing grading preprocessing on the independent audio data to obtain noise-reduced independent audio data; on the basis of the segmented audio, a time domain and frequency domain collaborative extraction algorithm is adopted, multi-dimensional time domain features and frequency features sensitive to audio quality difference are obtained, and a lightweight acoustic feature vector is obtained through an incremental principal component analysis method; based on the lightweight acoustic feature vector, utilizing a dual-threshold matching judgment mode to obtain similarity between features; and on the basis of the feature similarity, through feature dimension deviation verification, obtaining a quality inspection result containing a standard grade and feature anomaly positioning. Lightweight acoustic feature extraction of end side audio and video quality inspection is realized.
Owner:SHANGHAI SIOO INFORMATION TECH CO LTD

Personalized brain wave induction audio generation method and system

ActiveCN121122318ASpeech analysisTime domainNeural oscillation
The invention relates to the technical field of neuroacoustics, in particular to a personalized brain wave induced audio generation method and system, and the method comprises the steps: S1, extracting multi-dimensional features of an input audio signal, the extracted dimensions at least comprising a time domain feature, a frequency domain feature, a pitch feature, an emotion feature and a complex feature; s2, on the basis of the multi-dimensional features extracted in the step S1, parameters of four neuroacoustic induction mechanisms are calculated in parallel; and S3, according to the features obtained in the step S1 and the quadruple mechanism parameters obtained in the step S2, signals corresponding to the four mechanisms are fused through a weight distribution strategy and a signal synthesis formula to generate a final brain wave induction audio. The method has the advantages that four mechanisms of rhythm synchronization, harmonic resonance, neural oscillation coupling and binaural beat frequency effect are fused, parameters of all the mechanisms are dynamically adjusted by analyzing music characteristics, and efficient personalized brain wave induction is achieved.
Owner:FERD MANSON MULTIMEDIA TECH SHANGHAI

Temporomandibular joint sound signal identification and classification method based on deep learning neural network

PendingCN121281556ASpeech analysisTemporomandibular joint soundsData set
The invention provides a temporal-mandibular joint sound signal identification and classification method based on a deep learning neural network, and the method comprises the steps: collecting triaxial joint vibration signals during the movement of bilateral temporal-mandibular joints, carrying out the labeling processing of the collected triaxial joint vibration signals, and respectively labeling normal signals or abnormal signals, thereby obtaining a training data set; constructing a temporal-mandibular joint sound signal recognition and classification model based on the deep learning neural network, wherein the temporal-mandibular joint sound signal recognition and classification model based on the deep learning neural network comprises a data preprocessing module, a feature extraction module and a recognition and classification module; obtaining a trained temporal-mandibular joint sound signal identification and classification model based on the deep learning neural network; inputting a triaxial joint vibration signal to be recognized into the trained temporal-mandibular joint sound signal recognition and classification model based on the deep learning neural network to obtain a final classification result; the method is higher in objectivity and accuracy and higher in robustness.
Owner:AFFILIATED STOMATOLOGICAL HOSPITAL OF NANJING MEDICAL UNIV

Content aware audio processing

PendingUS20260018179A1Speech analysisComfort noiseNoise
Content aware audio processing includes receiving, by a digital signal processor, a frame of audio data. In response to detecting that the frame of audio data is a silent frame, the digital signal processor selects a light graph from a plurality of graphs including the light graph and a full graph. Comfort noise is generated that corresponds to the silent frame. The comfort noise frame is processed through the light graph in place of the silent frame. The light graph is dedicated for processing comfort noise frames.
Owner:ADVANCED MICRO DEVICES INC

Intelligent Muting Of Participant Audio In Communication Sessions

An input audio signal associated with a communication session is received. A determination is made that a participant associated with the input audio signal is not audibly speaking within the input audio signal. In response to determining that the participant is not audibly speaking, an audio feed corresponding to the input audio signal is muted by rendering the audio feed not audible to at least one other participant device connected to the communication session.
Owner:ZOOM COMMUNICATIONS INC