Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Audio recognition" patented technology

Modular Audio Recognition Framework (MARF) is an open-source research platform and a collection of voice, sound, speech, text and natural language processing (NLP) algorithms written in Java and arranged into a modular and extensible framework that attempts to facilitate addition of new algorithms.

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

A method and system for remotely recognizing lip language by a drone

PendingCN122369454AData streamNoise
This invention provides a method and system for long-range lip-reading recognition using unmanned aerial vehicles (UAVs), belonging to the fields of artificial intelligence and UAV technology. The system utilizes a camera, IMU (Integrated Mutor Unit), and microphone mounted on the UAV to acquire video streams, IMU data streams, and audio streams. Motion compensation is applied to the video stream using IMU data to stabilize image frames. Then, image sequences of the target speaker's lips and nasal / jaw region are extracted and input into a lip-reading recognition model to obtain initial recognition text and visual confidence scores. Simultaneously, the signal-to-noise ratio (SNR) of the audio stream is calculated. When the visual confidence score is low and the SNR is high, the audio recognition process is triggered, generating audio recognition text and an audio confidence score. Finally, the visual and audio confidence scores, along with the lip movement-audio temporal matching degree, are weighted and fused to generate the final recognition text. This invention effectively solves the problems of video instability and background interference in long-range lip-reading recognition, improving recognition accuracy in complex environments.
Owner:FUZHOU PLANNING DESIGN & RES INST

Audio recognition method, apparatus, device, and storage medium

ActiveCN119905108BNoiseEngineering
This application relates to an audio recognition method, apparatus, device, and storage medium, and pertains to the field of vehicle technology. The method includes: in response to detecting an abnormal noise audio, determining an abnormal noise region based on the abnormal noise audio; if a fault code is detected, determining the faulty component corresponding to the fault code based on the fault code; if the region where the faulty component is located coincides with the abnormal noise region, determining that the abnormal noise audio is emitted by the faulty component. This is used to improve the accuracy of vehicle operation audio recognition.
Owner:CHONGQING CHANGAN TECH CO LTD

Video content recognition methods, devices, equipment, storage media, and software products

PendingCN122313344ANoise (video)Noise
This application provides a video content recognition method, apparatus, device, storage medium, and program product, belonging to the field of data processing technology. The method includes: acquiring a video to be recognized, comment data of the video to be recognized, and descriptive text, wherein the video to be recognized consists of video data and audio data; inputting the comment data into a semantic analysis model to obtain a judgment result of the background music volume; if the volume judgment result indicates a large volume, inputting the video data and audio data into an audio-video fusion extraction model to obtain fusion information text output by the audio-video fusion extraction model; using the fusion information text and the descriptive text to remove background noise from the audio data to obtain a noise-reduced frequency; and determining video content information based on the noise-reduced frequency and the video data. This method solves the problem of low accuracy in audio recognition and information extraction when background music is present.
Owner:CHENGDU TD TECH LTD

Emotion analysis and feedback device based on visual and audio recognition

PendingCN122296895Areduce distractionsAvoid collection blind spotsHuman bodyBlind zone
This invention discloses an emotion analysis and feedback device based on visual and audio recognition in the field of psychological instruments and equipment technology. The device includes: a mobile base; a functional panel, vertically fixed to the top of the mobile base; a positioning ring, vertically suspended to the side of the functional panel, with an internal opening larger than the frontal contour of the human body, used for centering the user and restricting their sitting posture; a connecting arm, one end connected to the lower end of the functional panel and the other end connected to the positioning ring, used to support the positioning ring in a suspended state; and a flexible covering assembly located above the connecting arm, its surface forming a sloped structure facing the functional panel, used to guide the user's gaze to focus on the functional panel. This device features a hollow, large-sized positioning ring, which can center and align the user's torso and correct their sitting posture, avoiding blind spots caused by posture deviation. Combined with multi-directionally deployed visual and audio acquisition components, it significantly improves the accuracy of emotion analysis.
Owner:ANHUI SUNSHINE HEART HEALTH TECH DEV CO LTD

Data processing method and apparatus

This specification provides a data processing method and apparatus, wherein the data processing method includes: preprocessing initial audio data to obtain audio data, and inputting the audio data into an audio recognition model to obtain initial text containing prosody identifiers; determining at least one text unit corresponding to the initial text, and determining time information corresponding to each of the at least one text unit based on the audio data; updating the initial text in the prosody identifier dimension based on the time information corresponding to each of the at least one text unit to obtain target text corresponding to the audio data, wherein the target text and the audio data are used to train an audio generation model.
Owner:BEIJING YUANLI WEILAI SCI & TECH CO LTD

Speaker recognition methods, devices, equipment, and media based on pre-trained speech models

This application discloses a speaker recognition method, apparatus, device, and medium based on a pre-trained speech model, relating to the field of audio recognition technology. The method includes: inserting a target adapter after the target layer of a pre-trained speech model, and processing the target hidden representation output by the target layer using the target adapter; fusing the adapted representations to obtain a fused feature representation, and obtaining a main task loss based on the fused feature representation; acquiring several speech segments, acquiring target speech features corresponding to each speech segment, and adjusting the distance between the target speech features in the target vector space to obtain a content-identity decoupling contrast loss; adjusting the parameters of the target adapter according to the main task loss and the content-identity decoupling contrast loss, and using the adjusted adapter for speaker recognition. By introducing the content-identity decoupling contrast loss, interference from changes in speech content on speaker embedding is suppressed.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Portable personal assistant system and method for sensory data storage, manipulation, and exchange

Ways to facilitate multimodal interaction and communication for users are provided. A personal assistant system for a user includes a portable device equipped with a three-dimensional (3D) camera for generating images of 3D space and a propulsion system for self-propulsion. The personal assistant system also comprises a speaker, a microphone, and a control system interfacing with these components. The control system features a navigation module for controlling the propulsion system, an audio recognition module for converting audio from the microphone, and an image recognition module for processing images from the 3D camera. A processor transforms these formats into a processor-output, while an artificial intelligence (AI) module analyzes the formats to identify actionable commands and convert them into processor-output. An output generating module then converts the processor-output into a user-friendly output. Use of the personal assistant system enhances user convenience, accessibility, and engagement in educational, professional, and personal pursuits.
Owner:RAVLYUK OLHA

Machine learning based video recognition method, device, server and storage medium

Embodiments of the present application disclose a video recognition method and device based on machine learning, a server and a storage medium. Embodiments of the present application can obtain a target video; obtain a source video corresponding to the target video, the target video being created by processing the source video; compare the target video and the source video in content to obtain a content type of the target video; when the content type of the target video is a funny content type, perform audio recognition on the target video and the source video to determine an audio type of the target video; when the audio type of the target video is a funny dubbing type, determine the target video as a funny dubbing video, so as to push the funny dubbing video to a user. Embodiments of the present application compare the target video with the source video in dimensions such as content and audio to identify whether the target video is a funny dubbing video created by processing the source video. Thus, the present application can accurately identify a funny dubbing video from a plurality of videos, and improves the efficiency of video recognition.
Owner:TENCENT TECH (BEIJING) CO LTD

Chatbot to assist in vehicle shopping

Methods for systems generating recommendations regarding a vehicle purchase are disclosed. An artificial intelligence (AI) or machine learning (ML) chatbot or voicebot is provided to receive input from a user, determine information regarding a type of vehicle based upon the input, and generate a total cost of ownership of the type of vehicle for presentation to the user. The chatbot may apply a natural language processing (NLP) algorithm to a text input to generate an intermediate input for use in determining the total cost of ownership. The voicebot may apply an audio recognition algorithm to an audio input to generate text, which may then be processed by the voicebot or the chatbot using an NLP algorithm to generate an intermediate input for use in determining the total cost of ownership.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Interaction control method and device based on agent telephone customer service, equipment, storage medium and program product

The application provides an interaction control method and device based on an agent telephone customer service, equipment, a storage medium and a program product, relates to the technical field of artificial intelligence, and the method comprises the following steps: acquiring a speech recognition result, semantic integrity information and user speaking state information output by an audio recognition module; inputting the speech recognition result, the semantic integrity information and the user speaking state information into a reinforcement learning model in a reinforcement learning agent to obtain a decision action output by the reinforcement learning model; and performing a calling operation on a business agent based on the decision action. By deploying the reinforcement learning agent in the voice customer service system, the application decides the calling time of the business agent according to the semantic integrity and the speaking state of the user voice, effectively avoids the mis-triggering problem caused by user pauses or voice punctuation, and improves the accuracy and fluency of voice interaction without modifying the existing business agent.
Owner:CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1

system

PendingJP2026105389ATelecommunicationsHome robot
Provide a system. 【Solution means】 An audio receiving means for receiving audio data and converting it into a digital format, An audio recognition means for converting the audio data into text data, A natural language processing means for analyzing the text data and extracting schedule information, A schedule registration means for automatically registering the schedule information in a calendar, A translation means for translating into a different language when the text data is a specific language, A notification means for notifying the user's portable information terminal of the converted text data and the translated data, A home robot device installed in a home environment for recording the audio data, A system including the above.
Owner:SOFTBANK GROUP CORP

Model training method, audio data processing method, corresponding device and product

The present disclosure provides a model training method, an audio data processing method, corresponding devices and products. The model training method comprises: for any round, inputting first audio training data corresponding to a preset application field into a first target labeling model and a first audio recognition model of the current round to obtain first predicted labeling data corresponding to the first audio training data; deleting target audio training data in the first audio training data according to the first predicted labeling data to obtain second audio training data; performing noise training on the first audio recognition model based on the second audio training data to obtain a second audio recognition model; training the first target labeling model based on the second audio training data and the second audio recognition model to obtain a second target labeling model of the current round; and performing model training of the next round when a preset convergence condition is not met. The embodiments of the present disclosure can make the target labeling model quickly applicable to the preset application field through training.
Owner:MOORE THREADS TECH CO LTD

System for recognizing soundscapes

A system for recognizing soundscapes includes audio extractors, audio sampling and modulation circuits, an audio recognition device, environment detectors, a data integration circuit, a wireless communication module, a wireless base station, and a cloud server. The audio extractors extract soundscape signals so that the audio sampling and modulation circuits generate audio modulation signals. The audio recognition device recognizes the audio modulation signals to generate audio features. Based on biological audio models respectively corresponding to different species, the audio recognition device determines the number of the species corresponding to the audio features, thereby adjusting the sampling rate, the working period, or both of the soundscape signal extracted by each audio extractor. The environment detectors detect different environment-related data. The data integration circuit synchronizes the environment-related data, the audio features, the species corresponding thereto and uploads them to the server through the base station and the communication module.
Owner:FAR EASTONE TELECOMMUNICATIONS CO LTD

Lifelog device utilizing audio recognition, and method therefor

ActiveUS12670208B2LifelogAudio frequency
The present invention relates to a lifelog device utilizing audio recognition and a method therefor, and to a device capable of recording and classifying audio lifelogs by means of an artificial intelligence algorithm. To this end, the lifelog device of the present invention comprises: an input unit for inputting lifelog data including an audio signal; and analysis unit for analyzing the inputted data; a determination unit for classifying the class of the data on the basis of the analyzed analysis value; and a recording unit for recording the inputted data and the classified class of the data.
Owner:COCHL INC

Model training method, audio data processing method, corresponding device and product

ActiveCN122116882BNoiseEngineering
The present disclosure provides a model training method, an audio data processing method, corresponding devices and products. The model training method comprises: for any round, inputting first audio training data corresponding to a preset application field into a first target labeling model and a first audio recognition model of the current round to obtain first predicted labeling data corresponding to the first audio training data; deleting target audio training data in the first audio training data according to the first predicted labeling data to obtain second audio training data; performing noise training on the first audio recognition model based on the second audio training data to obtain a second audio recognition model; training the first target labeling model based on the second audio training data and the second audio recognition model to obtain a second target labeling model of the current round; and performing model training of the next round when a preset convergence condition is not met. The embodiments of the present disclosure can make the target labeling model quickly applicable to the preset application field through training.
Owner:MOORE THREADS TECH CO LTD

Audio recognition method and device, electronic equipment and storage medium

This application relates to an audio recognition method, apparatus, electronic device, and storage medium, applied in the field of computer technology. The method includes: acquiring audio data to be recognized, the audio data including audio data of at least one speaker; inputting the audio data to be recognized into a voiceprint recognition model to extract target voiceprint features of a target speaker from the audio data; the voiceprint recognition model uses text data of a specified speaker in training sample data as a supervision signal to learn to mask voiceprint features unrelated to the text data during voiceprint feature extraction; and generating a recognition result for the audio data corresponding to the target speaker in the audio data to be recognized based on the target voiceprint features.
Owner:IFLYTEK CO LTD

Intelligent conference recording method and apparatus, intelligent device and storage medium

PCT designated stageWO2026137700A1Frame sequenceEngineering
The present application is applicable to the technical field of intelligent devices, and provides an intelligent conference recording method and apparatus, an intelligent device, and a storage medium. The method comprises: collecting a video frame sequence image and an on‑site audio of a conference site in real time; identifying whether a target speaking trigger posture exists in the video frame sequence image; on the basis of the identified target speaking trigger posture and a conference participant information record, determining a current speaker and identity information thereof; associating speaking content corresponding to the on-site audio collected from the moment when the target speaking trigger posture is identified with the speaker and the identity information thereof; and generating a conference record on the basis of the speaking content associated with all speakers and identity information thereof determined at the conference site.
Owner:SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD

Audio recognition method and apparatus, readable medium, and electronic device

Embodiments of the present disclosure relate to an audio recognition method and device, readable medium and electronic equipment. The method comprises: obtaining a target audio to be recognized; inputting the target audio into a pre-generated target audio recognition model to obtain a target text output by the target audio recognition model. The target audio recognition model can be a model pre-generated according to a sample data set, a first loss function and a second loss function. The sample data set can include a plurality of target sample audios and a target sample text corresponding to each target sample audio. The target sample audio can include a first sample audio and a second sample audio. The first loss function can be used to determine the text prediction accuracy of the first sample audio. The second loss function can be used to determine the text prediction accuracy of the second sample audio.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Audio denoising method and system based on joint optimization of video understanding and audio recognition

This invention discloses an audio denoising method and system based on joint optimization of video understanding and audio recognition, belonging to the field of audio denoising technology. The method includes: acquiring original audio signals and corresponding original video signals within the same time period; performing feature extraction and audio category recognition on the original audio signals to obtain an audio category temporal probability matrix; performing video semantic sound prediction on the original video signals to obtain a video semantic category probability matrix; performing cross-modal temporal alignment between the audio category temporal probability matrix and the video semantic category probability matrix to obtain a cross-modal difference guidance vector; fusing the original audio features, cross-modal difference guidance vector, three-state semantic control vector, and knowledge correction factor through multi-source feature cross-attention to obtain a joint representation; and inputting the joint representation into an audio denoising model to output a preliminary denoising result. This invention improves the semantic accuracy of the denoising result in complex environments.
Owner:SICHUAN HUSHAN ELECTRIC APPLIANCE

Conversational scene-based cue mining system, method, computer device and storage medium

The application discloses a conversation scene-based clue mining system, method, computer device and storage medium, and the system comprises an audio and video acquisition device, an audio recognition module, a video analysis module, a data fusion module, a clue mining module and a result presentation module. The audio and video acquisition device collects audio and video data in a conversation process, the audio recognition module converts the collected audio data into text and extracts audio features, the video analysis module performs facial expression recognition and body language analysis on the video data to extract video features, the data fusion module fuses and processes the audio feature data and the video feature data, the clue mining module mines data clues by using a machine learning algorithm after fusion processing, and classifies the mined clues; and the result presentation module visually presents the mined clues. The subjective factor influence can be reduced, the clue mining efficiency is improved, multi-modal information fusion is realized, and the advantages of visual result presentation and strong scalability are achieved.
Owner:HUADI COMP GROUP

Speech recognition method based on memristor and electronic device

This application discloses a memristor-based speech recognition method and electronic device, relating to the field of speech recognition technology. The method includes: acquiring the original pulse sequence of audio data; inputting the original pulse sequence into a speech recognition network; during the training process of the speech recognition network, the parameters of the memristor-based reservoir are kept fixed to introduce hardware errors from the memristor during training, and the key parameters of the classifier are iteratively corrected in an error-containing environment, enabling it to learn the ability to compensate for hardware errors. Based on this, the memristor-based reservoir first extracts features from the original pulse sequence to obtain first pulse features; then, the classifier performs audio classification based on the key parameters and the first pulse features, compensating for feature distortion caused by hardware errors through the key parameters, thereby improving the robustness of speech recognition and obtaining more accurate audio-recognized text.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Audio deepfake identification method and system based on inconsistency detection characterization

This invention provides a method and system for identifying deepfake audio based on representation inconsistency detection, belonging to the field of audio recognition technology. The method processes the audio waveform to be detected using a low-to-high consistency detection model, outputting the forgery probability of the detected audio waveform. The low-to-high consistency detection model includes a staged feature extraction module comprising a 24-layer XLS-R model. The first 21 layers of the XLS-R model output low-level acoustic representations, and the last 3 layers output high-level semantic representations. This invention proposes that the core detectable clue for deepfake audio lies in the mismatch between low-level acoustic representations and high-level semantic representations. This innovative discovery fundamentally reconstructs the focusing logic of the detection task, upgrading the traditional passive identification targeting specific flaws to active verification based on internal consistency. It possesses solid theoretical support and can overcome scenario limitations, providing stronger generalization ability and long-term application value for cross-domain and multi-type deepfake audio detection.
Owner:HEFEI FEIDU INFORMATION TECH CO LTD +1

An audio recognition method, an electronic device, and a readable storage medium

The application discloses an audio recognition method, an electronic device and a readable storage medium. The method comprises the following steps: acquiring preset audios, and extracting preset audio features corresponding to the preset audios respectively; performing clustering processing on the preset audio features to obtain a plurality of audio feature groups; selecting standard audio features in each audio feature group respectively, and constructing an audio feature library by using the standard audio features; acquiring a to-be-recognized audio sent by a terminal; wherein the to-be-recognized audio is acquired by a corresponding sound collecting device of the terminal; extracting to-be-recognized audio features of the to-be-recognized audio; determining a target audio feature most similar to the to-be-recognized audio features from the standard audio features in the audio feature library based on the to-be-recognized audio features; and sending target audio information corresponding to the target audio feature to the terminal. The method can greatly reduce the data amount of the audio feature library while ensuring the reliability of audio recognition through clustering and extraction of standard audio features.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Business process identification methods, devices, equipment and storage media

This invention discloses a business process identification method, apparatus, device, and storage medium. The business process identification method includes: acquiring an audio file to be identified; performing speech recognition on the audio file to obtain an audio recognition result; executing a preset business process chain to call a process state matcher to match the audio recognition result, thereby obtaining a business process identification result. This invention can improve the accuracy of business process identification results.
Owner:北京中关村科金技术有限公司

Method, system and computer storage medium for identifying negative information

PendingCN122435369APattern recognitionMedicine
The present application relates to the technical field of data processing, and relates to a negative information recognition method and system and a computer storage medium, comprising: performing text conversion on audio data in an acquired video to be detected to obtain target text; performing sentiment analysis and negative recognition processing on the target text to obtain an audio recognition result carrying a frame extraction review label; determining a corresponding suspected video range set in the video to be detected according to the audio recognition result carrying the frame extraction review label; performing frame extraction operation on the suspected video range set using a preset frame extraction strategy to obtain a key frame image set corresponding to the suspected video range; performing visual analysis on the key frame image set to recognize negative content of the key frame image set to obtain a target visual review result corresponding to the key frame image set; and performing fusion judgment on the audio recognition result and the target visual review result based on a preset fusion strategy to obtain a target judgment result for the video to be detected.
Owner:SHENZHEN ZERO INTELLIGENT TECHNOLOGY CO LTD