Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

271 results about "Audio recognition" patented technology

Modular Audio Recognition Framework (MARF) is an open-source research platform and a collection of voice, sound, speech, text and natural language processing (NLP) algorithms written in Java and arranged into a modular and extensible framework that attempts to facilitate addition of new algorithms.

Multi-mode collaborative awareness power station high-risk operation inspection method and system

The invention provides a multi-mode cooperative sensing power station high-risk operation inspection method and system, and relates to the technical field of video recognition, and the method comprises the steps: activating an unmanned plane and a quadruped robot after a power station operation task is started; starting a video acquisition unit, and establishing a synchronous video stream; carrying out fusion alignment with the global reference coordinate system through an external synchronization signal; inputting the video sequence of the fused view angle into a multi-view angle action behavior recognition network, and establishing a dangerous behavior grade score; auditory data and olfactory data of the quadruped robot are obtained, and linkage abnormity is established; and polling abnormity is reported according to linkage abnormity and dangerous behavior grade scores. Through the method and the device, the technical problem of low inspection efficiency caused by difficulty in comprehensively identifying potential risks in a dynamic environment due to limitation of a single sensing mode is solved, and the inspection efficiency of a power station is improved by fusing multi-mode data, timely finding and processing the potential risks and improving the accuracy of video and audio identification.
Owner:BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH

End-to-end dialect audio recognition method and system based on multi-layer information fusion

PCT designated stageWO2025255947A1Speech recognitionSpeech soundAudio frequency
An end-to-end dialect audio recognition method (10) based on multi-layer information fusion. The method (10) comprises: performing audio preprocessing on dialect audio to generate an acoustic feature (S102); inputting the acoustic feature into a coder, and then performing, by the coder, a progressive downsampling operation on the acoustic feature to generate multi-layer fine-grained acoustic features (S104); performing, by means of a layer adaptation module, multi-layer information fusion on the multi-layer fine-grained acoustic features to generate a fused acoustic feature (S106); performing, by means of a cross-attention mechanism, cross-fusion on the fused acoustic feature to generate a corrected acoustic feature (S108); and inputting the corrected acoustic feature into an end-to-end dialect recognition model to generate a dialect audio recognition result (S110). In the method, complex speech signals and multi-accent features can be efficiently captured and processed, and dialect audio can also be classified and decoded online in real time, thereby improving the accuracy and robustness of speech recognition.
Owner:SHANGHAI QIYUE INFORMATION TECH CO LTD

Audio recognition method and apparatus, device, storage medium and computer program product

Provided is an audio recognition method. The method includes that audio data is encoded to an audio encoding feature; the audio encoding feature is decoded to a first decoded feature; a preset word text and the first decoded feature are encoded to a first text encoding feature, the first text encoding feature including a feature representing semantics of the preset word text and a feature representing semantics of the audio data; and the first text encoding feature is decoded to a predicted audio text, where the predicted audio text represents the semantics of the preset word text and the semantics of the audio data. An audio recognition apparatus and a storage medium are also provided.
Owner:MASHANG CONSUMER FINANCE CO LTD

Audio processing model quantification method and system

The invention relates to the technical field of AI audio chips, and discloses an audio processing model quantification method and system, and the method comprises the steps: applying an incremental voltage pulse to a multi-layer memristor cross array, measuring a conductivity value, and generating a memristor quantification configuration table; programming the neuron units in the multi-layer memristor cross array to obtain quantized neuron units; calculating a synaptic weight matrix; the audio signal to be processed is input into the quantization neuron unit to execute pulse calculation, the audio features are obtained, and the audio recognition result is generated on the end-side processor, so that the problems of nonlinearity and element difference of the memristor are effectively solved, the robustness of the system is improved, and the optimal balance of precision and power consumption can be realized in different scenes.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Multi-modal emotion recognition method and device, electronic equipment, storage medium and product

The embodiment of the invention provides a multi-mode emotion recognition method and device, electronic equipment, a storage medium and a product, and relates to the technical field of emotion recognition. The method comprises the following steps: acquiring a to-be-recognized audio / video which comprises an audio stream and a video stream, segmenting the audio stream to obtain at least one audio segment, inputting each audio segment into an audio recognition model to obtain an audio recognition result, determining a corresponding video segment in the video stream according to a target audio segment of which the audio recognition result is an emotion result, and inputting the video segments into a video recognition model to obtain a video recognition result, and determining a target emotion result of the to-be-recognized audio and video based on the audio recognition result and the video recognition result. According to the embodiment of the invention, the video emotion recognition is used for assisting the audio emotion recognition to complete the emotion recognition of the audio and the video, errors possibly caused by single audio recognition are avoided, and the recognition accuracy can be improved.
Owner:BEIJING VISION WORLD TECH CO LTD

Audio identification method and device based on marking and backtracking correction

The invention relates to an audio recognition method and device based on marking and backtracking correction, and the method comprises the steps: segmenting a target audio, and enabling adjacent segments after segmentation to have a partial overlapping region; if a slice point exists in a non-mute segment of the target audio and the acoustic feature similarity of a preset number of frames before and after the slice point is smaller than a set similarity threshold value, the slice point is marked as a cut-off risk point, and the cut-off risk point is used for indicating that continuous semantics before and after the slice point has a cut-off risk; after a text corresponding to the audio of each slice is recognized through the speech recognition model, the language confidence degree of an overlapping area where the truncation risk point is located is determined, and the language confidence degree is used for indicating context language logic of the overlapping area; and if it is determined that the language confidence is lower than a set confidence threshold, correcting the recognition texts of the slices before and after the truncation risk point. According to the invention, the accuracy of audio recognition is improved.
Owner:FIBOCOM WIRELESS

Method and device for detecting and judging bullying behavior based on audio and behavior feature recognition

The invention provides a bullying behavior detection and judgment method and device based on audio and behavior feature recognition, and the method comprises the steps: extracting and analyzing the emotion, keyword and abnormal sound features through an audio recognition module, and judging whether a bullying-related language or audio behavior exists in real time; the behavior recognition module is used for behavior data processing and time sequence feature extraction, behavior classification is performed in combination with a GRU network, and detection and recognition of bullying related behaviors are completed; in combination with the recognition results of the audio recognition module and the behavior recognition module, whether the bullying behavior occurs in the current scene is judged, and if the judgment result is yes, an alarm is given and corresponding personnel are notified to go to the current scene; otherwise, recording the identification result to the cloud. The beneficial effects of the invention are that through multi-dimensional technical means and process design, the bullying behavior can be efficiently, accurately and sustainably monitored, pre-warned and recorded, and powerful technical guarantee is provided for building a safe and friendly environment.
Owner:TIANJIN JINHU DATA CO LTD

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Method and system for recognizing and positioning abnormal sound of livestock and poultry

The invention provides a livestock and poultry abnormal sound recognition and positioning method and system, and relates to the technical field of livestock and poultry breeding monitoring, and the method comprises the steps: extracting the sound characteristics of a livestock and poultry sound signal, carrying out the abnormal audio recognition according to the characteristic pattern of the sound characteristics, and obtaining the sound type of the sound signal; when the sound category is abnormal audio, calculating a generalized cross-correlation delay inequality to determine a sound source position range of the sound signal; and calling a particle swarm optimization algorithm to search the generalized cross-correlation delay inequality to obtain a distance value of an optimal sound source position of the sound signal, calling an SRP-PHAT algorithm to search to obtain a sound source position space point of the sound signal in a sound source position range, and finally determining a three-dimensional position coordinate when the livestock and poultry make a sound. Through the method and the device, the defects that livestock abnormal audio recognition is easily interfered by factors such as external environment noise and the like in the prior art, and the sound source positioning precision is limited in a multi-sound-source scene are overcome.
Owner:BEIJING RES CENT FOR INFORMATION TECH & AGRI

Screen operation method, system and equipment based on air blowing and medium

The invention relates to a blowing-based screen operation method, system and equipment and a medium, and belongs to the field of blowing detection. The method comprises the following steps: acquiring a blowing audio and a face image of a user, and performing face contour detection on the face image based on a face contour recognition algorithm; when the face contour of the user meets a preset detection condition, whether the mouth shape of the user in the face image is a blowing mouth shape or not is judged based on a blowing mouth shape recognition algorithm; when the user is in the blowing nozzle type, the identity of the user is recognized according to the blowing audio of the user based on a blowing voiceprint recognition algorithm; wherein the identity of the user comprises a registered user and a non-registered user; when the identity of the user is a registered user, acquiring a sight focus from the face image based on an eyeball focus recognition algorithm; and executing an operation of the corresponding screen area according to the sight focus. The accuracy of blowing behavior detection is greatly improved, and effective blowing interaction between the screen and the user is achieved.
Owner:ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1

Intelligent sound signal sensing method and system based on time-frequency feature fusion

The invention discloses an intelligent sound signal sensing method and system based on time-frequency feature fusion, and relates to the technical field of intelligent sound signal processing, and the method comprises the steps: obtaining an original audio signal; performing short-time Fourier transform on the original audio signal to obtain a plurality of frequency spectrum features; inputting the plurality of frequency spectrum features into a plurality of neural networks for processing, and extracting frequency domain depth features; processing the original audio signal by using a one-dimensional convolutional neural network, and extracting time domain depth features; splicing and fusing the frequency domain depth feature and the time domain depth feature in a channel dimension to obtain a time-frequency fusion feature; and inputting the time-frequency fusion feature into a full connection layer classifier to obtain a final audio scene recognition result. According to the invention, by designing a double-branch network architecture of time domain and frequency domain parallel processing, time domain waveform information and frequency domain amplitude and phase information are effectively fused, and in an end-to-end training scene, the accuracy of audio recognition is remarkably improved.
Owner:ZHEJIANG UNIV +1

Audio data processing method and device, equipment and storage medium

The invention relates to the field of audio data analysis, in particular to an audio data processing method and device, equipment and a storage medium. The method comprises the following steps: carrying out multi-angle audio continuous acquisition and scene noise suppression on a conference room through a surrounding microphone array, and generating a filtering optimization audio signal; performing three-dimensional time difference positioning calculation according to the filtered and optimized audio signal to obtain accurate sound source positioning information; personalized voiceprint analysis is carried out on the filtered and optimized audio signals, audio distribution modeling is carried out based on accurate sound source positioning information, and a multi-person conference audio field is constructed; performing parallel audio stream separation according to the multi-person conference audio field to generate an intelligent spliced audio segment; and performing deep semantic analysis and semantic logic correction on the intelligent spliced audio segment to generate an audio analysis result. According to the invention, rapid and accurate speaker audio recognition of a multi-person parallel conference is improved.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Harmful voice detection method in combination with side language information

The invention provides a harmful voice detection method in combination with side language information, and the method comprises the steps: collecting a multi-source voice sample, carrying out the preliminary screening and automatic marking of the multi-source voice data through a multi-modal model, and obtaining a preliminary marking sample; performing manual annotation and data reorganization based on the preliminary annotation sample to obtain a high-quality annotation data set, and performing audio unified preprocessing on the high-quality annotation data set to obtain standardized input; based on standardized input, self-supervised high-dimensional features are extracted, and a dual-task model topology is constructed to perform audio recognition; and for the dual-task model topology, carrying out source, category and joint capability training to obtain a joint optimization model so as to output a source label and a type label of harmful information of the to-be-detected audio. According to the method, the optimal performance of a text harmful scene can be considered through high-quality multi-dimensional data support and source and category dual judgment, and the paranguage harmful detection capability is remarkably improved.
Owner:ZHEJIANG UNIV +1

Data processing method, device and equipment and readable storage medium

The invention discloses a data processing method, apparatus and device, and a readable storage medium. The method comprises the steps of obtaining dubbing audio data of original text data; performing audio detection processing on the dubbing audio data according to a sound detection rule to obtain sound segments in the dubbing audio data; performing audio recognition on the sound fragment to obtain recognition text data of the sound fragment; obtaining a text similarity between the original text data and the recognition text data, and determining a quality evaluation value of the sound segment according to the text similarity; the quality evaluation value of the sound segment is used to assist in quality rating of the sound segment. The embodiment of the invention can be applied to various scenes such as the map field, the traffic field, the automatic driving field, the vehicle-mounted scene, the cloud technology, the artificial intelligence, the intelligent traffic and the auxiliary driving, and the corpus generation efficiency in the voice synthesis service is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Wind-noise-resistant self-adaptive volume adjusting method and system for riding earphone

The invention relates to the technical field of audio signal processing, and discloses a wind-noise-resistant self-adaptive volume adjustment method and system for a riding earphone, and the method comprises the following steps: S1, collecting environment sound and a source audio signal in parallel; s2, analyzing environment sound, and determining macroscopic adjustment intensity; s3, analyzing the source audio in parallel, and identifying a content type and a transient signal of the source audio; s4, based on the psychological acoustic model, generating a frequency-related target method tone quality compensation gain; s5, performing dynamic smoothing and frequency weighting on the target gain according to the content and the transient characteristics to generate an application gain; s6, fusing the macroscopic adjustment intensity and the application gain, and synthesizing a final gain parameter; and S7, applying the final gain to the source audio, and outputting the compensated audio. According to the method, the psychological acoustic model and audio content parallel analysis are fused, frequency-level accurate compensation is realized, the problems of traditional degradation and sudden hearing sense change are solved, and the audio definition and comfort are improved.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD

Classroom teaching comprehensive quality evaluation method and system based on teacher and student behavior multi-mode intelligent analysis

The invention discloses a classroom teaching comprehensive quality evaluation method and system based on teacher and student behavior multi-mode intelligent analysis, and the method comprises the steps: firstly obtaining a multi-view classroom teaching video, and constructing a multi-view teacher behavior data set, a teacher behavior audio data set and a student behavior data set; secondly, constructing a multi-view behavior recognition network based on a Mama structure, and respectively obtaining a teacher behavior visual recognition model, a teacher behavior audio recognition model, a teacher behavior multi-modal fusion model and a student learning state detection model through the three data sets; and then acquiring a multi-view classroom teaching video in real time, and based on the model, obtaining a comprehensive teacher behavior identification result and the concentration degree of the students. And finally, according to the teacher behavior comprehensive identification result and the concentration degree of the students, performing quantitative evaluation on the classroom teaching video, and outputting a classroom comprehensive quality evaluation result. According to the method, the problem of one-sided single-modal analysis is solved, and the balance between the multi-view behavior recognition precision and the running speed is realized.
Owner:HANGZHOU DIANZI UNIV

Video processing method and device, computer equipment and storage medium

The invention relates to a video processing method and device, computer equipment and a storage medium. The method comprises the following steps: carrying out background sound and human sound separation on original audio data corresponding to a to-be-processed video, carrying out line feature extraction on the separated human sound audio data to obtain line information with line timestamps, and segmenting the to-be-processed video into a plurality of video segments according to the line timestamps, and segmenting the original audio data into a plurality of audio clips, and synthesizing the face recognition result of each video clip and the audio recognition result of the audio clip corresponding to each video clip to determine a speaking object of each video clip. The recognition accuracy of the speaking object can be improved by eliminating the background sound, and the speaking object in each video clip is comprehensively determined in combination with three kinds of modal information of audio, video and line text, so that the detection accuracy of the speaking object can be greatly improved; the objective of the invention is to solve the problem of low accuracy of speaking object detection for video data in the prior art.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Audio recognition method and device and electronic equipment

The invention provides an audio recognition method and device and electronic equipment. The method comprises the steps of obtaining to-be-recognized audio; determining at least one first audio clip and at least one second audio clip from the to-be-recognized audio through a pre-trained first recognition model; the ratio of the song duration in the first audio clip meets a first specified condition, and the ratio of the song duration in the second audio clip meets a second specified condition; determining a first probability that the first audio clip belongs to the AI flipping audio and a second probability that the second audio clip belongs to the AI generated audio through a second identification model; determining an audio recognition result of the to-be-recognized audio according to the first probability and the second probability; the audio recognition result is AI audio turning and singing, AI generated audio or normal audio. According to the method, audio recognition is carried out on the singing sound segments and the non-singing sound segments, misleading of the non-singing sound segments to the recognition result is avoided, the recognition accuracy of the singing sound segments is improved, and then the recognition accuracy of the whole audio is improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Audio recognition method, method, apparatus for positioning target audio, and device

ActiveUS12361935B2Multi-channel direction findingSpeech recognitionSound sourcesAudio frequency
This application discloses a method for positioning a target audio signal by a computer device. The method includes: performing echo cancellation on the audio signals collected in a plurality of directions in a space, the audio signals comprising a target-audio direct signal; obtaining weights of a plurality of time-frequency points in the echo-canceled audio signals, a weight of each time-frequency point indicating a relative proportion of the target-audio direct signal in the echo-canceled audio signals at the time-frequency point; obtaining a weighted audio signal energy distribution of the audio signals in the plurality of directions by using the weights of the plurality of time-frequency points in the echo-canceled audio signals; and obtaining a sound source azimuth corresponding to the target-audio direct signal in the audio signals by using the weighted audio signal energy distribution of the audio signals in the plurality of directions.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Auditory coal gangue identification method based on array multiple beams

The invention provides a coal gangue identification method, which comprises the following steps of: firstly, determining the layout of an auditory sensor array; then, coal gangue audio signals are collected, and array element receiving signals are obtained; then, implementing a beam forming algorithm on each array element receiving signal to obtain an audio feature representing coal gangue difference information in a beam forming signal; processing the obtained audio features to obtain a multi-dimensional tensor matrix x containing the coal gangue audio features; and finally, by taking the multi-dimensional tensor matrix chi as input, identifying the coal gangue by using a deep learning network to obtain a coal gangue identification result. According to the auditory coal gangue recognition method provided by the invention, the auditory sensor array is adopted, and the influence of factors such as noise and underground equipment interference on coal gangue audio recognition can be remarkably reduced by performing beam forming on the audio signals of the target area; and the classification identification model is combined to carry out intelligent identification on the coal gangue, so that the coal gangue identification accuracy is improved.
Owner:HUANENG COAL TECH RES CO LTD +1

Audio recognition model training method and audio recognition method

The present application discloses an audio recognition model training method and an audio recognition method, including: obtaining a plurality of historical audio data and inputting the plurality of historical audio data into an audio recognition model to be trained; extracting frequency domain features from each historical audio data to obtain a plurality of initial audio features; calculating, based on each initial audio feature, a set of target weights that matches the initial audio feature; obtaining a preset parameter set and calculating, based on the preset parameter set and each set of target weights respectively, a plurality of fusion parameters; processing, based on each fusion parameter, the initial audio feature that matches each fusion parameter to obtain a plurality of target audio features; training the audio recognition model to be trained based on the plurality of target audio features to obtain a target audio recognition model. The technical solution of the present application calculates the weights of each audio respectively, improving the recognition accuracy of the model.
Owner:TP-LINK INT CHENGDU CO LTD

Voice command recognition method, device, equipment and storage medium for intelligent elevator

The present disclosure relates to a method, apparatus, device and storage medium for recognizing voice commands of an intelligent elevator, the method comprising: collecting audio emitted by a user in an elevator car; recognizing command words in the audio; determining whether a first audio segment before a target audio segment and / or a second audio segment after the target audio segment in the audio are emitted by the user who emitted the target audio segment, the target audio segment being an audio segment containing a command word in the audio; if not, determining that the command word is valid, and executing the corresponding elevator command according to the command word. The present disclosure takes into account that the elevator commands spoken by the user are usually concise and have fewer words. If, in the collected audio, the audio segment before the command word and / or the audio segment after the command word are spoken by the same user as the command word, it is considered that the user is chatting, and the intention of speaking the command word to take the elevator is not credible. Otherwise, the command word is considered valid, thereby reducing the probability of the intelligent elevator system being mistakenly awakened by voice commands.
Owner:SOUNDAI TECH CO LTD

Automatic string music tremolo detection method

The invention belongs to the technical field of audio recognition, and particularly relates to a string music tremolo automatic detection method which comprises the following steps: S1, preprocessing and framing an input audio, and extracting time sequence features; s2, obtaining an emotion intensity score through an emotion intensity evaluation function; s3, mapping the emotion intensity score into an expected tremor speed, and obtaining a stable tremor speed through exponential smoothing and amplitude limiting; s4, generating a dynamic threshold value changing along with time; and S5, the main tremolo detector calculates the criterion in parallel, compares the criterion with a dynamic threshold value on a synchronization time axis to determine frame-level tremolo, connects adjacent frames to obtain a tremolo segment, and drives threshold value self-adaption by emotional intensity, so that the detection criterion is synchronized with a playing situation, wrong division and over-division are reduced, and the boundary consistency and the stability and accuracy of tremolo estimation are improved.
Owner:SHANGQIU NORMAL UNIVERSITY

Binary natural voice audio homework correction method and system in AI auxiliary confrontation mode

The invention relates to the technical field of information processing, in particular to a binary natural voice audio homework correction method and system in an AI auxiliary confrontation mode, firstly, a confrontation mode and audio homework parameters are set, and the confrontation mode is one of a man-machine confrontation mode, a double-person multi-resistance mode, a three-person confrontation mode and a multi-person confrontation mode with more than three persons; the audio operation parameters comprise an audio type, an audio speed and audio time; determining a to-be-read text according to the audio operation parameters, obtaining an adversarial audio, performing scoring by using the audio recognition model, and obtaining adversarial result data according to scores, and the audio recognition model adjusts the standard audio of the to-be-read text according to the audio feature parameters of the confrontation participants, compares the audio of the confrontation participants with the adjusted standard audio, and scores and outputs the audio of each confrontation participant according to a comparison result. According to the invention, through the AI auxiliary confrontation mode, on the premise that the binary natural voice audio homework is corrected rapidly and accurately, the active activity of the audio homework participants is improved.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

A method and system for remotely recognizing lip language by a drone

PendingCN122369454AData streamNoise
This invention provides a method and system for long-range lip-reading recognition using unmanned aerial vehicles (UAVs), belonging to the fields of artificial intelligence and UAV technology. The system utilizes a camera, IMU (Integrated Mutor Unit), and microphone mounted on the UAV to acquire video streams, IMU data streams, and audio streams. Motion compensation is applied to the video stream using IMU data to stabilize image frames. Then, image sequences of the target speaker's lips and nasal / jaw region are extracted and input into a lip-reading recognition model to obtain initial recognition text and visual confidence scores. Simultaneously, the signal-to-noise ratio (SNR) of the audio stream is calculated. When the visual confidence score is low and the SNR is high, the audio recognition process is triggered, generating audio recognition text and an audio confidence score. Finally, the visual and audio confidence scores, along with the lip movement-audio temporal matching degree, are weighted and fused to generate the final recognition text. This invention effectively solves the problems of video instability and background interference in long-range lip-reading recognition, improving recognition accuracy in complex environments.
Owner:FUZHOU PLANNING DESIGN & RES INST

Image capturing apparatus having audio recognition, control method thereof, and storage medium

An image capturing apparatus obtains audio of an utterance that occurs in a vicinity of the image capturing apparatus, captures an image, and controls image transmission such that in response to a determination that an expression indicating a particular person is included in the audio of the utterance, transmits a first image related to the obtainment of the audio of the utterance among captured images to an external apparatus associated with the expression indicating the particular person. In addition, a second image captured in the external apparatus and related to playing of the first image is received from the external apparatus.
Owner:CANON KK

Machine learning based video recognition method, device, server and storage medium

Embodiments of the present application disclose a video recognition method and device based on machine learning, a server and a storage medium. Embodiments of the present application can obtain a target video; obtain a source video corresponding to the target video, the target video being created by processing the source video; compare the target video and the source video in content to obtain a content type of the target video; when the content type of the target video is a funny content type, perform audio recognition on the target video and the source video to determine an audio type of the target video; when the audio type of the target video is a funny dubbing type, determine the target video as a funny dubbing video, so as to push the funny dubbing video to a user. Embodiments of the present application compare the target video with the source video in dimensions such as content and audio to identify whether the target video is a funny dubbing video created by processing the source video. Thus, the present application can accurately identify a funny dubbing video from a plurality of videos, and improves the efficiency of video recognition.
Owner:TENCENT TECH (BEIJING) CO LTD

Audio recognition method and device, electronic equipment and storage medium

The invention provides an audio recognition method and device, electronic equipment and a storage medium, and belongs to the technical field of audio processing, and the method comprises the steps: carrying out the Fourier transform of a non-voice signal, obtaining a noise frequency domain signal, and calculating the first power spectrum density of a background noise signal in the non-voice signal; performing Fourier transform on the voice signal to obtain a voice frequency domain signal, and calculating a second power spectral density corresponding to the initial voice signal based on the voice frequency domain signal and the first power spectral density; calculating a power ratio of an initial voice signal to a background noise signal in the voice signals; performing signal enhancement on the initial voice signal based on the power ratio, the voice frequency domain signal, the first power spectral density and the second power spectral density to obtain a target voice signal; and identifying the target voice signal based on the target model. According to the audio recognition method and device, the electronic equipment and the storage medium provided by the invention, the audio recognition precision can be improved.
Owner:BEIJING SUPERHEXA CENTURY TECH CO LTD

Audio processing method, device and system and storage medium

The invention relates to the technical field of audio processing, and provides an audio processing method, device and system and a storage medium, and the method comprises the steps: receiving a mixed audio signal which at least comprises two kinds of audio signals; identifying the mixed audio signal through at least one audio identification module to identify a target audio signal in the mixed audio signal; and carrying out audio processing on the target audio signal by adopting corresponding preset processing rules through the at least two audio processing modules. By adopting an intelligent identification and differential processing mechanism, on the premise of keeping the integrity of the original audio, different target audio signals are subjected to respective corresponding better audio processing, for example, in a game scene, environment sound effects such as shocking explosive sound and fine footstep sound are enhanced in an immersive low-frequency mode, and the effect of improving the sound quality is achieved. And meanwhile, the teammate voice instruction is kept clear and transparent without interference, so that the audio processing effect is greatly improved.
Owner:GUANGDONG DINGCHUANG SMART MANUFACTURING CO LTD

Speech determination model training method, object speech extraction method, device, electronic equipment and storage medium

The present disclosure relates to a speech determination model training method, an object speech extraction method, a device, an electronic device and a storage medium. The method comprises: obtaining a second speech set, the second speech set comprising a plurality of second sample speeches and a second object identifier corresponding to each second sample speech, each second sample speech comprising a fifth speech and a sixth speech, the fifth speech comprising at least speech corresponding to a fifth speech non-associated object identifier, and the sixth speech being silent audio; performing speech extraction on the fifth speech in each second sample speech based on a speech extraction model to obtain a seventh speech corresponding to each second sample speech; and training a candidate speech determination model based on the seventh speech and the sixth speech to obtain a target speech determination model. The present application increases the training samples by adding the speech corresponding to the fifth speech non-associated object identifier and the silence, trains the candidate speech determination model, obtains a model suitable for richer scenarios, and reduces the error rate of audio recognition.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD