Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

403 results about "Vocal sound" patented technology

The human voice consists of sound made by a human being using the vocal tract, such as talking, singing, laughing, crying, screaming, etc. The human voice frequency is specifically a part of human sound production in which the vocal folds (vocal cords) are the primary sound source.

Electric cooker anti-noise voice interaction system based on multi-mode perception and control method

The invention relates to the technical field of intelligent household appliances and man-machine interaction, and discloses an electric cooker anti-noise voice interaction system based on multi-mode perception and a control method, and the system comprises a multi-mode perception and collection unit which synchronously collects millimeter wave radar echo signals and acoustic signals; the signal preprocessing and feature extraction unit is used for extracting a user physiological vibration signal and a three-dimensional space position vector from the radar signal and extracting an acoustic energy envelope and a sound source direction vector from the acoustic signal; the time-space consistency verification unit and the voice gating and recognition unit are used for carrying out time synchronization verification by calculating the correlation between the physiological vibration signal and the acoustic energy envelope, and distinguishing real human voice from an environment false trigger source; meanwhile, space consistency verification is carried out by comparing the user direction of radar positioning with the sound source direction of acoustic positioning, so that non-target human voice interference is eliminated. According to the invention, the anti-interference capability and reliability of voice interaction in a real home environment are improved through double verification on a physical level.
Owner:LINGNAN NORMAL UNIV

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Human voice activity detection method and device, computer equipment and storage medium

The invention relates to the field of audio processing, and discloses a human voice activity detection method and device, computer equipment and a storage medium. The method comprises the following steps: collecting audio and video data of a user, and dividing the audio and video data into a video stream and an audio stream; detecting the video stream by using a FaceMesh model to obtain a lip opening degree and a head attitude angle, and determining a lip motion state according to the lip opening degree and the head attitude angle; blocking the audio stream to obtain a plurality of audio blocks, performing noise reduction on each audio block to obtain a plurality of noise-reduced audio blocks, and detecting each noise-reduced audio block by using a silhoVAD model to obtain a human voice activity probability and an audio cache queue; and inputting the lip motion state and the human voice activity probability into a state machine, introducing an audio cache queue by the state machine, and outputting a speaking identifier and a corresponding audio clip. Speaking recognition is cooperatively completed in combination with machine vision and hearing signals, and the accuracy of whole human voice activity detection is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

A Human-Computer Voice Interaction Control Method and System Based on Smart TV

This application relates to a human-computer voice interaction control method and system based on a smart TV. The method includes acquiring voice data within a preset range, processing the voice data to obtain voice feature data carrying control commands; performing feature analysis on the voice feature data using a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results; sending a secondary confirmation request to the user based on the control command comparison results, and combining the confirmation voice information from the user feedback to perform command recognition evaluation and correct command deviation processing to obtain the corrected control command; performing function switching processing on the TV according to the correct control command, and optimizing the display effect by linking and adjusting related devices based on the program display requirements after the switch to obtain human-computer voice interaction control data. This application has the effect of improving the intelligence of voice interaction control of smart TVs.
Owner:GUANGZHOU XIANYOU INTELLIGENT TECH CO LTD

Desktop robot microphone array sound source positioning system and method

The invention relates to the technical field of robots, and particularly discloses a desktop robot microphone array sound source positioning system and method. Comprising a main control chip, acquisition microphones and a power supply module, an ESP32 main control board is connected with a circular array formed by the four acquisition microphones through double I2S bus pins, the array diameter of the four acquisition microphones is 50 mm, the included angle between every two adjacent acquisition microphones is 90 degrees, the acquisition microphones are grounded through GND lines, and the power supply module is connected with the ESP32 main control board. And the ESP32 main control board provides 3.3 V voltage for the acquisition microphone through a power supply line. A four-microphone circular array with the diameter of 50 mm and ESP32 double I2S synchronous acquisition are adopted, 360-degree full horizontal positioning is achieved, the hardware size is small, the miniaturization requirement of a desktop accompanying robot is met, a VAD voice activity detection module is introduced, algorithm operation is triggered only when human voice is detected, power consumption in a dormant state is small, and the endurance of the robot is prolonged.
Owner:HANGZHOU XINGMENGDAO TECHNOLOGY CO LTD

Music source separation method and system based on hybrid expert self-attention network

The invention discloses a music source separation method and system based on a hybrid expert self-attention network, and belongs to the technical field of audio signal processing and deep learning. The method comprises the following steps: receiving a time domain mixed audio signal, and obtaining a complex frequency spectrum through short-time Fourier transform; the method comprises the following steps: estimating a complex ideal proportion mask through an improved separator network MoEFormer, respectively modeling a time domain dependency relationship and a frequency domain dependency relationship in the separator network by adopting an axial attention mechanism, and replacing a feedforward network in a standard Transformer with a hybrid expert layer; the hybrid expert layer comprises an expert network specially designed for drum sound, Bass and human voice characteristics, and a self-adaptive weight distribution mechanism based on gating routing; and finally, reconstructing the separated time domain signal through inverse short-time Fourier transform. Through expert specialization and feature adaptive fusion, the problem of multi-sound-source feature confusion is effectively solved, and low calculation complexity is kept while the separation precision is improved.
Owner:JINLING INST OF TECH

system

We provide the system. [Solution] A means for analyzing voice and video information of a person acquired by the user and generating a three-dimensional model of the said person, A means for reproducing the voice of the person based on the aforementioned audio information, A means to enable dialogue with the person in a virtual space using a generated three-dimensional model and reproduced voice, A system that includes this.
Owner:SOFTBANK GROUP CORP

Data processing method and computer equipment

The invention provides a data processing method and computer equipment, which can be applied to the field of audio and video processing, and comprises the following steps: determining first azimuth information of an abnormal event occurring in a physical space through a shooting device, and guiding the shooting device and a pickup device to perform linkage parameter adjustment according to the first azimuth information so as to focus a target area when the abnormal event occurs, the purpose of focusing is to only collect audios and videos in a target area. And when the sound source sounds in the target area, determining a first sub-area (the area where the person in the sound production state in the abnormal event is located) from the target area, and adjusting the pickup parameter again to enable the pickup range to be limited in the first sub-area. According to the method, firstly, the sound curtain of the area where the abnormal event is located is constructed through the parameter coarse adjustment process (namely, only audio and video in the target area are collected, and sound outside the target area is shielded), when the human voice is recognized in the target area, the parameters are further finely adjusted to achieve accurate directional pickup of the human voice in the abnormal event, and the audio and video evidence obtaining quality is effectively improved.
Owner:HUAWEI TECH CO LTD

Speech recognition methods, devices, electronic equipment, media and software products

This application discloses a speech recognition method, apparatus, electronic device, medium, and program product, belonging to the field of audio technology. The method includes: acquiring an identification error corresponding to a secondary path of an audio device, the identification error being used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals; if the identification error is greater than or equal to a first threshold, acquiring a first spectral ratio, the first spectral ratio being the ratio of the spectra of the signals collected by at least two microphones; if the first spectral ratio is greater than or equal to a second threshold, determining that the signals collected by at least two microphones were emitted by a user wearing the audio device.
Owner:VIVO MOBILE COMM CO LTD

An adjustable frequency response microphone with integrated XLR and USB outputs

This invention discloses an adjustable frequency response microphone integrating XLR and USB outputs, comprising a microphone body, a dynamic microphone driver, a passive filter, a pass-through / preamp circuit, an XLR XLR interface, a USB module, and a function control circuit. The passive filter employs an LC circuit structure, supporting three frequency modes: low-frequency cutoff, mid-frequency boost, and flat response, effectively filtering out ambient noise and enhancing vocal quality. The pass-through / preamp circuit offers both pass-through and preamp modes, with adjustable preamp gain to accommodate dynamic microphone drivers of varying sensitivities, and can be externally or internally powered by phantom power. The USB module integrates a Type-C interface and an audio chip, supporting analog-to-digital / digital-to-analog conversion, real-time headphone monitoring, and external power supply. The function control circuit integrates an encoder and a touch panel, enabling adjustments such as volume, mixing, and mute, and enhances the interactive experience with an RGB lighting module. This invention combines the advantages of professional XLR output with convenient USB output, making it suitable for professional recording, live streaming, gaming, and other scenarios.
Owner:ZHAOQING HEJIA ELECTRONICS CO LTD

Method for spatializing a sound stream including a sound object in a motor vehicle

A method for spatializing a sound stream including a sound object in a motor vehicle comprises the steps of: extracting (203) the voice using an artificial intelligence model to separate the sound stream into two tracks, a first track including the sound object and a second track including the other sound elements of the sound stream; spatializing (205) the first track so that the sound object is perceived by a user of the motor vehicle as originating from a predefined location in the vehicle; and simultaneously broadcasting (209) the first and second tracks in the vehicle. A spatialization device and a motor vehicle including the device are also described. Figure to be published with the abbreviation: Fig 2
Owner:STELLANTIS AUTO SAS +1

Conference picture processing method, device, storage medium and system

This application relates to a conference screen processing method, device, storage medium, and system. It involves acquiring multiple audio data streams using a data acquisition device, determining the sound source location information based on the first audio data stream where the human voice energy exceeds a first threshold, acquiring the image corresponding to the sound source location information, identifying the speaker in the image whose human voice energy meets preset conditions, and obtaining the facial deflection angle between a preset position representing the speaker's face and the data acquisition device. The speaker's image and facial deflection angle are then sent to a data processing device. The data processing device can display the image of the speaker with the smallest facial deflection angle using a display device. This eliminates the need for the data processing device to expend significant computing resources for image recognition, reducing its computational resource consumption and improving the efficiency of displaying the speaker's face at the conference.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Noise reduction devices

To provide a relatively small noise reduction device that can effectively reduce the frequency range of human speech. [Solution] The noise reduction device 100 comprises a ventilation passage 24 for ensuring the wearer's breathing, with the mouth covered by a covering part (protective part 10), and a pair of noise reduction units connected to the ventilation passage 24. One of the pair of noise reduction units is a first noise reduction unit 30 that reduces low-frequency sounds from the mouth guided through the ventilation passage 24 using an active noise cancellation method, and the other of the pair of noise reduction units is a second noise reduction unit 40, 50 that reduces high-frequency sounds higher than low-frequency sounds from the mouth guided through the ventilation passage 24 using an interference tube method or a resonance tube method.
Owner:CASIO COMPUTER CO LTD

A smart accompaniment singing method and system

ActiveCN121122218BEngineeringAudio frequency
This invention belongs to the field of audio processing technology and provides an intelligent accompaniment singing method and system. The method can collect the user's vocal pitch and voiceprint information, then divide the song selected based on the song selection command into vocal data and accompaniment data. The pitch of the accompaniment data is adjusted by comparing the vocal data with the user's vocal pitch, and the audio is played based on the adjusted accompaniment data. The user's singing data is continuously acquired and monitored, and when the user's singing data meets preset conditions, accompaniment vocals generated using the voiceprint information and vocal data are added to the audio. This proposed method can tailor the song's pitch to the user's individual needs, adjusting the accompaniment pitch to a suitable range for the user, ensuring easier singing. It can also selectively add accompaniment vocals based on the user's singing performance, thereby guiding the user to adjust their singing style, improve their singing level, and enhance their singing experience.
Owner:SHENZHEN WANSHENG CULTURE TECH CO LTD

Method, apparatus, device, medium and program product for processing call audio

Embodiments of the application disclose a call audio processing method and device, equipment, medium and program product, and belong to the technical field of spatial audio. The method comprises the following steps: a first terminal collects real-time human voice of a user during a call through at least two microphones, and obtains at least two call audios; spatial audio data is generated based on the at least two call audios, wherein the spatial audio data refers to audio data of real-time position information of the user in a space; and the spatial audio data is sent to a second terminal. The method realizes the restoration of a space field in which original call audio is located when a terminal at one end plays call audio from another end during real-time video call.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Generative audio anonymization reconstruction method and device based on sound source separation and semantic preservation, equipment and program product

The invention discloses a generative audio anonymization reconstruction method, device, equipment and program product based on sound source separation and semantic preservation, and relates to the technical field of voice privacy protection and audio signal processing. The method comprises the following steps: acquiring an original audio, and carrying out sound source separation on the original audio to obtain at least one speaker sound track and an environment background sound track; performing authentication on each speaker sound track, and determining an authentication result of each speaker sound track; wherein the authentication result of the speaker sound track comprises an authorized person sound track and an unauthorized person sound track; carrying out anonymization processing on the unauthorized human voice track to obtain an anonymized human voice track; and carrying out re-synthesis on the anonymized human voice track and the environment background audio track to obtain an anonymized scene audio track. According to the technical scheme provided by the embodiment of the invention, an audio stream breakage phenomenon can be avoided, the scene continuity of the audio is improved, and the intelligibility and the overall quality of the audio are further improved.
Owner:SHENZHEN JIAYZ PHOTO IND LTD

Audiovisual content rendering with display animation suggestive of geolocation at which content was previously rendered

Techniques have been developed to facilitate the capture of performances on handheld or other portable computing devices and, in some cases, the pitch-correction and mixing of such vocal performances with backing tracks for audible rendering on such devices. Captivating visual animations and / or facilities for listener comment and ranking are provided in association with an audible rendering of a performance, e.g., a vocal performance captured and pitch-corrected at another similarly configured mobile device and mixed with backing instrumentals and / or vocals. Geocoding of captured vocal performances and / or listener feedback may facilitate animations or display artifacts in ways that are suggestive of a performance or endorsement emanating from a particular geographic locale on a user manipulable globe. In this way, implementations of the described functionality can transform otherwise mundane mobile devices into social instruments that foster a unique sense of global connectivity and community.
Owner:SMULE INC

Video translation method and device

The invention provides a video translation method and device, and the method comprises the steps: determining a voice audio segment corresponding to each target subtitle text in a subtitle set of a target language through the subtitle set of the target language and an audio file of a to-be-translated video; then determining an emotion category of a voice audio segment corresponding to the target subtitle text and a long audio of a role to which the target subtitle text belongs; then, based on the target subtitle text, the duration of the target subtitle text, the emotion category of the voice audio fragment corresponding to the target subtitle text and the long audio of the role to which the target subtitle text belongs, generating a target voice audio fragment corresponding to the target subtitle text; and finally, generating a translated video based on the target human voice audio segments corresponding to all the target subtitle texts, thereby remarkably improving the quality and availability of automatic dubbing, and effectively improving the watching experience.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Sound box volume intelligent adjustment method and device combining environmental noise monitoring and voice recognition

The present application relates to a sound box volume intelligent adjustment method combining environmental noise monitoring and voice recognition, comprising the following: obtaining a first decibel value and a second decibel value of adjacent sampling intervals without human voice interference respectively; calculating the difference between the second decibel value and the first decibel value to obtain a decibel change value; judging the environmental noise change based on the decibel change value, if the environmental noise change is larger or smaller, triggering the voice recognition function, obtaining the feedback voice of the user and based on this to adaptively adjust the sound box volume. The present application uses the difference between the second decibel value and the first decibel value without human voice interference to approximately estimate the environmental noise change, when the environmental noise changes, the voice recognition function is triggered, the voice interaction with the user is carried out, and the sound box volume is adaptively adjusted, on the one hand, the volume can be adjusted in time when the noise fluctuates to adapt to the user's listening to the sound box content, on the other hand, the interaction with the user increases the user's experience.
Owner:FOSHAN CHANSTEK SOUND EQUIP CO LTD

KTV song requesting system and method for requesting songs based on microphone voice instruction analysis

The invention relates to the technical field of intelligent voice recognition, and discloses a KTV song requesting system and method for requesting songs based on microphone voice instruction analysis, and the system comprises a wireless microphone, a microphone receiver and a song requesting set top box. The wireless microphone is used for distributing a human voice signal into a sound amplification signal flow and an instruction signal flow which are parallel in real time during a key triggering period; the microphone receiver encapsulates the instruction signal flow into a data frame with a channel state identifier and audio data, and is connected to the song requesting set top box through a special data transmission line; the song requesting set top box comprises a synchronous sampling module and an instruction analysis engine, the synchronous sampling module is used for synchronously sampling the echo reference signal according to the channel state identifier, and the instruction analysis engine performs adaptive echo cancellation processing and semantic recognition on the audio data based on the echo reference signal so as to execute song requesting operation. The method is applied to the system. Interference such as background music can be effectively eliminated, and the voice instruction recognition accuracy is improved.
Owner:HUNAN SIFANG FRIENDS TECHNOLOGY CO LTD

Human voice main melody extraction method and device, electronic equipment and storage medium

The embodiment of the application provides a kind of human voice main melody extraction method and device, electronic equipment and storage medium, belong to the field of financial technology.The method comprises: obtaining harmonic splicing data;Harmonic splicing data is input to the preset original melody extraction model to obtain audio saliency data by saliency calculation, and audio saliency data is discriminated to obtain human voice discrimination information;According to audio saliency data and human voice discrimination information, sample audio data is extracted to obtain human voice main melody sequence;Human voice main melody sequence, human voice discrimination information and preset main melody reference sequence, discrimination reference information are calculated to obtain target loss data;According to target loss data, the original melody extraction model is adjusted to obtain the target melody extraction model;Target splicing data is input to the target melody extraction model to extract the target main melody sequence.The embodiment of the application can improve the extraction effect of human voice main melody.
Owner:PING AN TECH (SHENZHEN) CO LTD

Oral doctor's order intelligent processing method based on ai chest plate recorder system and related device

The application provides an oral medical order intelligent processing method based on an AI chest badge recorder system and related devices, collects audio data of an emergency scene obtained by a medical staff wearing an AI chest badge, and respectively obtains a doctor's dictation voice signal and a nurse's recitation voice signal through noise reduction enhancement and human voice separation processing; the doctor's voice recognition matching and identity authorization determination are completed relying on a pre-stored voiceprint feature library, and after authorization, the two voice signals are respectively transcribed into medical order text and recitation text; consistency verification is carried out on the two texts, and after the verification is passed, structured medical order data is generated and delivered to an execution terminal to assist nurses to standardize medical operations. The application discards the traditional mode of artificial memory, oral check and after-the-fact paper supplement, improves the accuracy of oral medical order checking, the efficiency of circulation and the standardization of execution, ensures the safety of emergency diagnosis and treatment, realizes the traceability of the whole process of medical order, and adapts to the application requirements of emergency and emergency rapid treatment.
Owner:SHENZHEN PEOPLES HOSPITAL

Method and equipment for displaying digital human sign language

The invention provides a method and equipment for displaying digital human sign languages. The method can be applied to a first device, and the method can comprise the steps that first audio data are acquired, and the first audio data comprise audio data of human voice; sending the first audio data to a second device; receiving a first action parameter sequence of the digital human sign language sent by the second equipment; the first action parameter sequence is rendered according to the display speed, a multi-frame image of the digital sign language is obtained, the display speed is related to the completion time of rendering a second action parameter sequence corresponding to second audio data by the first device, and the second audio data is audio data of human voice obtained before the first device obtains the first audio data. In the technical scheme, the device can generate the digital human sign language actions in real time according to the human voice when the user watches a video or a television program, and plays the digital human sign language actions in real time, so as to provide sign language translation service for the hearing impaired user in real time.
Owner:HUAWEI TECH CO LTD

Method and apparatus for performing speech enhancement, storage medium, device, and product

PendingUS20260004788A1Speech analysisNeural learning methodsPhonetic environmentNoise
A speech enhancement method, apparatus, and computer-readable storage medium for training neural networks to enhance speech quality. The method obtains a training set containing training samples, each comprising a sample reference speech, a sample comparison speech from the same sound-producing object, and a mixed speech combining interfering human voice, ambient noise, and the sample comparison speech. Sample voiceprint vectors are extracted from reference speech and sample audio features from mixed speech. A speech enhancement network processes these inputs to output predicted audio features, which are compared against comparison audio features to determine training loss values. The network's weight parameters are iteratively updated based on these loss values until training completion, enabling effective speech enhancement through voiceprint-guided processing.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Railway vehicle noise distinguishing and extracting device and method

The invention relates to the technical field of railway vehicle noise analysis and processing, and particularly discloses a railway vehicle noise distinguishing and extracting device and method.The device comprises an acquisition module, a processing module and an output module, and the acquisition module is used for acquiring carriage mixed noise in the running process of a railway vehicle; the processing module is used for performing frequency decomposition, noise classification and independent loudness analysis of various types of classified noise on the collected mixed noise of the carriage; the output module is used for outputting and / or storing the loudness values and the frequency characteristics of various types of noise after classification; according to the method, wheel track noise, passenger voice and train station reporting voice can be accurately separated and subjected to loudness analysis in real time, the problem that the passenger voice and the train station reporting voice are difficult to separate and distinguish is solved, and a scientific noise monitoring and management means is provided for rail transit operation enterprises; the passenger comfort level is improved, the station reporting volume and the operation service quality are optimized, and complaint and transformation cost caused by noise is reduced.
Owner:HEFEI RAIL TRANSIT GROUP OPERATION CO LTD

Audio noise reduction method and system for Bluetooth headset

PendingCN121665155AMicrophonesSignal processingNoiseSpectral subtraction
The invention relates to the technical field of voice enhancement, in particular to an audio noise reduction method and system for a Bluetooth headset, and the method comprises the steps: analyzing the feature condition of noise influence in a mobile scene, including the feature change condition of a signal obtained by a microphone under the condition of pedestrian voice noise interference; according to the method, the spectral subtraction factor of the spectral subtraction method is adjusted in a targeted manner by integrating the characteristic difference conditions of the audios obtained by the reference microphone and the main microphone when the audios are moved to different scenes, so that the filtering error occurring during the audio noise reduction of the Bluetooth headset in the moving scene is avoided, and the audio noise reduction level in the call process of the Bluetooth headset is further improved.
Owner:DONGGUAN YUANZE ACOUSTIC TECH CO LTD

A method for implementing a headphone transparent mode

The application discloses a method for realizing a headphone transparent mode. The prior art method needs to use a device other than the headphone to complete required path measurement, and requires complicated hardware resources, environment and process. The method uses an external microphone and an error microphone of the headphone to complete path measurement of a primary channel, uses a loudspeaker of the headphone and the error microphone to complete path measurement of a secondary channel, establishes a first-stage filter of the transparent mode based on an N-order FIR filter, establishes a cost function, and uses a genetic algorithm to solve coefficients of the FIR filter; a high shelf filter is used to establish a second-stage filter of the transparent mode, and to complete suppression of potential abnormal high-frequency components. The two-stage filters jointly establish a complete transparent mode filter. The application uses conditions of the headphone itself to complete necessary path measurement, concentrates the FIR filter on fitting of effective frequency bands in environmental human voices, and reduces difficulty and complexity of required resources of the transparent mode design.
Owner:HANGZHOU NATCHIP SCI & TECH CO LTD

Method and device for automatically splitting audio of original film for film dubbing

PendingCN122658338ATime domainSystems analysis
The application provides a film dubbing-oriented original film audio automated split processing method and device. The method calls original composite audio track data and performs overlap sampling division to generate a time-domain audio frame sequence and construct an acoustic analysis initial representation data table; injects the same into an acoustic source deep decoupling model to generate a time-frequency domain acoustic mask weight sequence, and combines the initial representation data table to generate a background sound environment energy feature parameter sequence; constructs a role-identified voice split attribute mapping detail table for recording voice attribution; and generates automated film dubbing split audio control instruction stream data for dubbing system analysis based on the mapping detail table and the background sound sequence. Through deep decoupling and role identification mapping, the application solves the technical defects of low audio split purity and inability to realize role-level automatic attribution in traditional schemes, and significantly improves the automation degree and audio track processing quality of film dubbing.
Owner:CHINA FILM IND GROUP CO LTD BEIJING ARTIFICIAL INTELLIGENCE RESEARCH & APPLICATION BRANCH

Intelligent voice interaction system and method

The application provides a kind of intelligent voice interaction system and method, it is related to electronic information technical field, including: acquisition calibration module, for by multichannel acoustic sensor array real-time acquisition mixed audio signal in environment, and set four reference positioning points located at corner point on sensor array, form quadrilateral structure;Based on the area characteristics of quadrilateral generation geometric correction value, and synchronous extraction noise frequency band feature and speech short-time energy distribution characteristics, and utilize geometric correction value to extract the feature and calibrate;Noise identification module is used for based on noise frequency band feature and speech short-time energy distribution characteristics, identify noise type, including wide-band high-intensity noise, low-frequency persistent noise, multi-source human voice interference and burst transient noise, generate label vector, and dynamically adjust noise reduction strategy according to label vector.The application is collected and calibrated audio signal by multichannel acoustic sensor array, improves the accuracy and fluency of voice interaction.
Owner:FUJIAN REIDA PRECISION

A blockchain storage method and system based on audio authenticity identification technology

The present application relates to the technical field of audio discrimination, and particularly relates to a blockchain storage method and system based on audio authenticity discrimination technology; an audio feature extraction model is used to extract audio features of training audio data; an audio authenticity discrimination model is constructed, and the audio authenticity discrimination model is trained by using the audio features of the training audio data; whether target audio data is real human voice is determined by using the trained audio authenticity discrimination model, and when the target audio data is real human voice, the target audio data is sent to an audio management server for storage, so that the safety of the stored audio data is ensured; storage code is sent to a blockchain server for storage, so that the storage code can be prevented from being tampered with; before the target audio data is accessed on a user terminal, authenticity verification of whether the target audio data is tampered with is required, so that the authenticity and safety of the target audio data are further ensured.
Owner:HUNAN MANGO INTELLIGENT MEDIA TECH DEV CO LTD