Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1217 results about "Vocal sound" patented technology

The human voice consists of sound made by a human being using the vocal tract, such as talking, singing, laughing, crying, screaming, etc. The human voice frequency is specifically a part of human sound production in which the vocal folds (vocal cords) are the primary sound source.

Intelligent microphone pickup and speech enhancement method, system and device

The invention relates to an intelligent microphone pickup and speech enhancement method, system and device, and the method comprises the steps: obtaining an audio signal collected by a dual-silicon microphone array, and carrying out the cross-correlation analysis of the audio signal, and obtaining a corresponding human voice signal correlation feature; performing sound source positioning analysis on the audio signal based on the human voice signal correlation feature to obtain a corresponding target human voice signal; performing frequency characteristic analysis on the target human voice signal, and performing segmented dynamic gain processing on the signal according to a preset frequency band range to obtain a corresponding voice enhancement signal; performing noise component adaptive filtering processing on the target human voice signal to obtain a corresponding noise suppression parameter; and inputting the speech enhancement signal and the noise suppression parameter into a preset echo cancellation model for joint optimization to obtain a corresponding output human voice signal. According to the invention, the coupling problem of multiple acoustic interferences can be effectively solved.
Owner:SHENZHEN SHIDU DIGITAL TECH CO LTD

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD

Voice processing method and device based on voiceprint feature screening, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health, voice navigation and the like, and discloses a voice processing method, device, equipment and medium based on voiceprint feature screening. The method comprises the following steps: screening voice of a near-field main speaker based on voiceprint similarity and signal intensity, executing duration filtering and confidence verification, dynamically constructing a voiceprint feature library, generating a voiceprint mask matrix based on the voiceprint feature library, performing frequency band suppression on a to-be-processed voice signal, and outputting a purified voice signal. According to the method, the voiceprint library is dynamically constructed, the voiceprint mask matrix is generated based on the voiceprint feature library, frequency band suppression is performed on the to-be-processed voice signal, non-target voiceprints are effectively shielded, the voice signal quality is improved, and thus target voice is accurately captured in a high-noise environment.
Owner:平安科技(上海)有限公司

Multi-modal time sequence alignment AI video translation method and system

The invention relates to the technical field of subtitle translation, in particular to a multi-modal time sequence alignment AI video translation method and system, and the method comprises the steps: 1, carrying out the multi-modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint feature data and visual time sequence data; 2, performing cross-language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint feature data and the target language text; generating a mouth shape animation matched with the target language voice based on the lip key point data and the limb action time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the mouth shape animation and the limb action sequence through a cross-modal time sequence encoder, and dynamically adjusting the layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.
Owner:HANGZHOU BAOMIHUA TECH CO LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

KTV intelligent light control method and device

The invention discloses a KTV intelligent light control method and device. The method comprises the following steps: S1, extracting song feature data from a song requesting system, wherein the song feature data comprises low-frequency energy EL, a rhythm feature Tbeat, an emotion label Cemtion and a refrain time point tcorus; s2, generating a light theme template based on the Cemtion and the Tbeat, and setting a switching time point; s3, separating the microphone signal X (t) into a human voice track V (t) and an accompaniment track A (t) through a pre-trained real-time source separation model; s4, adjusting spotlight parameters according to the volume VdB and pitch change rate # imgabs0 # of the human voice track V (t), and driving an atmosphere lamp according to the frequency spectrum of the accompaniment track A (t); s5, acquiring a personnel distribution thermodynamic diagram H (x, y) through an infrared thermal imaging sensor, and acquiring an action amplitude parameter Mlevel through a millimeter wave radar; and S6, fusing personnel distribution and motion amplitude data. According to the invention, KTV intelligent light control can be carried out more intelligently.
Owner:CHENGDU YINYUE CHUANGXIANG TECH CO LTD

Real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning

The invention discloses a real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning. The method comprises the following steps that S1, user data are collected and preprocessed; s2, extracting feature vectors of facial expressions and limb actions; s3, generating speech synthesis data by using the improved GE2E network and a preset target speech text; s4, generating a synthesized voice audio based on the voice synthesis data; s5, generating a three-dimensional digital human motion sequence according to the facial expression feature vector and the body motion feature vector; s6, performing timestamp alignment on the three-dimensional digital human action sequence and the synthesized voice audio, and constructing a synchronous output stream; and S7, rendering the synchronous output stream, and performing three-dimensional visual output. According to the method, the improved GE2E network, deep learning and sound cloning methods are fused, three-dimensional virtual human voice action synchronous control is achieved, and the method has the advantages of being high in real-time performance, high in immersion and natural in interaction.
Owner:HANGZHOU SECOND LIFE TECH CO LTD

Intelligent video editing method fusing human face features and human voice features

The invention relates to an intelligent video editing method fusing human face features and human voice features, which comprises the steps of input preprocessing, multi-modal analysis, fusion scoring and automatic editing, and adopts a multi-modal mode for analysis, so that the recall rate of key segments is improved, and the efficiency of editing is improved. A personalized editing strategy is supported, facial expression changes and voice emotion peak values are aligned through dynamic time warping, a CLIP-like structure is used for training a bimodal encoder, the human face and the voice are mapped to a unified vector space, the association weight of the human face and the voice is automatically learned, and the intelligent degree of editing is improved.
Owner:BEIJING HEJUHUITONG E-COMMERCE CO LTD

Karaoke all-in-one machine intelligent tuning method based on AI sound effect optimization

The invention relates to the technical field of artificial intelligence and digital audio signal processing, and particularly discloses a Karaoke all-in-one machine intelligent tuning method based on AI sound effect optimization, and the method comprises the steps: obtaining an original mixed audio signal sung by a user, and carrying out the preprocessing; processing the preprocessed audio by using a multi-source audio separation model based on a deep neural network, and extracting a human voice track and accompaniment components; extracting time-frequency domain features from the human voice track, and performing quantitative evaluation on the voice tension degree of the user in combination with the trained voice state recognition model; dynamically adjusting an equalizer and a dynamic range compression parameter according to an evaluation result to realize adaptive sound effect optimization; synthesizing the optimized voice and accompaniment and outputting the synthesized voice and accompaniment to monitoring equipment; and finally, establishing a personalized user behavior feature vector by collecting user singing performance and historical tuning data, and performing online fine tuning on a system strategy based on a deep Q network reinforcement learning mechanism to continuously optimize a tuning effect.
Owner:SHENZHEN MSAI TECH CO LTD

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Single-channel music separation method based on U-shaped network

The invention discloses a single-channel music separation method based on a U-shaped network, which is used for separating human voice from accompaniment. According to the method, a U-shaped network model comprising an encoder, a connection layer, a bottleneck layer and a decoder is constructed, feature extraction is performed on input through a plurality of continuous encoding units, and then the input enters the bottleneck layer formed by a multi-scale feature extraction module, a dual-path module and a multi-scale feature extraction module. And a connection layer formed by the time-frequency enhancement module and the fusion module is used for connecting the coding unit and the bottleneck layer or the coding unit and the decoding unit, and inputting into the next decoding unit. And finally, the decoder outputs the predicted human voice mask and accompaniment mask. A time-frequency enhancement module is introduced into a jump connection layer, and input music signal features on each frequency box are weighted in a time dimension, so that a network model pays attention to more important time points, and the time sequence modeling capability is improved. And meanwhile, a multi-scale feature extraction module is introduced into a bottleneck layer, so that the music signal feature extraction capability of the model is enhanced.
Owner:HANGZHOU DIANZI UNIV

Audio separation method and related equipment

The invention provides an audio separation method and related equipment. The audio separation method and the related equipment are suitable for a scene in which background music is separated from songs. A to-be-processed mixed audio signal is obtained, and the mixed audio signal is a unified audio signal formed by mixing a plurality of audio signals together through a sound mixer or a sound console. Wherein the plurality of audio signals includes a human voice signal and a background sound (e.g., music sound) signal. And inputting the mixed audio signal into a diffusion model, extracting a human voice signal in the mixed audio signal, and subtracting the human voice signal from the mixed audio signal to obtain a background sound signal in the mixed audio signal. According to the scheme, the generation type model, namely the diffusion model, which is more advanced compared with a deep neural network is adopted to learn distribution of the human voice signals, and compared with a traditional method, the human voice can be estimated more accurately, so that the human voice included in the mixed audio signals is extracted more accurately; and thus, high-quality background sound is separated from the mixed audio signal.
Owner:HONOR DEVICE CO LTD

Double-earphone sound effect adjusting method and system, storage medium and program product

The invention discloses a double-earphone sound effect adjusting method and system, a storage medium and a program product, and relates to the field of loudspeakers, and the method comprises the steps: collecting an environment audio signal through a binaural microphone, and executing spectrum analysis to obtain a frequency distribution characteristic of the environment audio signal; based on a preset human voice frequency range, extracting an audio signal conforming to the human voice feature from the frequency distribution feature as a target audio; calculating the time difference and the intensity difference of the target audio between the left ear and the right ear to obtain the human sound source direction and the relative distance in the environment; acquiring head attitude sensor data, and calculating a matching included angle between the head orientation of the user and the human sound source direction based on the head attitude sensor data; when the matching included angle is smaller than a preset threshold value, increasing the audio signal gain of the head facing the direction; and when the conversation voice of the user is detected, the audio signal gain is kept unchanged. By implementing the application, the interaction efficiency when the double earphones are used can be improved.
Owner:BESING TECH SHENZHEN CO LTD

Audio and video synchronization detection method and device, electronic equipment, program product and medium

The invention discloses an audio and video synchronous detection method and device, electronic equipment, a program product and a medium, and relates to the field of artificial intelligence. Human voice audio and human face video in a to-be-detected video file can be segmented into audio clips and video clips according to preset clip duration; and time sequence feature extraction can be performed on each audio clip to obtain the audio feature of each audio clip, and time sequence feature extraction can be performed on each video clip to obtain the video feature of each video clip. Then, the feature similarity between each audio clip and each video clip can be calculated, and when it is determined that the target feature similarity between the target audio clip and the target video clip at the same segmentation position is not larger than the feature similarity between the target audio clip and at least one other video clip, the target audio clip and the at least one other video clip can be segmented. And judging that the audio and video of the to-be-detected video file are asynchronous, thereby improving the accuracy of audio and video synchronous detection by means of time sequence feature extraction, feature similarity calculation and feature similarity comparison.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Artificial intelligence-powered music registry, collaboration, and workflow management system

An AI-powered music registry, collaboration, and workflow management platform that addresses the challenges faced by music industry participants in the digital age. The system comprises a segmentation and hashing subsystem for musical pieces, segments, and isolated elements, enabling the evaluation of uniqueness and the consideration of individual creators' contributions. An artificial intelligence (AI) and machine learning (ML) subsystem is employed for extracting and isolating individual instruments, vocals, and performer contributions, while a component-level tracking module enables enhanced crediting and royalty distribution.
Owner:QOMPLX INC

Electric cooker anti-noise voice interaction system based on multi-mode perception and control method

The invention relates to the technical field of intelligent household appliances and man-machine interaction, and discloses an electric cooker anti-noise voice interaction system based on multi-mode perception and a control method, and the system comprises a multi-mode perception and collection unit which synchronously collects millimeter wave radar echo signals and acoustic signals; the signal preprocessing and feature extraction unit is used for extracting a user physiological vibration signal and a three-dimensional space position vector from the radar signal and extracting an acoustic energy envelope and a sound source direction vector from the acoustic signal; the time-space consistency verification unit and the voice gating and recognition unit are used for carrying out time synchronization verification by calculating the correlation between the physiological vibration signal and the acoustic energy envelope, and distinguishing real human voice from an environment false trigger source; meanwhile, space consistency verification is carried out by comparing the user direction of radar positioning with the sound source direction of acoustic positioning, so that non-target human voice interference is eliminated. According to the invention, the anti-interference capability and reliability of voice interaction in a real home environment are improved through double verification on a physical level.
Owner:LINGNAN NORMAL UNIV

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Intelligent volume adjusting method and system based on environmental sound

The invention discloses an intelligent volume adjusting method and system based on environment sound, and the method is applied to a wireless earphone, and comprises the steps: determining a target prompt word according to an initial prompt word registration form submitted by a user; environment sound sampling data collected by wireless earphones are divided into environment noise data and human voice data, and the wireless earphones comprise a left earphone worn on a left ear and a right earphone worn on a right ear; dividing the human voice data into a plurality of groups of character voice data in one-to-one correspondence with the plurality of groups of characters according to voice features in the human voice data; determining target character sound data corresponding to the phrase whose similarity with the target prompt word is greater than a preset threshold value when identifying that the phrase whose similarity with the target prompt word is greater than the preset threshold value exists in the plurality of groups of character sound data; determining the orientation of human voice data corresponding to the target human voice data; and reducing the volume of the left earphone or the right earphone according to the orientation. According to the invention, the use convenience of the wireless earphone is improved.
Owner:BESING TECH SHENZHEN CO LTD

Audio recognition and noise reduction processing method and system for power grid regulation and control workbench environment

The invention discloses an audio recognition and noise reduction processing method and system for a power grid regulation and control workbench environment, and the method comprises the steps: obtaining pure audio signals of all power grid regulation and control workers in different working scenes and different background noises in a dispatching desk, carrying out the cross-domain fusion of two features of a time domain and a frequency domain, and taking the two features as a training set to train a speaker model; voice signals are collected in real time, and pure noise and silence in the voice signals are removed; according to the noise level estimation and the noise power, carrying out weighting processing on the voice signal after the pure noise and the mute are removed; recognizing and separating human voice fragments by adopting a continuous down-sampling and resampling algorithm with multi-resolution characteristics; judging whether the human voice fragment contains a target speaker or not through a speaker model; and generating a plurality of voice segments based on the human voice segments, and extracting target speaking human voice through clustering and a speaker model. The method adapts to environmental changes of regulation and control dispatching desks at all levels, environmental noise can be remarkably reduced, and accurate recognition and extraction of human voices are achieved.
Owner:STATE GRID FUJIAN ELECTRIC POWER CO LTD +3

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Rehabilitation training system for hearing disorder patient based on virtual reality

The invention discloses an auditory disorder patient rehabilitation training system based on virtual reality. The system comprises a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module and a virtual reality auxiliary module. The invention belongs to the technical field of auditory rehabilitation training, and particularly relates to an auditory disorder patient rehabilitation training system based on virtual reality, which adopts a multi-stage cross-modal attention-enhanced voice audiovisual characterization method, converts sound signals into visual light waves and tactile vibration data through a deep learning model, and displays the visual light waves and tactile vibration data on the basis of the visual light waves and tactile vibration data. High-precision and real-time multi-modal data fusion is realized; basic auditory reconstruction is carried out by adopting an audio frequency generation dynamic adjustment method combined with voiceprint dynamic mapping, personalized sound scene mapping is effectively and dynamically generated, and a patient is helped to reconstruct basic perception ability for sound; a cross-modal contrast learning method is adopted to assist lip language synchronous training, and voice context background prompts are provided in the rehabilitation training process.
Owner:川北医学院附属医院

Noise reduction regulation and control method for earphone and noise reduction earphone

The invention belongs to the technical field of earphone noise reduction, and provides a noise reduction regulation and control method for an earphone and a noise reduction earphone. The method comprises the following steps: acquiring a first environment sound signal of an environment area, and extracting human voice features from the first environment sound signal to construct a human voice feature template library; in response to a noise reduction regulation and control instruction, collecting a second environment sound signal of the environment area, converting the time domain signal into a frequency domain signal through Fourier transform, and extracting a frequency spectrum feature from the frequency domain signal; performing matching calculation on the spectrum features and a human voice feature template library to determine environment voice signals, human voice signals and noise signals; an active noise reduction algorithm is adopted to generate offset sound waves opposite to the noise signals in phase, the human voice signals are amplified, and the offset sound waves and the amplified human voice signals are output through a loudspeaker of the earphone. According to the invention, effective noise filtering and clear transparent transmission of specific human voice are realized, smooth communication can be realized without taking off the earphone, and the use experience is improved.
Owner:GUANGZHOU SOUNDBOX ACOUSTIC TECH

Human voice activity detection method and device, computer equipment and storage medium

The invention relates to the field of audio processing, and discloses a human voice activity detection method and device, computer equipment and a storage medium. The method comprises the following steps: collecting audio and video data of a user, and dividing the audio and video data into a video stream and an audio stream; detecting the video stream by using a FaceMesh model to obtain a lip opening degree and a head attitude angle, and determining a lip motion state according to the lip opening degree and the head attitude angle; blocking the audio stream to obtain a plurality of audio blocks, performing noise reduction on each audio block to obtain a plurality of noise-reduced audio blocks, and detecting each noise-reduced audio block by using a silhoVAD model to obtain a human voice activity probability and an audio cache queue; and inputting the lip motion state and the human voice activity probability into a state machine, introducing an audio cache queue by the state machine, and outputting a speaking identifier and a corresponding audio clip. Speaking recognition is cooperatively completed in combination with machine vision and hearing signals, and the accuracy of whole human voice activity detection is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Audible voice interaction method and system

The invention belongs to the technical field of robot and human voice interaction, and particularly relates to a human voice interaction method and system. The method comprises the following steps: step 1, an intelligent agent enables radio duration to realize self-adaption according to environment loudness; 2, designing a multi-agent framework with a large model and a small model cooperating with each other; and step 3, based on the framework designed in the step 2, analyzing the radio in the step 1, and using a retrieval enhancement technology based on a tree-shaped document to improve the question and answer reliability of the intelligent agent and realize the voice interaction of the intelligent agent. The method is used for solving the problems that in the prior art, most voice interaction technologies can only achieve single-round interaction and cannot remember context content of interaction, and pure human voice is difficult to obtain when the traditional voice interaction technology is directly transplanted to a robot which can move and makes noise by itself.
Owner:HARBIN INST OF TECH +1

Speech enhancement method and device based on full-band information and sub-band information

The invention provides a speech enhancement method and device based on full-band information and sub-band information, and belongs to the technical field of speech processing.The speech enhancement method comprises the steps that windowing Fourier transform is conducted on an audio signal; extracting full-band features and sub-band features from the frequency domain representation result of the audio signal; performing full-band mask prediction on the full-band features to obtain a 256-dimensional full-band mask value; splicing the full band mask values of the 256 dimensions and the Log Spectrogram features of the 256 dimensions, and dividing the spliced features into 16 groups of 32-dimensional sub-band features according to a frequency band; performing sub-band mask prediction on the 16 groups of 32-dimensional sub-band features to obtain 16 groups of 16-dimensional sub-band mask values; according to the 16 groups of 16-dimensional sub-band mask values, obtaining a 256-dimensional target mask value; and obtaining an enhanced audio signal according to the 256-dimensional target mask value. According to the voice enhancement method provided by the invention, the damage to human voice can be reduced while the effective noise reduction of the audio can be realized, and the voice enhancement effect can be effectively improved.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

earphone

1. The name of the product of this design: earphones. 2. Purpose of this design product: It has noise reduction and intelligent voice transmission functions and can be used as a hearing protector. For example, it can be used as an intelligent labor protection headset in a noisy production environment. 3. The key point of the design of this product lies in its shape. 4. The picture or photo that best illustrates the key points of the design: Design 1 Stereoscopic Figure 1. 5. Designate Design 1 as the base design.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

A Human-Computer Voice Interaction Control Method and System Based on Smart TV

This application relates to a human-computer voice interaction control method and system based on a smart TV. The method includes acquiring voice data within a preset range, processing the voice data to obtain voice feature data carrying control commands; performing feature analysis on the voice feature data using a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results; sending a secondary confirmation request to the user based on the control command comparison results, and combining the confirmation voice information from the user feedback to perform command recognition evaluation and correct command deviation processing to obtain the corrected control command; performing function switching processing on the TV according to the correct control command, and optimizing the display effect by linking and adjusting related devices based on the program display requirements after the switch to obtain human-computer voice interaction control data. This application has the effect of improving the intelligence of voice interaction control of smart TVs.
Owner:GUANGZHOU XIANYOU INTELLIGENT TECH CO LTD

Cross-Language Voice Similarity Analysis

A system includes a hardware processor and a memory storing a cross-language voice similarity analyzer (analyzer). The hardware processor executes the analyzer to generate an embedding vector representation of an audio sample of a human voice in a feature space including existing embedding vectors corresponding respectively to different reference voices, decompose the embedding vector representation to identify a linear or non-linear combination of vocal component vectors corresponding to the human voice, each vocal component vector representing a respective predetermined voice characteristic descriptor, increase the dimensionality of the linear or non-linear combination of the vocal component vectors to match the dimensionality of the embedding vector representation to provide a reconstructed embedding vector representation of the human voice, and identify, by comparing the reconstructed embedding vector representation with one or more of the existing embedding vectors, one of the reference voices as a match for the audio sample of the human voice.
Owner:DISNEY ENTERPRISES INC

Cabin voice recognition method and device based on ASR and VAD and medium

The invention provides a cabin voice recognition method and device based on ASR and VAD, and a medium, and relates to the technical field of voice recognition, and the method comprises the steps: obtaining a to-be-recognized voice segment of a preset time length; dividing the voice segment to be recognized into a plurality of continuous voice frames with the same duration to obtain a voice frame list SA; inputting the SA into a preset VAD model to obtain a voice confidence list SA 'corresponding to the SA; performing a first sliding window operation on SA '; obtaining the number QN of human voice confidence coefficients greater than a preset first human voice confidence coefficient threshold value in the first sliding window; if QN / m is greater than eta 1, determining that the voice frame corresponding to the first personal voice confidence in the corresponding first sliding window is a human voice frame; determining a human voice segment corresponding to the to-be-recognized voice according to the human voice frame and the non-human voice frame in the SA; on the basis of improving the accuracy and the stability of human voice detection, the method has relatively high expandability and applicability.
Owner:MOBILE TECH COMPANY CHINA TRAVELSKY HLDG +1

Distributed discernment system

An example of a distributed discernment system including a discernment server and a communications interface permitting bi-directional communications to and from the discernment server; and a plurality of human interface devices, each including a speaker, a microphone, a processor running a local processing program, and a system interface permitting bi-directional communications between the human interface device and the discernment server, where the diagnostic program running on the discernment server is adapted to generate interview instructions provided to the interface devices and the interface devices are adapted receive interview instructions from the discernment server, present a verbal question to a human interviewee; receive and process sensor data from the microphone to determine whether the microphone sensor data corresponds to a complete human voice response to the presented verbal question.
Owner:DISCERN SCI INT INC