Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1044 results about "Vocal sound" patented technology

The human voice consists of sound made by a human being using the vocal tract, such as talking, singing, laughing, crying, screaming, etc. The human voice frequency is specifically a part of human sound production in which the vocal folds (vocal cords) are the primary sound source.

Voice interaction method and system of AI intelligent robot

The invention relates to the technical field of voice interaction, particularly discloses an AI intelligent robot voice interaction method and system, and aims to solve the problems of low voice interaction accuracy, insufficient reliability and lack of authority control in a complex noise environment. A dynamic noise feature library containing steady-state noise, impact noise and human voice interference features and a pre-stored gesture instruction library are constructed, audio signals are collected in real time, low-frequency-band, middle-frequency-band and high-frequency-band differential noise reduction is executed, Mel-frequency cepstral coefficient features are extracted, noise scenes are matched, corresponding voice recognition models are switched, and voice recognition is achieved. And calculating a confidence value of the voice instruction, outputting multi-modal verification data in combination with a dynamic confidence threshold, and outputting an authority control signal through voiceprint matching, authority verification and instruction consistency judgment. Through multi-modal fusion, dynamic adaptation and authority control, the voice recognition accuracy and interaction safety in a complex noise environment are remarkably improved, and the method is suitable for scenes such as factory intelligent inspection.
Owner:HANGZHOU SOHA TECH CO LTD

Voice processing method and device based on voiceprint feature screening, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health, voice navigation and the like, and discloses a voice processing method, device, equipment and medium based on voiceprint feature screening. The method comprises the following steps: screening voice of a near-field main speaker based on voiceprint similarity and signal intensity, executing duration filtering and confidence verification, dynamically constructing a voiceprint feature library, generating a voiceprint mask matrix based on the voiceprint feature library, performing frequency band suppression on a to-be-processed voice signal, and outputting a purified voice signal. According to the method, the voiceprint library is dynamically constructed, the voiceprint mask matrix is generated based on the voiceprint feature library, frequency band suppression is performed on the to-be-processed voice signal, non-target voiceprints are effectively shielded, the voice signal quality is improved, and thus target voice is accurately captured in a high-noise environment.
Owner:平安科技(上海)有限公司

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Karaoke all-in-one machine intelligent tuning method based on AI sound effect optimization

The invention relates to the technical field of artificial intelligence and digital audio signal processing, and particularly discloses a Karaoke all-in-one machine intelligent tuning method based on AI sound effect optimization, and the method comprises the steps: obtaining an original mixed audio signal sung by a user, and carrying out the preprocessing; processing the preprocessed audio by using a multi-source audio separation model based on a deep neural network, and extracting a human voice track and accompaniment components; extracting time-frequency domain features from the human voice track, and performing quantitative evaluation on the voice tension degree of the user in combination with the trained voice state recognition model; dynamically adjusting an equalizer and a dynamic range compression parameter according to an evaluation result to realize adaptive sound effect optimization; synthesizing the optimized voice and accompaniment and outputting the synthesized voice and accompaniment to monitoring equipment; and finally, establishing a personalized user behavior feature vector by collecting user singing performance and historical tuning data, and performing online fine tuning on a system strategy based on a deep Q network reinforcement learning mechanism to continuously optimize a tuning effect.
Owner:SHENZHEN MSAI TECH CO LTD

Hearing aid intelligent noise reduction and human voice enhancement technology based on electroencephalogram signals

The invention relates to a hearing aid intelligent noise reduction and human voice enhancement system based on electroencephalogram signals, and belongs to the field of biomedical engineering and acoustic signal processing. The system comprises an electroencephalogram signal acquisition module, a multi-channel acoustic sensor array, an embedded neural signal processor, an adaptive beam forming module, a dynamic speech enhancement engine and a dual-mode output device, and constructs electroencephalogram-acoustics joint features by extracting an alpha / theta wave power ratio, a P300 component and auditory cortical Gamma phase synchronism. A deep network is driven to separate target voice, a wave beam direction and a frequency response curve are dynamically adjusted based on neural feedback, a closed-loop calibration unit is innovatively adopted, gain is reversely adjusted according to N1-P2 wave amplitude, heart rate variability and eye movement data are fused to optimize decisions, and when the signal-to-noise ratio is-5dB, the voice recognition rate reaches 89%, the auditory fatigue is reduced by 37%, and the decision conflict rate is smaller than 6%. The defects of attention blind area, noise separation failure and physiological adaptation of a traditional hearing aid are overcome. The system is suitable for the fields of hearing impairment rehabilitation, special communication and intelligent cabins.
Owner:MAXSON GLOBAL GROUP INC

Single-channel music separation method based on U-shaped network

The invention discloses a single-channel music separation method based on a U-shaped network, which is used for separating human voice from accompaniment. According to the method, a U-shaped network model comprising an encoder, a connection layer, a bottleneck layer and a decoder is constructed, feature extraction is performed on input through a plurality of continuous encoding units, and then the input enters the bottleneck layer formed by a multi-scale feature extraction module, a dual-path module and a multi-scale feature extraction module. And a connection layer formed by the time-frequency enhancement module and the fusion module is used for connecting the coding unit and the bottleneck layer or the coding unit and the decoding unit, and inputting into the next decoding unit. And finally, the decoder outputs the predicted human voice mask and accompaniment mask. A time-frequency enhancement module is introduced into a jump connection layer, and input music signal features on each frequency box are weighted in a time dimension, so that a network model pays attention to more important time points, and the time sequence modeling capability is improved. And meanwhile, a multi-scale feature extraction module is introduced into a bottleneck layer, so that the music signal feature extraction capability of the model is enhanced.
Owner:HANGZHOU DIANZI UNIV

Artificial intelligence-powered music registry, collaboration, and workflow management system

An AI-powered music registry, collaboration, and workflow management platform that addresses the challenges faced by music industry participants in the digital age. The system comprises a segmentation and hashing subsystem for musical pieces, segments, and isolated elements, enabling the evaluation of uniqueness and the consideration of individual creators' contributions. An artificial intelligence (AI) and machine learning (ML) subsystem is employed for extracting and isolating individual instruments, vocals, and performer contributions, while a component-level tracking module enables enhanced crediting and royalty distribution.
Owner:QOMPLX INC

Electric cooker anti-noise voice interaction system based on multi-mode perception and control method

The invention relates to the technical field of intelligent household appliances and man-machine interaction, and discloses an electric cooker anti-noise voice interaction system based on multi-mode perception and a control method, and the system comprises a multi-mode perception and collection unit which synchronously collects millimeter wave radar echo signals and acoustic signals; the signal preprocessing and feature extraction unit is used for extracting a user physiological vibration signal and a three-dimensional space position vector from the radar signal and extracting an acoustic energy envelope and a sound source direction vector from the acoustic signal; the time-space consistency verification unit and the voice gating and recognition unit are used for carrying out time synchronization verification by calculating the correlation between the physiological vibration signal and the acoustic energy envelope, and distinguishing real human voice from an environment false trigger source; meanwhile, space consistency verification is carried out by comparing the user direction of radar positioning with the sound source direction of acoustic positioning, so that non-target human voice interference is eliminated. According to the invention, the anti-interference capability and reliability of voice interaction in a real home environment are improved through double verification on a physical level.
Owner:LINGNAN NORMAL UNIV

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Video dubbing language conversion method and system and related equipment

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Owner:SHENZHEN MAIFENG TECH CO LTD

Noise reduction regulation and control method for earphone and noise reduction earphone

The invention belongs to the technical field of earphone noise reduction, and provides a noise reduction regulation and control method for an earphone and a noise reduction earphone. The method comprises the following steps: acquiring a first environment sound signal of an environment area, and extracting human voice features from the first environment sound signal to construct a human voice feature template library; in response to a noise reduction regulation and control instruction, collecting a second environment sound signal of the environment area, converting the time domain signal into a frequency domain signal through Fourier transform, and extracting a frequency spectrum feature from the frequency domain signal; performing matching calculation on the spectrum features and a human voice feature template library to determine environment voice signals, human voice signals and noise signals; an active noise reduction algorithm is adopted to generate offset sound waves opposite to the noise signals in phase, the human voice signals are amplified, and the offset sound waves and the amplified human voice signals are output through a loudspeaker of the earphone. According to the invention, effective noise filtering and clear transparent transmission of specific human voice are realized, smooth communication can be realized without taking off the earphone, and the use experience is improved.
Owner:GUANGZHOU SOUNDBOX ACOUSTIC TECH

Human voice activity detection method and device, computer equipment and storage medium

The invention relates to the field of audio processing, and discloses a human voice activity detection method and device, computer equipment and a storage medium. The method comprises the following steps: collecting audio and video data of a user, and dividing the audio and video data into a video stream and an audio stream; detecting the video stream by using a FaceMesh model to obtain a lip opening degree and a head attitude angle, and determining a lip motion state according to the lip opening degree and the head attitude angle; blocking the audio stream to obtain a plurality of audio blocks, performing noise reduction on each audio block to obtain a plurality of noise-reduced audio blocks, and detecting each noise-reduced audio block by using a silhoVAD model to obtain a human voice activity probability and an audio cache queue; and inputting the lip motion state and the human voice activity probability into a state machine, introducing an audio cache queue by the state machine, and outputting a speaking identifier and a corresponding audio clip. Speaking recognition is cooperatively completed in combination with machine vision and hearing signals, and the accuracy of whole human voice activity detection is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Audible voice interaction method and system

The invention belongs to the technical field of robot and human voice interaction, and particularly relates to a human voice interaction method and system. The method comprises the following steps: step 1, an intelligent agent enables radio duration to realize self-adaption according to environment loudness; 2, designing a multi-agent framework with a large model and a small model cooperating with each other; and step 3, based on the framework designed in the step 2, analyzing the radio in the step 1, and using a retrieval enhancement technology based on a tree-shaped document to improve the question and answer reliability of the intelligent agent and realize the voice interaction of the intelligent agent. The method is used for solving the problems that in the prior art, most voice interaction technologies can only achieve single-round interaction and cannot remember context content of interaction, and pure human voice is difficult to obtain when the traditional voice interaction technology is directly transplanted to a robot which can move and makes noise by itself.
Owner:HARBIN INST OF TECH +1

earphone

1. The name of the product of this design: earphones. 2. Purpose of this design product: It has noise reduction and intelligent voice transmission functions and can be used as a hearing protector. For example, it can be used as an intelligent labor protection headset in a noisy production environment. 3. The key point of the design of this product lies in its shape. 4. The picture or photo that best illustrates the key points of the design: Design 1 Stereoscopic Figure 1. 5. Designate Design 1 as the base design.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

A Human-Computer Voice Interaction Control Method and System Based on Smart TV

This application relates to a human-computer voice interaction control method and system based on a smart TV. The method includes acquiring voice data within a preset range, processing the voice data to obtain voice feature data carrying control commands; performing feature analysis on the voice feature data using a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results; sending a secondary confirmation request to the user based on the control command comparison results, and combining the confirmation voice information from the user feedback to perform command recognition evaluation and correct command deviation processing to obtain the corrected control command; performing function switching processing on the TV according to the correct control command, and optimizing the display effect by linking and adjusting related devices based on the program display requirements after the switch to obtain human-computer voice interaction control data. This application has the effect of improving the intelligence of voice interaction control of smart TVs.
Owner:GUANGZHOU XIANYOU INTELLIGENT TECH CO LTD

Cross-Language Voice Similarity Analysis

A system includes a hardware processor and a memory storing a cross-language voice similarity analyzer (analyzer). The hardware processor executes the analyzer to generate an embedding vector representation of an audio sample of a human voice in a feature space including existing embedding vectors corresponding respectively to different reference voices, decompose the embedding vector representation to identify a linear or non-linear combination of vocal component vectors corresponding to the human voice, each vocal component vector representing a respective predetermined voice characteristic descriptor, increase the dimensionality of the linear or non-linear combination of the vocal component vectors to match the dimensionality of the embedding vector representation to provide a reconstructed embedding vector representation of the human voice, and identify, by comparing the reconstructed embedding vector representation with one or more of the existing embedding vectors, one of the reference voices as a match for the audio sample of the human voice.
Owner:DISNEY ENTERPRISES INC

Multi-mode AI-driven short video automatic translation and speech synthesis system

The invention discloses a multi-mode AI-driven short video automatic translation and speech synthesis system, which relates to the technical field of audio and video and comprises video acquisition, processing, translation and synthesis modules and the like. The video acquisition end receives various audio and video data submitted by a user and detects violation content; the primary processing end identifies formats of compliance data, transcodes the compliance data, separates audio and video streams, and further separates human voices and background voices; the video translation end extracts a human voice signal to generate a target language text, generates an audio stream in combination with rhythm features and a human voice module, and fits the audio stream with background sound; a video synthesis end extracts a lip image frame sequence from a video frame, generates a matched lip vertex displacement coordinate according to an audio frame sequence and renders the matched lip vertex displacement coordinate, and finally compresses, packages and outputs an audio and video synchronization sequence; various AI technologies are applied to the modules, automatic translation and speech synthesis of short videos are achieved, and the generation efficiency and quality of cross-language video content are improved.
Owner:南京地平线网络科技有限公司

Game machine

To improve a game machine capable of outputting music pieces.SOLUTION: A game machine can execute multiple kinds of notice performances, can output an instrumental music piece during a period before development of a weak SP ready-to-win, can output a vocal music piece during a period after development of a strong SP ready-to-win, can execute a music piece limitation performance for limiting output of an instrumental music piece at the time of execution from among notice performances during a period before development of a weak SP ready-to-win, and can execute a music piece limitation performance for limiting output of a vocal music piece at the time of execution from among notice performances during a period after development of a strong SP ready-to-win.SELECTED DRAWING: Figure 14-68
Owner:SANKYO CO LTD

Signal processing method and device, model training method and device, equipment and storage medium

The embodiment of the invention provides a signal processing method and device, a model training method and device, equipment and a storage medium. The method comprises the steps of obtaining a mixed sound signal and a reference sound signal; filtering a linear echo component in the mixed sound signal based on the reference sound signal to obtain a first filtered signal; obtaining an output signal through a pre-trained nonlinear filtering model and the first filtering signal, and sending the output signal to the far-end terminal equipment; wherein the output signal filters out a nonlinear echo component of a target energy quantity level relative to the first filtering signal, the target energy quantity level is smaller than a preset energy level, and the preset energy level is determined based on the energy level of the human voice component in the first filtering signal. The nonlinear component in the sample filtering signal is filtered based on the nonlinear filtering model, and meanwhile, the filtering amount of the nonlinear component is controlled, so that excessive suppression of human voice components is avoided, and the voice communication quality in a complex environment is improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Speech enhancement method and system in dynamic noise environment

The invention relates to the technical field of speech analysis, in particular to a speech enhancement method and system in a dynamic noise environment. The method comprises the following steps: acquiring a voice signal within a preset time under a fixed voice frequency band, and determining a main channel and a secondary channel; framing and windowing the voice signal to obtain a voice field, obtaining a time-frequency matrix of the voice field, determining a feature vector according to a phase difference and an amplitude ratio of each time-frequency point, and obtaining a key human voice cluster based on feature vector clustering; in the key human voice cluster, direction consistency of time-frequency points is calculated based on complex numbers of the main channel and the secondary channel, each time-frequency point is compared with other time-frequency points to obtain voiceprint feature consistency, and a human voice evaluation index is formed based on the time-frequency points and the other time-frequency points; obtaining a frequency domain binary label matrix; and training a network based on the frequency domain binary label matrix, and completing voice enhancement through the network. According to the invention, the problems of many interference persons and insufficient noise interference separation capability are solved.
Owner:SUZHOU AUDITORYWORKS CO LTD

Intelligent chest card collection and analysis method and system based on AI

The invention relates to the technical field of AI large model analysis, and discloses an AI-based intelligent chest card collection and analysis method and system, and the method comprises the steps: achieving the communication between an equipment management service module and chest card equipment through an emqx service module, and carrying out the information configuration of the chest card equipment through the equipment management service module; sound information is collected through chest card equipment; and uploading the sound information to a quality inspection analysis platform, processing the sound information through the quality inspection analysis platform to obtain an audio transliteration text, and analyzing the audio transliteration text based on an AI large model to obtain a service judgment result. Compared with a traditional keyword matching mode, the AI analysis is carried out in a mode of transferring the text through the audio, the ability of understanding semantics and hidden information can be improved, and the accuracy of judging the service quality can also be improved. The intelligent chest card recognizes personnel position information and personnel face information through the structured light module, and the definition of human voice collection can be greatly improved.
Owner:SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD

Multi-voice separation method based on lightweight dual-path Transform network

The invention discloses a multi-voice separation method based on a lightweight dual-path Transform network, and the method comprises the steps: collecting audio multi-voice data, and carrying out the preprocessing of the data, and forming a data set; the method comprises the following steps: constructing a dual-path Transform network model DPTNet, and introducing a recurrent neural network to optimize the dual-path Transform network model DPTNet; and training the dual-path network model DPTNet, and performing engineering deployment based on the trained model. The method is beneficial to obtaining higher-quality audio fingerprint recognition capability, sound source separation capability and voice enhancement function, can be used for tracking and positioning the position of a sound source, helps positioning and tracking related applications, can be expanded to the medical field, can be used for heart sound segmentation, namely, recognition of specific signals of the heart, and can be applied to the field of medical science. The method helps to diagnose cardiovascular and other medical problems, and has technical innovation and practical application value.
Owner:NANTONG UNIV

Intelligent quality inspection system for high-noise environment language service quality

The invention provides a high-noise environment language service quality intelligent quality inspection system. The intelligent quality inspection system for the language service quality in the high-noise environment comprises a noise sensing module for extracting noise features by performing short-time Fourier transform and spectrogram analysis on an input voice signal; the self-adaptive speech enhancement module is used for dynamically combining the subunits based on the noise characteristics so as to enhance human voice signals; and the noise conditional automatic voice recognition module is used for establishing an acoustic model based on the noise features and the enhanced human voice signals. The intelligent quality inspection system for the language service quality in the high-noise environment not only improves the reliability of recognition, but also can reduce recognition errors caused by noise, lays a solid foundation for multi-dimensional quality inspection, has the advantages of real-time performance, multiple dimensions, automation and self-adaption, and is suitable for popularization and application. And the monitoring efficiency and accuracy of the voice service quality in the high-noise environment are remarkably improved.
Owner:ZHONGYI CLOUD (BEIJING) INTERNET OF THINGS TECH CO LTD

Multi-mode video subtitle identification method and system, electronic equipment and storage medium

The invention provides a multi-mode video subtitle recognition method and system, electronic equipment and a storage medium, and relates to the technical field of video processing, and the method comprises the steps: carrying out the audio and video track separation of a to-be-recognized video, and obtaining an audio file and a video file; carrying out human voice track and background sound track separation on the audio file to obtain a human voice track audio; performing subtitle recognition on the human voice track audio by adopting an automatic voice recognition method with a timestamp to obtain a first subtitle text; performing subtitle area detection on the video file according to the visual language model to obtain a subtitle area external frame; performing frame-by-frame subtitle recognition on the video file by adopting an optical character recognition method according to the external frame of the subtitle area to obtain a second subtitle text; and performing subtitle fusion on the first subtitle text and the second subtitle text according to the time axis to obtain a subtitle recognition result. According to the invention, the completeness and accuracy of subtitle recognition are improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Signal processing method, model training method, apparatus, device, and storage medium

PCT designated stageWO2025200881A1Speech analysisNonlinear filterVoice communication
A signal processing method, a model training method, an apparatus, a device, and a storage medium. The signal processing method comprises: acquiring a mixed sound signal and a reference sound signal (S101); on the basis of the reference sound signal, filtering a linear echo component in the mixed sound signal to obtain a first filtered signal (S102); and obtaining an output signal by means of a pre-trained nonlinear filtering model and the first filtered signal, and sending the output signal to a remote terminal device, wherein for the output signal, a nonlinear echo component of a target energy level is filtered out relative to the first filtered signal, the target energy level is less than a preset energy level, and the preset energy level is determined on the basis of the energy level of a human voice component in the first filtered signal (S103).A nonlinear component in a sample filtered signal is filtered out on the basis of a nonlinear filtering model, and the amount of the nonlinear component filtered out is controlled, avoiding over-suppression of the human voice component, and improving the voice communication quality in complex environments.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Audio noise reduction method, system and device based on AI model and storage medium

The invention discloses an audio noise reduction method, system and device based on an AI model and a storage medium, and relates to the technical field of voice noise reduction, and the method comprises the steps: receiving a live video stream, and segmenting the live video stream into image frame data and audio stream data; the image frame data and the audio stream data are input to a preset scene recognition model, the recognition result of the current scene is obtained, and the preset scene recognition model comprises a character recognition sub-model, a voice recognition sub-model and a classification output module; and determining and executing a current noise reduction strategy according to the identification result to obtain an audio output signal, the current noise reduction strategy being a human voice enhancement strategy, an ambient sound enhancement strategy or a fusion enhancement strategy. According to the invention, scene recognition is carried out in a multi-modal fusion mode, different noise reduction strategies are adaptively executed according to the scene recognition result, noise optimization can be carried out for different scenes, and the audio output effect is improved.
Owner:NANJING PUTIAN TELEGE INTELLIGENT BUILDING

Voice wake-up interaction method and system based on microphone array

The invention discloses a voice wake-up interaction method and system based on a microphone array, and the method comprises the steps: carrying out VAD processing, so as to judge whether a target audio segment has voice or not; voice and noise source directions are obtained; selecting beam parameters of voice and noise in combination with a pre-designed fixed beam; whether GSC module processing is carried out or not is selected according to the difference between the beam parameters of the voice and the noise, so that an enhanced audio signal is obtained, a wake-up task is carried out, and a final voice wake-up result is obtained; and according to whether the wake-up is successful, determining whether to lock the voice beam direction in the current interaction stage for enhancement, thereby preventing interference of voice in other directions on subsequent interaction tasks. According to the voice enhancement mode based on the microphone array, sound source positioning can be realized, interference in a non-target direction can be suppressed, the voice quality in the target direction can be improved, the wake-up success rate can be effectively improved when the voice enhancement mode is applied to voice wake-up, and then the experience of back-end voice interaction is improved.
Owner:PANOVASIC TECHNOLOGY CO LTD

Game machine

To improve a game machine capable of outputting music pieces.SOLUTION: A game machine can execute multiple kinds of notice performances, can output an instrumental music piece during a period before development of a weak SP ready-to-win, can output a vocal music piece during a period after development of a strong SP ready-to-win, can execute a music piece limitation performance for limiting output of an instrumental music piece at the time of execution from among notice performances during a period before development of a weak SP ready-to-win, and can execute a music piece limitation performance for limiting output of a vocal music piece at the time of execution from among notice performances during a period after development of a strong SP ready-to-win.SELECTED DRAWING: Figure 14-68
Owner:SANKYO CO LTD

Audio processing method, device and system

The invention discloses an audio processing method, device and system, which are applied to the technical field of audio processing and vehicles. The method comprises the following steps: acquiring a singer voice audio of a singer singing for a target song, and acquiring an original singer voice audio and an accompaniment audio of the target song; a first audio is controlled to be played in the first sound area, a second audio is controlled to be played in the second sound area, the first audio comprises an original singer sound audio, the second audio comprises a singer sound audio and an accompaniment audio, the first sound area is a space area where a singer is located, and the second sound area comprises a space area where the singer and a listener are located. Therefore, the singer can hear the original singing voice in the karaoke scene, but the listener is not interfered by the original singing voice, and the experience of the listener is improved while the better karaoke experience of the singer is ensured.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Microphone far-field pickup amplitude dynamic range control method

The invention relates to the technical field of audio signal processing, and discloses a microphone far-field pickup amplitude dynamic range control method, and the method comprises the steps: processing synchronous bone conduction and air conduction signals; respectively extracting feature vectors representing near-field interference and far-field human voice from the bone conduction signal and the air conduction signal; inputting the feature vector into an artificial intelligence decision engine, and carrying out collaborative analysis to generate a gain strategy considering interference suppression and human voice maintenance; and the gain of the air conduction signal is adjusted according to the strategy. The system comprises a signal preprocessing module, a first feature extraction module, a second feature extraction module, an artificial intelligence decision engine and a gain adjustment and reconstruction module. Unambiguous observation is carried out on near-field interference by using the bone conduction signal, the gain control problem when far-field weak human voice and near-field strong interference coexist is solved, and the pickup quality can be ensured while signal clipping can be effectively prevented.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD +1