Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Voice source" patented technology

The voice source contains important lexical and non-lexical information. The non-lexical information can convey, for example, prosodic events, emotional status, as well as cues pertaining to the uniqueness of the speaker’s voice. In engineering applications, there is a need for a more accurate source model that could model different voice qualities.

Robot motion control method and system based on voice interaction

InactiveCN120461427AProgramme-controlled manipulatorSimulationVoice source
The invention discloses a robot motion control method and system based on voice interaction, and relates to the technical field of robot motion control, and the method comprises the steps: extracting each segment of independent audio signal and keywords conforming to a preset voice library from mixed multi-channel audio based on a blind source separation technology; and selecting the keyword with the highest confidence coefficient as an execution command word, obtaining a standard track of an action corresponding to the execution command word, generating an executable simulation track of the execution command word in combination with real-time environment information, and controlling the robot to execute motion according to the simulation track after judging that the simulation track is safe and feasible. According to the method, the multi-channel mixed audio is separated into the independent signals, the execution command word is determined after the voice sources of different users are distinguished, the motion trail generated by simulating the conditions of the corresponding preset action and the enforceable environment is utilized, and the multi-dimensional characteristics of the motion trail and the standard trail of the preset action are compared and analyzed; and evaluating the feasibility of the motion trail corresponding to the execution command in the environment.
Owner:PUYANG VOCATIONAL & TECHN COLLEGE

Voice masking method, device and system and vehicle

A voice masking method, device and system are suitable for the field of intelligent vehicles, and the method comprises the following steps: determining a voice source position and a target masking position; receiving a sound signal from the voice source position, and detecting whether a voice signal exists in the sound signal so as to generate a detection result; if the detection result is that a voice signal exists in the sound signal, generating masking sound for the voice signal, and outputting the masking sound to a first loudspeaker at the voice source position and a second loudspeaker at the target masking position; and if the detection result is that the voice signal does not exist in the sound signal, no masking sound is generated. According to the invention, the requirement of passengers for private voice communication in the vehicle cabin space can be met, and the information safety is improved.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Conference control method and device based on AI vision and multi-device cooperation, equipment and storage medium

The invention discloses a conference control method, device and equipment based on AI vision and multi-equipment cooperation and a storage medium, and relates to the technical field of computer vision, and the conference control method based on AI vision and multi-equipment cooperation comprises the steps: recognizing a conference room environment through a convolutional neural network, and determining the position information of conference participants; sound source positioning is carried out on a voice source of a microphone array according to the position information, and the position of a spokesman is determined; controlling a camera to face the position of the spokesman so as to enable a display to display the spokesman; and under the condition that the display displays the spokesman, recognizing the speaking content of the spokesman through voice recognition and semantic analysis, and adjusting the content of the projector according to the speaking content. The conference efficiency is improved, the sense of participation and interactivity of conference participants are enhanced, intelligent cooperation and automatic control of multiple devices in the conference room are achieved, and powerful support is provided for smooth proceeding of the conference.
Owner:广州潮创科技有限公司

Facility agriculture voice interaction system based on large model

The invention discloses a facility agriculture voice interaction system based on a large model, and relates to the technical field of voice interaction. Comprising a voice signal acquisition and enhancement module, a voice recognition and semantic standardization processing module, a feature extraction and emergency degree quantification module, an intelligent evaluation and emergency instruction judgment module and a priority scheduling and response execution module, voice signals are continuously monitored through a multi-microphone array of the wearable terminal, the direction of a voice source is enhanced by using a far-field beam forming technology, the signals are processed through a time domain / frequency domain noise reduction algorithm, and finally voice streams are cached in real time with millisecond-level sampling precision. Through the multi-microphone array and far-field beam forming technology, voice capture is optimized, noise is reduced, real-time voice recognition and semantic analysis of the edge server are combined, the priority of emergency instructions is quantified, it is ensured that resources are rapidly scheduled to process emergency tasks, the agricultural production efficiency is improved, and crop losses are reduced.
Owner:CHONGQING XINDA ZHISHENG TECH CO LTD

Cross-scene forged voice detection method based on multi-level voice representation

PendingCN121884830ASpeech recognitionVoice sourceSupervised learning
The invention discloses a cross-scene forged voice detection method based on multi-level voice representation. According to the method, a voice sample is obtained from a to-be-detected voice source and input into a forged voice detection model subjected to end-to-end joint training for detection, and the model comprises a self-supervised learning voice representation module, a hierarchical time attention network and a lightweight classifier. The self-supervised learning speech representation module is used for extracting multi-level speech representations of speech samples, the hierarchical time attention network fuses the multi-level speech representations to obtain discriminant features, and the lightweight classifier outputs authenticity scores and forgery scores based on the discriminant features. And determining a detection conclusion of the voice sample according to the score result. Through hierarchical fusion and end-to-end joint training of multi-level speech representation, the stability and generalization ability of forged speech detection under a cross-scene condition are improved, and the detection accuracy and reliability are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Electroencephalogram auditory speech extraction model training method and device, equipment and storage medium

The invention relates to an electroencephalogram auditory speech extraction model training method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring electroencephalogram data and mixed voice data; performing feature extraction on the mixed voice data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used for representing the change trend of the electroencephalogram signal data along with time, and the electroencephalogram space features are used for representing the change features of the electroencephalogram signal data among different electroencephalogram channels, then voice embedding features are obtained through fusion, and first prediction audio data of the target sound source are separated from the mixed voice data; and calculating the loss of the main model according to the first predicted audio data and the original audio data of the target sound source so as to obtain a trained electroencephalogram auditory speech extraction model. By adopting the method, the voice source concerned by the listener can be accurately identified.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Training method and system for self-supervised training predictor for speech separation

ActiveCN115762557Bimprove performanceSelf-supervised training feature improvementSpeech recognitionSound sourcesVoice source
The embodiment of the present application provides a training method and system of a self-supervised training predictor for speech separation. The method comprises: extracting self-supervised training features of each single person voice source speech by using a pre-training model; extracting shallow features for speech representation and deep features for context information in the self-supervised training features, and determining the shallow features and the deep features of each single person voice source speech as training labels of the self-supervised training predictor; inputting training mixed speech generated by each single person voice source speech into the self-supervised training predictor to obtain estimated features of each single person voice source speech; and training the self-supervised training predictor based on a loss function determined based on the estimated features and the training labels corresponding to each single person voice source speech. The self-supervised training predictor is trained and applied in a speech separation model in the embodiment of the present application, so that the accuracy of the self-supervised training features is improved, the performance of the speech separation system is improved, and the model parameters and the calculation complexity are reduced.
Owner:AISPEECH CO LTD

A multi-screen control sovereignty dynamic determination method based on in-vehicle voice source positioning

PendingCN122420709AVoice sourceSpeech sound
This invention relates to the field of dynamic control sovereignty determination technology, specifically to a multi-screen dynamic control sovereignty determination method based on in-vehicle voice source localization. The method includes acquiring the voice source localization result of the current in-vehicle voice command, generating the target semantic mapping path of the voice command, constructing control sovereignty candidates, and forming a control mapping matrix by combining the current screen status information. The method selects the screen node with the optimal response link as the current voice sovereignty screen, monitors the control behavior feedback information of the sovereignty screen node, and if no behavior feedback matching the voice command is detected within a preset time window, a sovereignty correction mechanism is triggered, and the sovereignty is re-determined based on the next-level preferred node in the mapping path. This invention improves the system's interactive fault tolerance capability under dynamic, variable, and abnormal conditions, and solves problems such as voice control link breakage, retry failure, and user experience interruption in existing technologies.
Owner:EAST EVER TECH CO LTD

Method for generating a spatial voice signal and a device thereof, method for receiving a spatial voice signal and a device thereof

A method for generating a spatial voice signal performed by a voice pickup device, the method comprises: obtaining, by a plurality of microphones associated with the voice pickup device, a plurality of audio signals; determining a number of voice sources based on the plurality of audio signals; extracting, for each of the voice sources, a voice signal based on the plurality of audio signals, so as to generate a plurality of voice signals corresponding to the number of voice sources; mapping the plurality of voice signals into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal, wherein for each of the voice sources, a time delay exists between the first synthesis signal and the second synthesis signal, and the time delay is associated with a relative location between the voice source and the plurality of microphones; transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal.
Owner:HARMAN INT IND INC +1

Multilingual speech and semantic intelligent translation method and system applied to exhibition scene

The invention discloses a multilingual speech semantic intelligent translation method and system applied to an exhibition scene, and belongs to the technical field of machine translation, and the method comprises the following steps: S1, obtaining a multi-person question judgment result; s2, if the multi-person questioning judgment result is multi-person questioning, audio identification information is obtained through analysis, and otherwise, the audio identification information is directly obtained through analysis; s3, obtaining each storage question keyword, each contrast question keyword and a key matching weighting factor corresponding to each question keyword; s4, obtaining a same-group evaluation result, if the same-group evaluation result is the same group, analyzing to obtain the comprehensive matching similarity of the storage groups, otherwise, analyzing the comprehensive matching similarity of each parallel storage group; s5, obtaining a comprehensive matching judgment result, if the comprehensive matching judgment result is unqualified, performing secondary refining processing to obtain a question and answer, and otherwise, directly obtaining the question and answer; and S6, voice broadcasting is carried out, and accurate separation and language recognition of voice sources of different questioning users are achieved.
Owner:ZHEJIANG HUIZHAN ELF TECHNOLOGY CO LTD

Robust binaural beam forming method

The invention provides a robust binaural beam forming method. The robust binaural beam forming method comprises the steps of 1, calculating noisy signals received by all microphones in a binaural hearing aid; 2, calculating a covariance matrix of noisy signals and total noise; 3, calculating the binaural cue loss of the expected voice source and the jth interference source; step 4, calculating the worst binaural cue retention condition of the expected voice source; 5, the formula (13) in the step 4 is simplified by making the two norm of e smaller than or equal to a constant eta and utilizing a Cauchy-Schwarz inequality; 6, listing a cost function according to the total noise covariance matrix Rnn obtained in the step 2 and the Pa obtained in the step 5; and step 7, decomposing the matrix # imgabs0 # through a singular value decomposition method, and solving filters wL and wR by using a CVX toolbox in Matlab. According to the invention, under the condition of HRTF mismatch, the spatial cues are reserved as much as possible, and balance between spatial cue reservation and noise reduction effect is realized through parameter adjustment.
Owner:INNER MONGOLIA UNIVERSITY

Intelligent air conditioner control method and device based on voice information

The invention discloses an intelligent air conditioner control method and device based on voice information. The method comprises the steps that the voice information of a target area is obtained; extracting at least one piece of target voice information in the voice information, wherein the target voice information comprises one or more of voice source position information sending the voice information, voice source equipment information of the voice information and voice source user information of the voice information; determining voice demand information for sending the voice information according to the target voice information; and air conditioner control parameters of the intelligent air conditioner are generated based on the voice demand information, and the intelligent air conditioner is controlled to execute control operation matched with the air conditioner control parameters. It can be seen that the accuracy of controlling the air conditioner through voice can be improved, the intelligence of controlling the air conditioner through voice can be improved, and then the convenience, comfort and use experience of controlling the air conditioner through voice can be improved for a user.
Owner:FOSHAN VIOMI ELECTRICAL TECH

A large model-based human-computer voice precise interaction method

The present application belongs to the technical field of voice interaction, and particularly relates to a human-machine voice precise interaction method based on a large model. In view of the problems of voice source confusion, low instruction recognition rate and insufficient interaction safety caused by the complex acoustic environment in the vehicle cabin, the present application fuses reverberation features and harmonic attenuation features to construct acoustic fingerprints, and realizes high-precision sound source positioning in combination with seat occupancy state verification; performs priority filtering and semantic pre-alignment on non-main driver voice, and eliminates fragmented interference; introduces a large language model to standardize the mapping of structured text into instructions conforming to syntax, format and semantic specifications, and embeds a dynamic permission verification mechanism based on sound source position. The scheme significantly improves the voice instruction recognition accuracy and system robustness, effectively suppresses false triggering, enhances user privacy protection, provides an efficient and precise voice interaction experience for the intelligent cabin, and has outstanding technical practical value and industrial application prospects.
Owner:JINAN BOSAI NETWORK TECH CO LTD

AI voice transcription input method and system based on engineering geology term model

The invention relates to the technical field of geological information, in particular to an AI voice transcription input method and system based on an engineering geological term model, and the method comprises the steps: processing a field environment through the combination of dual-microphone noise reduction and acoustic echo cancellation, and obtaining pure voice source data containing lithology and parameters; calling a terminology library and a model, translating the terminology library and the model into a tagged text, and correcting an easy-to-confuse expression; identifying association by means of a field association degree algorithm and a tree structure, and judging integrity; missing information is supplemented and recorded, and double input is supported through voice prompt and highlight pop-up window guidance; the data and the image are bound through a unique identifier, structured storage is carried out, and triple verification is carried out; the system comprises a data acquisition and transfer module, a dynamic interaction and prompt module, a multi-modal fusion and storage module and a data detection module which are respectively responsible for voice transfer, semantic completion, data storage and quality verification. According to the invention, precise transfer and automatic processing of field voice are realized, the term recognition accuracy is improved, the labor cost is reduced, and the data reliability is guaranteed.
Owner:POWERCHINA BEIJING ENG CORP

End-to-end speech separation algorithm based on speech language model

PendingCN121583282ASpeech analysisAudio restorationVoice source
The invention discloses an end-to-end speech separation algorithm based on a speech language model, and the algorithm comprises the steps: discretizing a continuous audio into a 32-order discrete codebook sequence through a residual vector quantization coder-decoder, and introducing a transcription start symbol lt through an SOT strategy; sOSgt, SOSgt; a special separator is lt; sCgt; and a termination symbol lt; eOSgt, EOSgt; splicing a multi-person voice sequence; extracting audio depth features by using a pre-trained WavLM model, and guiding an autoregression decoder to output a separated zero-order codebook sequence in combination with a cross attention mechanism; predicting a high-order codebook sequence step by step through a non-autoregression model, configuring an independent embedding layer to fuse low-order information, and introducing a task embedding mechanism to optimize modeling; based on a special separator lt; sCgt; and slicing the multi-order discrete codebook sequence, and outputting an independent voice source through an Encodec decoder. According to the method, the intelligibility of voice separation and the audio restoration quality can be effectively improved, the decoding speed is high, the subjective hearing experiment result and the downstream task performance are excellent, the scene that the number of speakers is unknown is supported, and the industrialization application prospect is wide.
Owner:SHANGHAI JIAOTONG UNIV

Circuit of novel bubble toy

A circuit of a novel bubble toy is provided, including a power module E, an execution module, a manual control module, a signal amplification module, and a voice source signal input module. The power module is a 4.5V direct current power source. The power module E includes a port VCC, a port GND, and a port MIC. In the designed circuit, controlling a motor through a switch is reserved. Meanwhile, a voice can be input to a microphone MIC, so that the motor can also be controlled to operate through the voice, thus achieving the same effect as controlling the motor through the switch. For a user, there is one more control mode and one more playing method.
Owner:LIN HUAZUN

Speaker positioning method and device and storage medium

The invention discloses a speaker positioning method and device and a storage medium, and the method comprises the steps: collecting the micro-Doppler features of a target region, and extracting the throat vibration frequency information from the micro-Doppler features, so as to recognize whether an effective sound source exists in the target region or not; under the condition that the existence of the effective sound source is detected, collecting a depth point cloud for the target area, and at least extracting the three-dimensional coordinates of the mouth of the speaker by executing key point detection; whether synchronous consistency exists between the throat vibration frequency information and the mouth three-dimensional coordinates or not is verified; and when the synchronization consistency exists, determining the positioning information of the speaker according to the three-dimensional coordinates of the mouth. Therefore, the accuracy and the anti-interference capability of voice source positioning are improved by utilizing dual verification of physical vibration characteristics and spatial geometric information, and an efficient and stable speaker positioning effect can be kept in a complex and noisy environment.
Owner:AISPEECH CO LTD

Method and apparatus for speech source separation based on convolutional neural network

This article describes a method for speech source separation based on a convolutional neural network (CNN), the method comprising the following steps: (a) providing multiple frames of time-frequency transforms of an original noisy speech signal; (b) inputting the time-frequency transforms of the multiple frames into an aggregated multi-scale CNN having multiple parallel convolution paths; (c) extracting and outputting features from the input time-frequency transforms of the multiple frames through each parallel convolution path; (d) obtaining an aggregate output of the outputs of the parallel convolution paths; and (e) generating an output mask for extracting speech from the original noisy speech signal based on the aggregate output. This article also describes an apparatus for CNN-based speech source separation and a corresponding computer program product, the computer program product comprising a computer-readable storage medium having instructions, the instructions being suitable for performing the method when executed by a device having processing capabilities.
Owner:DOLBY LABORATORIES LICENSING CORP

Voice instruction recognition method and system for composite robot

The invention discloses a voice instruction recognition method and system for a composite robot, and the method comprises the steps: obtaining a plurality of voice signals, carrying out the preprocessing of the voice signals, obtaining independent voice instruction signals corresponding to different voice sources, extracting the voice features of the independent voice instruction signals, generating an instruction text, and carrying out the recognition of the voice instruction of the composite robot. And analyzing instruction semantics of the instruction text, matching the instruction semantics with pre-stored action task priority information, obtaining priority information of action tasks corresponding to the instruction semantics, respectively evaluating corresponding confidence coefficients, and calculating a multi-source priority information fusion score through a dynamic confidence coefficient weighted fusion mode. And selecting and executing an action task according to the multi-source priority information fusion score. According to the method, the literal meaning of an instruction can be understood, and the real urgency behind the instruction can be perceived, so that a decision conforming to an actual situation can be quickly and accurately made in an emergency, and the operation safety and the task continuity are ensured.
Owner:WUXI INSTITUTE OF TECHNOLOGY

Voice masking method, apparatus, and system, and vehicle

A voice masking method, apparatus, and system are provided, which are applicable to the field of intelligent vehicles. The method includes: determining a voice source location and a target masking location; receiving a sound signal from the voice source location, and detecting whether a voice signal exists in the sound signal, to generate a detection result; and if the detection result is that the voice signal exists in the sound signal, generating a masking sound for the voice signal, and outputting the masking sound to a first speaker at the voice source location and a second speaker at the target masking location; or if the detection result is that no voice signal exists in the sound signal, skipping generating the masking sound. According to this application, a requirement of an occupant for private voice communication in vehicle cockpit space can be met, and information security can be improved.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

A method and system for assisting the operation of a military operator

The application discloses a kind of army traffic service's business auxiliary method and system, it is related to voice data analysis technical field, including the multi-voice source separation processing of voice information, the voice source of the simultaneous speech of multiple people in call is separated, to accurately analyze the content that each person said, provide reliable basis for subsequent voice analysis and text analysis etc..The text content of each call is identified, and the key information is extracted from the text content.According to the situation of key word, the key information is extracted.The association between multiple calls on the call information time axis is identified, and multidimensional evaluation is carried out, to improve the effectiveness and reliability of traffic service business, to ensure the quality of army traffic communication, to integrate all the evaluation indexes of associated call events on the call information time axis, to evaluate the quality of traffic personnel's service, to improve the efficiency and accuracy of traffic service, to ensure the effective circulation of army information communication.
Owner:SHAANXI LINGFANG TECH CO LTD

Speech enhancement and separation method and device based on directional convolution beamforming

ActiveCN119943085BSpeech analysisHigh level techniquesDistortion freeNoise
The present disclosure provides a speech enhancement and separation method and device based on directivity convolution beamforming, which comprises the following steps: firstly, constructing a history observation signal matrix from a microphone array through a multi-channel observation signal, and constructing a Kalman gain linear prediction error model according to the history observation signal matrix; secondly, constructing a minimum variance distortionless response beamforming model based on a directivity beamformer and a maximum null beamformer, and estimating a speech plus noise covariance matrix; thirdly, using a time-varying variance based on a separated speech source estimation, combining the Kalman gain linear prediction error model and the minimum variance distortionless response beamforming model to establish a directivity convolution beamforming model, and completing the enhancement and separation of the speech signal through an alternating iteration method. The present disclosure can better suppress early reverberation and reduce late reverberation residues in a noisy reverberation environment; and through the construction of directivity gain and the real-time estimation of a noise covariance matrix, more robust speech separation performance is achieved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Device for processing voice and operation method thereof

Disclosed is a voice processing device. The voice processing device comprises a memory and a processor configured to perform sound source isolation on voice signals associated with the voices of speakers on the basis of the sound source positions of the respective voices. The processor is configured to: generate sound source position information indicating the sound source positions of the respective voices using the voice signals associated with the voices; generate isolated voice signals associated with the voices of the respective speakers from the voice signals on the basis of the sound source position information; and match the isolated voice signals and the voice source position information and store the same in the memory.
Owner:AMOSENSE CO LTD

Microphone sound source array positioning system and method

The invention provides a microphone sound source array positioning system and method. The microphone sound source array positioning system comprises a microphone array unit, a sound wave processing unit and a sound source recognition unit, the microphone array unit comprises a microphone algorithm module and a microphone array module; the sound source identification unit comprises a sound source geometric model construction module, a sound source geometric model identification module, a sound wave radiation identification module and a sound wave radiation distinguishing module; according to the invention, the three-dimensional microphone array is matched with the analysis of the reflected sound wave path, the high-precision sound source positioning in the vehicle is realized, the three-dimensional position of the sound source is accurately positioned by analyzing the path, delay and intensity of the sound wave and reconstructing the sound wave propagation trajectory, the direction and position of the sound source can be quickly identified without an additional sensor, and the positioning accuracy of the sound source is improved. Voice sources in all directions in the vehicle are distinguished through delay detection and intensity analysis, the complex in-vehicle space is effectively dealt with, and the intelligence of in-vehicle voice interaction and in-vehicle personnel distribution perception is improved.
Owner:HEBEI CHUGUANG AUTO PARTS CO LTD

Circuit of novel bubble toy

A circuit of a novel bubble toy is provided, including a power module E, an execution module, a manual control module, a signal amplification module, and a voice source signal input module. The power module is a 4.5V direct current power source. The power module E includes a port VCC, a port GND, and a port MIC. In the designed circuit, controlling a motor through a switch is reserved. Meanwhile, a voice can be input to a microphone MIC, so that the motor can also be controlled to operate through the voice, thus achieving the same effect as controlling the motor through the switch. For a user, there is one more control mode and one more playing method.
Owner:LIN HUAZUN

Human-machine voice precise interaction method based on large model

The invention belongs to the technical field of voice interaction, and particularly relates to a man-machine voice precise interaction method based on a large model. In order to solve the problems of voice source confusion, low instruction recognition rate, insufficient interaction safety and the like caused by a complex acoustic environment of a compartment, an acoustic fingerprint is constructed by fusing reverberation characteristics and harmonic attenuation characteristics, and high-precision sound source localization is realized by combining seat occupation state verification; performing priority filtering and semantic pre-alignment on the non-main driving voice, and eliminating fragmented interference; a large language model is introduced to standardize and map a structured text into an instruction conforming to grammar, format and semantic specifications, and a dynamic permission verification mechanism based on a sound source position is embedded. According to the scheme, the voice instruction recognition accuracy and the system robustness are remarkably improved, false triggering is effectively inhibited, user privacy protection is enhanced, efficient and accurate voice interaction experience is provided for the intelligent cockpit, and the technical practical value and the industrial application prospect are prominent.
Owner:JINAN BOSAI NETWORK TECH CO LTD

Voice analysis support device, voice analysis support method and voice analysis support program

To provide a voice analysis support device enabling semantic classification of voices according to specific intentions of an analyst in voice analysis such as VOC.SOLUTION: A voice analysis support device for analyzing voice information relating to products or services collected from voice owners, includes acquisition means for acquiring information relating to specific intentions of an analyst regarding a product or a service, acquisition means for acquiring voice information, extraction means for extracting phrases representing voices regarding the product or the service from the voice information, classification means for classifying the extracted phrases into semantic categories represented by the phrases on the basis of the information relating to the specific intentions, and storage means for storing the voice information in association with the phrases and the semantic categories into which the phrase is classified.SELECTED DRAWING: Figure 6
Owner:GENERIC SOLUTION CORP

Artificial voice generation system

According to the present invention there is provided an artificial voice generation system designed to restore or augment speech for individuals experiencing voice loss or desiring an alternative vocal quality. The system comprises a controllable air pump, an airflow member delivering air to the user's oral cavity, and a sound generation member— such as a replaceable membrane cartridge—configured to vibrate and produce sound in response to airflow. A user interface allows manual or automatic modulation of airflow and voice parameters, enabling real-time adjustment of pitch, loudness, and voice quality, including male, female, or non-binary characteristics. The system may include sensors for pressure and airflow, anti-jamming features, and can be adapted for use with or without a neck stoma, making it suitable for a wide range of users, from laryngectomy patients to those with temporary or chronic voice loss. By providing a customisable, high-quality artificial voice source with improved naturalness and intelligibility, the invention addresses limitations of existing devices such as electrolarynxes and tracheoesophageal prostheses, offering a versatile and user-friendly solution for voice rehabilitation and enhancement.
Owner:LARONIX PTY LTD

Speech enhancement and separation method and device based on directional convolution beam forming

ActiveCN119943085ASpeech analysisHigh level techniquesDistortion freeNoise
The invention provides a speech enhancement and separation method and device based on directional convolution beam forming, and the method comprises the steps: constructing a historical observation signal matrix through a multi-channel observation signal according to a microphone array, and constructing a Kalman gain linear prediction error model according to the historical observation signal matrix; constructing a minimum variance undistorted response beam forming model based on a directional beam former and a maximum null beam former, and estimating a voice and noise covariance matrix; a directional convolution beam forming model is established by using the time-varying variance based on separated voice source estimation, a simultaneous Kalman gain linear prediction error model and a minimum variance undistorted response beam forming model, and enhancement and separation of voice signals are completed in an alternate iteration mode. According to the invention, the early reverberation can be better inhibited and the late reverberation residue can be reduced in the noise-containing reverberation-containing environment; and through construction of directional gain and real-time estimation of a noise covariance matrix, more robust voice separation performance is realized.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Voice processing method and device

The invention provides a voice processing method and device, and relates to the technical field of data processing. The method comprises the following steps: in response to a to-be-processed voice source from a wake-up voice area, respectively carrying out online voice recognition and offline voice recognition on the to-be-processed voice to obtain a first online voice recognition result and a first offline voice recognition result; performing online and offline natural language understanding on the first online voice recognition result and the first offline voice recognition result respectively; sending the first online natural language understanding result and the first offline natural language understanding result to a natural language understanding result queue; and performing arbitration processing on the natural language understanding result based on the natural language understanding result queue, and outputting an arbitration processing result. According to the voice processing of the wake-up voice region, the offline ASR and online ASR parallel can be realized, the situation of time waste caused by the fusion of the offline ASR and the online ASR is avoided, and the voice processing efficiency is improved. Individual offline or individual online results may be judged to be available by arbitration processing.
Owner:BEIJING CHJ AUTOMOTIVE TECH CO LTD