Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

912 results about "Speech characteristics" patented technology

Speech characteristics are features of speech that in varying may affect intelligibility. They include: Articulation. Pronunciation. Speech disfluency. Speech pauses. Speech pitch. Speech rate.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and / or configure various settings associated with processing the speech on the ear-worn device.
Owner:FORTELL RESEARCH INC

Artificial intelligence speech recognition system

The invention discloses an artificial intelligence speech recognition system, and the system comprises a multi-modal feature extraction module which employs an improved Conformer architecture to synchronously extract the time-frequency features and text embedding vectors of speech signals; the joint training module is used for performing joint optimization on ASR and NMT loss functions through an adversarial training strategy, learning voice recognition and machine translation tasks at the same time through joint training, and completing direct mapping from voice features to a target language; the context perception translation engine is used for integrating an attention mechanism of a pre-training language model, carrying out deep coding on the extracted speech features and generating cross-language semantic representation; the self-adaptive post-processing module is used for dynamically optimizing an output result by adopting a reinforcement learning framework, dynamically adjusting the output result according to a reward function, and optimizing translation quality and a speech synthesis effect; the dynamic language recognition module is a real-time language classifier based on a Wave2Vec 2.0 framework and is used for recognizing the language of the input voice in real time; and the incremental field adaptation module is used for quickly updating a field term library by using a LoRA fine tuning technology.
Owner:ANKANG UNIV

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-mode emotion recognition method, system, electronic device and storage medium

Disclosed are a multi-mode emotion recognition method, a system, an electronic device, and a storage medium. The method includes obtaining a spectrogram of a voice to be recognized and a corresponding text and inputting the spectrogram and the text into a multi-mode emotion recognition model to obtain an emotion recognition result output by the multi-mode emotion recognition model. The multi-mode emotion recognition model is trained based on a sample spectrogram, and a corresponding sample text, and a sample emotion recognition result, and is configured to extract a feature from the spectrogram and the text by a self-attention mechanism to obtain the voice features and the text feature, fuse the text feature and voice feature to obtain a multi-mode fusion feature, and make an emotion classification decision to obtain an emotion recognition result based on the text feature, the voice feature, and the multi-mode fusion feature.
Owner:HUAZHONG NORMAL UNIV

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

LED display screen interaction method and system supporting voice interaction

The invention provides an LED display screen interaction method and system supporting voice interaction, and the method comprises the steps: carrying out the voice feature extraction processing through obtaining a continuous voice data stream sent by a user, and obtaining an acoustic feature sequence and a semantic association feature set of a voice segment; and calling a pre-trained voice semantic understanding model to perform joint semantic analysis processing on the acoustic feature sequence and the semantic association feature set, generating a semantic understanding result of the voice segment, and determining a user interaction intention type corresponding to the continuous voice data stream and key information positioning features of the user interaction intention in the voice content based on the semantic understanding result. And generating an LED display control instruction containing an interaction content identifier based on the user interaction intention type and the key information positioning feature, and sending the LED display control instruction to the target LED display screen to execute an interaction display operation. The real intention and key information in the continuous voice of the user can be accurately understood, the accuracy and naturalness of voice interaction between the LED display screen and the user are remarkably improved, and the interaction experience of the user is effectively improved.
Owner:SHANXI LAMPSON TECHNOLOGY CO LTD

Speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis

The invention relates to a speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis, and relates to the field of speech recognizing.The speech recognition and transcription method comprises the steps that a target speech signal and auxiliary modal information of synchronous visual information and text context information are obtained firstly, and the speech signal is segmented and recognized to obtain speech feature vectors; the method comprises the following steps: extracting text context information to obtain a text auxiliary feature vector, carrying out multi-modal fusion on the text auxiliary feature vector and the text auxiliary feature vector to generate fusion feature representation so as to carry out voice transcription to obtain an initial transcription text, and carrying out sentiment analysis according to visual information and the initial transcription text to generate a sentiment feature tag; and finally, optimizing and correcting the initial transliteration text based on the label to obtain a target transliteration text, thereby solving the technical problems that the speech recognition transliteration is difficult to adapt to dialect diversity and the recognition accuracy and robustness are insufficient due to neglect of emotion information, and improving the recognition accuracy and robustness through fusion of multi-modal information and emotion analysis. The voice content can be recognized more accurately, the transcription text can be optimized, and the accuracy and quality of voice recognition transcription are improved.
Owner:山西益通电网保护自动化有限责任公司

Speech coding method and decoding method based on time sequence modeling and related devices

The embodiment of the invention provides a voice coding method and decoding method based on time sequence modeling and a related device. The voice coding method comprises the following steps: acquiring target voice; extracting voice features corresponding to the voice waveform of the target voice through an input cavity convolution extraction module; performing down-sampling processing on the voice features through a down-sampling module to obtain a down-sampling sequence; performing long time sequence modeling on the down-sampling sequence through a long and short time memory neural network added with exponential activation and matrix operation to obtain a coding sequence; and converting the feature dimension of the coding sequence into 128 dimensions through the first convolutional layer to obtain a target code stream. According to the embodiment of the invention, the feature extraction module, the down-sampling module, the long and short term memory neural network and the output cavity convolution layer code the target voice, so that the compression of the voice with the extremely low bit rate can be realized, and the reconstruction of the voice with the extremely low bit rate and high quality under 150bps is realized.
Owner:SUN YAT SEN UNIV

Method for driving emotion interaction of intelligent device based on multi-modal understanding

The invention relates to the technical field of data processing, in particular to a method for driving emotion interaction of an intelligent device based on multi-modal understanding, and aims to eliminate illumination and noise interference and output a standardized face video stream, an effective voice segment and a touch thermodynamic diagram through an environment adaptive acquisition module. The feature extraction module extracts facial action optical flow features, voice Mel-frequency cepstral coefficient vectors and tactile pressure gradient parameters. The cross-modal correlation model adopts a tensor decomposition algorithm to calculate a space-time correlation matrix of visual and voice features, and the tactile feature weight is dynamically adjusted in combination with environmental parameters. According to the response strategy, an intervention scheme is retrieved based on a graph database, emotion confirmation statements, guide statements and behavior suggestions are fused to generate multi-mode response, and PID adjustment of the temperature control device and tactile pulse output of the vibration device are synchronously driven. And the feedback evaluation module verifies the emotion recognition consistency through a Pearson's correlation coefficient, triggers conflict sample separation storage and model increment training, and realizes closed-loop optimization.
Owner:BEIJING HAOXINQING MOBILE MEDICAL TECH CO LTD

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Man-machine interaction method for electronic screen control

The invention belongs to the field of man-machine interaction, and particularly discloses a man-machine interaction method for electronic screen control, which comprises the following steps: acquiring touch track data of a user according to a touch sensor, and performing feature extraction on the touch track data to obtain a touch feature vector; the method comprises the following steps: acquiring gesture action data of a user according to image acquisition equipment, and performing key point detection on the gesture action data to obtain a gesture feature vector; the method comprises the following steps: acquiring voice instruction data of a user according to an audio acquisition device, and performing acoustic feature extraction on the voice instruction data to obtain a voice feature vector; acquiring current environment parameters in real time through an environment detection module, wherein the environment parameters comprise illumination intensity, environment noise level and distance between a user and a screen; the objective of the invention is to solve the problem of insufficient reliability of a man-machine interaction mode in a complex environment in the prior art.
Owner:SHENZHEN SAIBO YUHUA ELECTRONIC TECH CO LTD

Speech encoder training method and apparatus, device, medium, and program product

Disclosed are a speech encoder training method performed by a computer device. The method includes: masking a first sub-feature representation at a first feature position in a first text feature representation to obtain a first masked feature representation; performing feature prediction on a masked first feature position in the first masked feature representation based on a first speech feature representation to obtain a first predicted feature representation; and training a first speech encoder based on a difference between the first predicted feature representation and the first sub-feature representation to obtain a second speech encoder. The first speech encoder is trained by combining data in a speech modality with data in a text modality, and information included in the data in the text modality is adopted so that the first speech encoder can learn relatively high-level semantic representations of speech, thereby improving the prediction accuracy of representations.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Fan blade voiceprint monitoring method and device based on deep learning

The embodiment of the invention provides a fan blade voiceprint monitoring method and device based on deep learning, and the method comprises the steps: obtaining noise reduction voice signal data, obtaining a multi-scale voice feature matrix, carrying out the fusion of a fan blade rotation phase label and an initial high-dimensional acoustic embedding vector, and obtaining a phase-guided acoustic embedding vector; inputting the locally enhanced acoustic embedding vector into a multi-head self-attention module of an improved Conform network, and combining position coding to obtain a global context acoustic embedding vector; inputting the first feed-forward module into a second feed-forward module of the improved Conform network to obtain a Conform output embedded sequence; inputting the Conformer output embedding sequence into a frame-level attention mechanism to obtain a frame-level weighted embedding sequence, inputting the segment-level weighted acoustic embedding vector into a joint loss optimization network, outputting a voiceprint feature vector through the joint loss optimization network, performing similarity comparison on the voiceprint feature vector and a preset fan blade operation state voiceprint library, and obtaining a voiceprint feature vector; and judging the current blade acoustic state category.
Owner:DATANG ENVIRONMENT IND GRP

Artificial intelligence-based speech emotion recognition method, device, equipment and medium

The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, and medium for speech emotion recognition based on artificial intelligence. The method performs frame segmentation and windowing processing on speech information to be recognized to obtain a speech frame sequence, extracts a speech feature tensor and a text feature tensor of the speech information to be recognized, aligns the speech feature tensor and the text feature tensor and performs feature extraction to obtain multimodal features, performs average pooling processing and global maximum pooling processing on the speech frame sequence using a local window to obtain enhanced speech features, performs feature fusion on the enhanced speech features and the multimodal features, determines a fusion result, and obtains an emotion recognition result based on the fusion result. The method obtains low-level enhanced speech features through pooling processing, and performs feature fusion on the enhanced speech features and the multimodal features, thereby avoiding the degradation problem of deep networks, effectively improving the accuracy of speech emotion recognition while improving the generalization ability of the model.
Owner:PING AN TECH (SHENZHEN) CO LTD

Internet of Things smart home voice terminal control method and system based on cloud platform

The invention relates to an Internet of Things smart home voice terminal control method and system based on a cloud platform, and belongs to the technical field of smart home voice control. The method comprises the following steps: by taking home position data of a user currently performing voice interaction as priori knowledge, performing weight dynamic adjustment on acquired voice data to complete adaptive beam forming processing; feature extraction is carried out on the voice data after beam forming processing, voice features are packaged into an encrypted data packet, and the encrypted data packet is sent to a cloud platform through an optimal transmission path; and constructing an instruction library associated with information on the cloud platform, and matching a device instruction in the instruction library through semantics analyzed by the voice features. According to the method, dual-frequency WiFi positioning and sound source positioning are fused, self-adaptive beam forming is driven by priori knowledge, reverberation elimination and weight dynamic adjustment are combined, and the voice definition and recognition rate under strong noise are remarkably improved; an instruction library is mined and constructed through association rules, and accurate instruction matching is realized by adopting multi-dimensional arbitration.
Owner:GUANGZHOU HAIMESI CLOUD TECHNOLOGY CO LTD

Smart classroom interaction analysis method based on double-layer architecture voice segmentation

The invention provides a smart classroom interaction analysis method based on double-layer architecture voice segmentation, and relates to the technical field of voice segmentation, and the method specifically comprises the following steps: extracting the voice features of a voice signal through employing a Mel-frequency cepstrum coefficient MFCC; designing a text-enhanced multi-scale time sequence-based perception time delay neural network, performing coarse screening on voice features, and dividing an audio clip into a single-speaker clip and a multi-speaker clip; and inputting the coarsely screened multi-speaker segment into a sliding window segmentation model SW-NIF fused with adjacent window information, and positioning speaker conversion points in the multi-speaker segment. And training the constructed model on the data set and verifying the model. According to the technical scheme, the problems that in the prior art, the segmentation problem of the classroom audio is neglected, and only the classroom audio is simply segmented for subsequent tasks, so that speakers in audio clips are mixed, and the analysis effect is affected are solved.
Owner:SHANDONG UNIV OF SCI & TECH

Normalizing flows with neural splines for high-quality speech synthesis

Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing generative text-to-speech models. The techniques include identifying a mapping of speech characteristics (SC) on a target distribution of a latent variable using a non-linear transformation for at least a subset of the SC. Parameters of the non-linear transformation are determined using a neural network that approximates a statistics of the SC with a statistics predicted for the SC based on the identified mapping and the target distribution of the latent variable.
Owner:NVIDIA CORP

Automatic emergency braking threshold value adjusting method and system based on in-cabin multi-mode information

The invention relates to an automatic emergency braking threshold value adjusting method and system based on in-cabin multi-mode information, and relates to the technical field of auxiliary driving, the method comprises the steps that visual information and audio information of a driver and at least one passenger in a cabin are obtained, the visual information comprises facial expressions and head postures, and the audio information comprises audio information of the driver and the at least one passenger; the audio information comprises voice features and non-voice sounds; respectively calculating a driver state coefficient and a passenger state coefficient based on the visual information and the audio information; fusing the driver state coefficient and the passenger state coefficient to obtain an in-cabin safety situation coefficient; and according to the in-cabin safety situation coefficient, a triggering threshold value of the automatic emergency braking system is adjusted in a self-adaptive mode. According to the method, multi-modal information such as vision and audio in the cabin is fused, and state analysis of all passengers is combined, so that intelligent self-adaptive adjustment of the AEB braking threshold value can be realized, and the sensing accuracy and decision reliability of an emergency braking system in a complex scene are improved.
Owner:ZHIJI AUTOMOTIVE TECH CO LTD

Intelligent conference voice transcription and summary generation method and system for voiceprint recognition and hot word optimization in power industry

The invention discloses an intelligent conference voice transcription and summary generation method and system for voiceprint recognition and hot word optimization in the power industry. The method comprises the following steps: acquiring voice samples in advance to construct an encrypted voiceprint database; mining terminologies from multi-source data in the power industry and dynamically allocating weights to establish a hot word library; performing unified decoding and adaptive preprocessing on the input audio; after speech features are extracted, speaker separation and recognition are achieved through deep embedding clustering; integrating the hot word bank, and generating a transliteration text with a speaker tag by adopting a transform-based field adaptive model; and carrying out multi-granularity semantic understanding and hierarchical abstract generation by utilizing the pre-training model in the power field, and automatically outputting a structured conference summary. According to the method, the problems that the accuracy of professional term recognition in the power industry is low, separation of multiple speakers is difficult, and summary generation depends on manpower are effectively solved, and the efficiency and the intelligent level of conference recording are remarkably improved.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Alzheimer disease recognition method and system based on voice features

The invention relates to the technical field of speech analysis, in particular to an Alzheimer's disease recognition method and system based on speech features, and the method comprises the following steps: extracting frame-level parameters to construct a sequence, recognizing sparse and fractured sections, generating an abnormal trend, and completing speech feature recognition. According to the method, multiple parameters such as the mean value of amplitude absolute values, the maximum difference value and the minimum difference value are serialized and integrated, a double analysis mechanism for the sparsity and the jump of the voice amplitude fluctuation is formed by combining multi-section continuous ratio comparison and mutation trend positioning, the overlapping degree of trend indexes in adjacent frame sections is calculated, and sites in a trend structure are extracted; a trend structure line of the time sequence is established, directional change and point location density of the trend structure line are extracted, quantitative classification of abnormal trends in the frame sequence is completed, cross-scale feature coupling recognition from voice micro fluctuation to time sequence trends is achieved, the discrimination degree and accuracy of voice features of the Alzheimer's disease are improved, and the recognition accuracy of the Alzheimer's disease is improved. And the stability and the discrimination efficiency of the identification result are obviously improved.
Owner:WUXI NO 2 PEOPLES HOSPITAL

Voice signal processing method and device, equipment, storage medium and computer program product

The invention relates to the technical field of voice communication, in particular to a voice signal processing method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: performing real-time noise classification on a received original voice signal based on a preset deep neural network noise recognition model; according to the noise classification result, selecting a corresponding combined noise reduction model to perform multi-stage noise reduction processing on the original voice signal to obtain a noise-reduced voice signal, the combined noise reduction model comprising one or more of a deep neural network-based noise reduction module, a Wiener filtering module and a spectral subtraction module; and on the basis of time domain feature extraction and frequency domain feature extraction, voice feature enhancement processing is performed on the noise-reduced voice signal to obtain the target voice signal, so that the transmission quality of the voice signal is improved.
Owner:SHENZHEN DINSTAR TECH

Audio processing method and device, storage medium and electronic device

The invention discloses an audio processing method and device, a storage medium and an electronic device, and relates to the technical field of data processing, and the audio processing method comprises the steps: obtaining voice feature data and semantic feature data of original audio data; the original audio data is audio data extracted from the target audio and video data; generating segmentation position data according to the voice feature data and the semantic feature data; and generating audio paragraph data and paragraph visualization data according to the segmentation position data, and providing the paragraph visualization data to a user interaction interface. According to the technical scheme of the embodiment of the invention, the utilization rate of the audio information can be improved, the data processing cost in an application scene is remarkably reduced, and the data application efficiency and scene adaptability are improved.
Owner:QINGDAO HAIER TECH +2

Voiceprint model training, voiceprint extraction method, device, equipment and storage medium

The present invention discloses a method, apparatus, device, and storage medium for training and extracting a voiceprint model. The method comprises: determining a voiceprint model and a classification model; extracting at least two voice features from a voice signal, wherein the voice signal is labeled with the user to which it belongs; initially training the voiceprint model and the classification model based on the at least two voice features with the goal of classifying the user; if the initial training is completed, continuing to train the voiceprint model and the classification model based on the at least two voice features with the goal of classifying the user and constraining the voiceprints of the same user, wherein the voiceprint model is used to extract the voiceprint from the at least two voice features, and the classification model is used to predict the user to which the voice signal belongs and is discarded upon completion of continued training. Converging the voiceprint model with the goal of constraining the voiceprints of the same user can improve the performance of the voiceprint model in scenarios such as environmental noise, cross-channel devices, and counterfeit attacks, improve the accuracy of the voiceprint, and thus improve the robustness of the voiceprint model.
Owner:BIGO TECH PTE LTD

Array microphone noise reduction recording method based on cascade noise reduction and blind source separation

The invention relates to an array microphone noise reduction recording method based on cascade noise reduction and blind source separation, which belongs to the technical field of voice signal processing and recording, and comprises the following steps: configuring a multi-channel array microphone, ensuring that the amplitude and phase of a channel signal are consistent, and collecting an original multi-channel voice signal; a weighted kernel function blind source separation algorithm is adopted, signal-to-noise ratio distribution characteristics of signals are extracted, kernel function weights are given, and target voice and interference signal components are obtained through decoupling of an independent component analysis model; executing target-oriented adaptive cascade noise reduction, locking the voice of a keynote speaker through directional pickup, reducing noise, filtering out reverberation, and enhancing the voice of a far-field target by combining a voice mask neural network with a far-field pickup algorithm in sequence; and processing the target voice through voice feature perception lossless coding and storing the target voice. According to the invention, stable acquisition of multi-channel signals, accurate separation of mixed signals and layered suppression of noise reverberation are realized, the signal-to-noise ratio and definition of far-field voice are significantly improved, and the method is suitable for single-person speaking or multi-person dialogue scenes.
Owner:SHANGHAI RONGDA DIGITAL TECH CO LTD

Speech recognition method and device, equipment, storage medium and program product

The invention provides a voice recognition method and device, equipment, a storage medium and a program product, and relates to the technical field of image processing. The speech recognition method comprises the following steps: acquiring speech to be recognized, and extracting speech features based on the speech to be recognized; obtaining video content corresponding to the to-be-recognized voice, and extracting video features based on the video content; acquiring a historical voice recognition text of the to-be-recognized voice, and extracting historical text features based on the historical voice recognition text; obtaining a first multi-modal fusion feature based on the voice feature and the historical text feature; obtaining a second multi-modal fusion feature based on the video feature and the historical text feature; and generating a speech recognition text corresponding to the speech to be recognized based on the first multi-modal fusion feature and the second multi-modal fusion feature.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

Speech video synthesis method and system

The invention provides a speech video synthesis method and system, and the method comprises the steps: inputting speech audio into a pre-trained audio-motion encoder, extracting a speech feature sequence from the speech audio, and generating a motion sequence of a face according to the speech feature sequence; inputting a single two-dimensional face picture shot from a first visual angle into a pre-trained picture encoder, and extracting three-dimensional identity features of an object contained in the face picture; inputting the three-dimensional identity features and the emotion labels into a pre-trained emotion mapping layer, and fusing to obtain emotion-containing three-dimensional identity features; the action sequence and the emotion-containing three-dimensional identity features are input into a video generation network based on a neural radiation field, a speech video of the object matched with the speech audio at a required second view angle is synthesized through the neural radiation field and camera parameters, and the second view angle can be adjusted within a preset range through the camera parameters.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Speech recognition system based on improved Transform architecture

The invention belongs to the field of artificial intelligence and voice recognition, and particularly relates to a voice recognition system based on an improved Transform architecture, which comprises a self-positioning module used for receiving an original audio signal, outputting a self-supervised voice feature vector and a traditional audio feature vector in parallel, and sending the self-supervised voice feature vector and the traditional audio feature vector to a feature normalization conversion module; the feature normalization conversion module is used for mapping the self-supervised voice feature vector and the traditional audio feature vector to a standard speaker feature space and outputting a normalized feature; the perception modeling module performs multi-scale time sequence coding through an improved Transform structure, and outputs a voice semantic probability distribution sequence; the CTC loss module is used for optimizing the acoustic model according to the voice semantic probability distribution sequence; the collaboration unit is used for receiving multiple paths of original audio features, screening credible channels from the obtained synchronization features, and outputting corrected features; and the fusion filtering module is used for receiving the local features and the corrected features, generating global probability distribution through attention weight fusion, and decoding the global probability distribution into a final text sequence.
Owner:SHENYANG LIGONG UNIV