Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50 results about "Modal voice" patented technology

Modal voice is the vocal register used most frequently in speech and singing in most languages. It is also the term used in linguistics for the most common phonation of vowels. The term "modal" refers to the resonant mode of vocal folds; that is, the optimal combination of airflow and glottal tension that yields maximum vibration.

Privacy information desensitization method and system for voice generation type large model

The invention discloses a privacy information desensitization method and system for a voice generation type large model, and relates to the technical field of artificial intelligence. Input voice data is discretized, Gaussian noise disturbance sensitive features are injected, and a cross-modal voice generation model is constructed in combination with three-stage training; and meanwhile, a cross-modal privacy enhancement mechanism is applied to detect and fuzzify sensitive information in real time in an output stage. According to the invention, the adaptability of a large-scale voice generation model in a privacy protection scene is improved, and comprehensive protection of user privacy is realized. And moreover, the capability of extracting user privacy information by an adversarial attacker is effectively limited, and the risk of sensitive information leakage in the data transmission, storage and generation process of the voice generation model is reduced. On the premise that privacy is ensured, the voice generation model can still keep high-quality generation performance, generated voice output has high naturalness and accuracy, and actual application requirements are met.
Owner:ZHEJIANG UNIV

AI multi-mode voice interaction method based on vehicle-mounted intelligent terminal and electronic equipment

The invention provides an AI multi-mode voice interaction method based on a vehicle-mounted intelligent terminal and electronic equipment, and the method comprises the steps: synchronously collecting voice information, facial micro-expressions and hand touch actions of a user through a microphone array, an in-vehicle camera and a steering wheel touch sensor on the vehicle-mounted intelligent terminal, and forming multi-mode interaction data; fusing the multi-modal interaction data, and judging a current driving scene through a pre-trained driving scene recognition model; dynamically adjusting voice interaction parameters according to the driving scene judgment result; recognizing a voice instruction of the user based on the adjusted voice interaction parameter; and in combination with instruction preferences in historical interaction data of the user, intention completion is performed on the recognized voice instruction, and a personalized execution instruction is generated. According to the invention, the recognition accuracy and interaction efficiency of the voice instruction are improved.
Owner:SHENZHEN BEIBO INTELLIGENT TECH

Multi-mode voice interaction system and method for intelligent cockpit

The invention discloses a multi-mode voice interaction system and method for an intelligent cabin. Firstly, various vehicle working condition parameters in a CAN bus are read in real time, a wavelet domain Wiener filtering and judgment statistic enhancement strategy is dynamically adjusted, and self-adaptive voice purification based on vehicle state prior is achieved. Secondly, a bidirectional long-short-term memory network is adopted as a semantic understanding core, semantic feature vectors are extracted from enhanced signals through segmented maximum pooling operation, and the problem of semantic confusion caused by manifold collapse in a traditional method is effectively solved. And the feature vector and a safety state signal are synchronously input into an improved LLM-CoT engine and a dynamic arbiter, so that context accurate understanding and safety-up priority response are realized. According to the method, vehicle working condition perception and biLSTM depth time sequence modeling are deeply adapted to a vehicle-mounted complex acoustic scene, the error rate of semantic understanding is remarkably reduced in actual measurement under multi-person dialect dialogue, and the accuracy, robustness and safety of intelligent cabin interaction are systematically improved.
Owner:SOUTHEAST UNIV

Method for realizing personalized singing synthesis model training through multi-mode voice driving

The invention relates to the technical field of voice signal processing, and discloses a method for realizing personalized singing synthesis model training through multi-modal voice driving, and the method comprises the following steps: obtaining multi-modal input data, including text, reference audio, speaker and emotional features; processing the data to obtain coding features; the inter-modal redundant information is evaluated and suppressed through redundant perception coding; compressing coding features by using an information bottleneck model, and retaining effective information; fusing the compressed features to generate personalized singing features; the input decoding module generates Mel spectrum features; and converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. And converting the Mel spectrum into an audio waveform through a vocoder, and outputting the audio waveform. According to the method, multi-modal data can be effectively fused, the individuation and emotion expression ability of the singing sound is improved, redundant information interference is reduced, the quality and efficiency of singing sound synthesis are improved, and finally, high-fidelity individualized singing sound is output.
Owner:SHENZHEN ZHONGLU CULTURE COMMUNICATION CO LTD

Spoken English pronunciation quality evaluation method based on multi-mode speech feature analysis

The invention belongs to the technical field of speech analysis, and discloses a spoken English pronunciation quality evaluation method based on multi-modal speech feature analysis, which comprises the following steps: acquiring a spoken English speech signal of a target user, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features; performing time-frequency domain corresponding relation analysis and feature extraction on the multi-mode speech features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, constructing a multi-dimensional representation model based on a multi-scale matching result, and calculating a fine-grained quality score of a phoneme unit in each multi-dimensional representation model in combination with rhythm and rhythm parameters in the multi-modal speech features, the rhythm coherence score and the overall fluency score are fused to generate a comprehensive pronunciation quality evaluation result and a visual diagnosis report of pronunciation deviation; the oral English pronunciation evaluation method realizes comprehensive and refined evaluation of oral English pronunciation, and provides a scientific guidance basis for personalized language learning.
Owner:ZHANG ZHOU HALTH VOCATIONAL COLLEGE

Video generation method and system based on voice analysis, and storage medium

The invention relates to the technical field of artificial intelligence and multimedia, and particularly discloses a video generation method based on voice analysis, and the method comprises the following steps: analyzing an input voice, and extracting a multi-modal voice feature; inputting the multi-modal speech features into a pre-trained scene association model, and outputting a scene label set; on the basis of dialect features in the input voice, region categories are recognized through a dialect classifier, and a corresponding visual element library is loaded from a culture database according to the region categories; selecting a scene template according to scene type labels in the scene label set, and selecting a character action template in combination with emotion type labels and interaction object relation labels; and calculating time sequence distribution of the video elements based on the speech speed change parameters, and rendering the scene template, the character action template and the rhythm of the input speech according to the time sequence distribution through a time sequence alignment algorithm to generate a target video. The method can improve the accuracy of the video generated based on the voice.
Owner:SHENZHEN SMART INSURANCE TECH CO LTD

Multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and medium

ActiveCN120954388ASpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The invention discloses a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and a medium, and relates to the technical field of multi-modal voice interaction.The method comprises the steps that a text token of a text training sample is obtained, a corresponding voice token is constructed, and pre-training data used for converting the text token into the voice token is obtained; in combination with the multi-modal input sample and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation; constructing and pre-training a basic model by using the pre-training data; a multi-modal voice interaction large model is constructed based on a pre-training basic model, and fine tuning data is used for training, so that voice acoustic features can be regulated and controlled based on multi-modal input, and voice is output. According to the method, through alignment and staged training of the text token and the voice token, fine regulation and control of voice acoustic characteristics are realized, long voice continuity and interaction naturalness are improved, and a model is efficiently endowed with voice interaction capability of controllable timbre and emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Multi-mode voice interaction method and device, intelligent equipment and readable storage medium

The invention provides a multi-mode voice interaction method and device and intelligent equipment, is suitable for the technical field of intelligent voice interaction, is applied to the intelligent equipment, and comprises the following steps: in response to a detected voice activity, extracting first voice data in the voice activity, and obtaining video data shot synchronously with the first voice data, a plurality of different users are shot in the video data. And screening out a target user sending the first voice data from the video data. And performing semantic integrity analysis on the text content corresponding to the first voice data. And when the semantic integrity analysis result is that the semantics is incomplete, continuously acquiring second voice data of the target user for a plurality of times, and generating a corresponding target statement with complete semantics after the first voice data and the second voice data are combined. Generating reply data according to the target statement, and outputting the reply data through voice. According to the embodiment of the invention, accurate, coherent and real-time voice interaction with the user needing interaction can be realized.
Owner:浙江人形机器人创新中心有限公司

Multi-modal speech recognition method based on large model, storage medium, electronic equipment and product

The invention relates to the technical field of voice recognition, and particularly provides a multi-modal voice recognition method based on a large model, a storage medium, electronic equipment and a product, and the method can comprise the steps: carrying out the preprocessing of an original voice signal of a user, and obtaining a processed voice signal; inputting the voice coding data and the historical dialogue data corresponding to the processed voice signal into a large language model to obtain a text vector corresponding to the processed voice signal; performing feature extraction on the processed voice signal to obtain a voice feature vector; processing a target vector sequence formed by splicing the voice feature vector and the text vector by using a pre-trained voice recognition module to obtain a text sequence; wherein the voice recognition module comprises a plurality of encoder layers and a plurality of decoder layers which are trained in advance; and cleaning and formatting the text sequence to obtain text data corresponding to the original voice signal. Some embodiments of the present application can improve the accuracy of speech recognition.
Owner:LINGXI TECHNOLOGY CO LTD

Multi-modality voice recognition device and method

A multi-modality voice recognition device is provided. The multi-modality voice recognition device includes a video encoder trained to receive a lip video to a video encoder to extract visual feature information for voice recognition, an audio encoder trained to receive a voice to extract voice feature information for voice recognition, a modality reconstructor trained to reconstruct the voice feature information from the visual feature information to generate reconstruction voice feature information, a random selector configured to randomly output one of the voice feature information and the reconstruction voice feature information, and a video-audio decoder trained to receive a multi-modality feature, where the visual feature information is connected to an output of the random selector, to output a character string which is a voice recognition result.
Owner:ELECTRONICS & TELECOMM RES INST

Multi-mode voice triggering for audio devices

The invention relates to multi-mode voice triggering for audio devices. Particular implementations of the subject technology provide systems and methods for multi-mode voice triggering of audio devices. An audio device may store a plurality of speech recognition models, each of which is trained to detect a single corresponding trigger phrase. In order that an audio device can detect a particular one of a plurality of trigger phrases without consuming processing and / or power resources to run a speech recognition model that can distinguish different trigger phrases, the audio device preloads a selected one of the speech recognition models for an expected trigger phrase into a processor of the audio device. The audio device may select one of the speech recognition models for the expected trigger phrase based on a type of companion device communicatively coupled to the audio device.
Owner:APPLE INC

Shopping guide ability improving system and method based on AI voice work cards

The invention discloses a shopping guide capability improving system and method based on an AI voice work card, and belongs to the technical field of artificial intelligence and intelligent shopping guide. The system comprises a multi-mode voice acquisition and feature extraction module, a crown cancellation model dynamic extraction and scoring module, a dialogue quality evaluation and intention analysis module and a closed-loop feedback optimization and partner training scene generation module, and the four modules form a deep coupling closed-loop cooperative system. The system extracts efficient verbal skills from crown sales samples through an innovative adaptive weight calculation algorithm, quantifies customer intention intensity through a multi-dimensional fusion method, improves shopping guide ability through interactive partner training, realizes system adaptive optimization through a closed-loop feedback mechanism, and improves customer experience. According to the invention, automatic inheritance of sales crowd experience and scientific evaluation and systematic improvement of shopping guide capability are realized, customer satisfaction and sales conversion rate are significantly improved, and the method is suitable for industries requiring shopping guide service, such as extensive home and retail.
Owner:MO DOU (DONGGUAN) INFORMATION TECHNOLOGY CO LTD

English oral pronunciation quality evaluation method based on multi-modal speech feature analysis

The present application belongs to the technical field of speech analysis, and discloses an English oral pronunciation quality evaluation method based on multi-modal speech feature analysis, which comprises the following steps: obtaining the speech signal of the target user's English oral speech, and performing multi-domain decomposition on the speech signal to obtain multi-modal speech features; performing time-frequency domain corresponding relationship analysis and feature extraction on the multi-modal speech features to obtain a pronunciation detail feature sequence; performing multi-scale matching on the pronunciation detail feature sequence and a preset standard pronunciation template, and constructing a multi-dimensional representation model based on the multi-scale matching result; calculating the fine-grained quality score of the phoneme unit in each multi-dimensional representation model in combination with the prosody rhythm parameters in the multi-modal speech features; and fusing the prosody coherence score and the overall fluency score to generate a comprehensive pronunciation quality evaluation result and a visual diagnostic report of pronunciation deviation; the present application realizes comprehensive and refined evaluation of English oral pronunciation, and provides scientific guidance basis for personalized language learning.
Owner:ZHANG ZHOU HALTH VOCATIONAL COLLEGE

Method and system for improving voice inter-translation quality of small languages based on large language model

The invention discloses a method and system for improving the voice inter-translation quality of a small language based on a large language model, and belongs to the technical field of intelligent voice translation services. The method comprises the following steps of language information input, speech detection and acquisition, semantic feature and acoustic feature extraction, automatic speech recognition (ASR) model recognition, large language model (LLM) translation and error correction, multi-modal speech synthesis (TTS) model fusion and target language speech signal output. The system comprises the following modules: a language information input module, a voice detection and acquisition module, a voice coding module, an automatic voice recognition module, a large language model translation and error correction module, a multi-modal voice synthesis module and a voice output module. According to the method, the accuracy and the naturalness of the voice inter-translation of the minority language can be remarkably improved, the method adapts to the characteristics of the multi-language minority language of the East Alliance, the industrial-grade application requirements are met, and the problems of low recognition accuracy and acoustic feature loss of the voice inter-translation of the minority language are solved.
Owner:GUANGXI DONGXIN YITONG TECH CO LTD

Multimodal speech separation method, training method, and related apparatuses

ActiveCN115620723BModal voiceSpeech sound
The application discloses a multi-modal speech separation method, a training method and related devices. The method comprises: obtaining audio and video data containing a target object; wherein the audio and video data contains lip video data of the target object; inputting the audio and video data into a trained multi-modal speech separation network to obtain audio data related to the lip video data of the target object; wherein a plurality of training samples for training the multi-modal speech separation network are divided into a plurality of subsets based on a first loss obtained after the training samples pass through the multi-modal speech separation network, and the multi-modal speech separation network is trained again based on at least part of the subsets. In the foregoing manner, the application can improve the accuracy of multi-modal speech separation.
Owner:IFLYTEK CO LTD

Multi-mode voice wake-up method and device and electronic equipment

The embodiment of the invention provides a multi-mode voice wake-up method and device and electronic equipment. The method is applied to the technical field of intelligent voice, and comprises the steps of providing a basis for subsequent accurate processing by acquiring voice, gesture and mouth shape multi-mode input data and adding a timestamp, avoiding a high misjudgment rate of a single voice mode in a complex environment, greatly improving wake-up accuracy, and when a mode conflict exists, improving wake-up accuracy. By identifying the confidence of each modal input data and combining the historical data to make a decision, the possibility of wrong wake-up operation caused by a single modal identification error is reduced, the reliability of wake-up decision is improved, if modal conflicts do not exist, the wake-up scene of the target equipment is accurately determined according to the multi-modal input data, and the wake-up efficiency of the target equipment is improved. Through accurate identification and adaptive processing of different wake-up scenes, the target device can be stably and accurately waken up in various complex environments, and the environmental adaptability of the target device is effectively enhanced.
Owner:NINGBO TELIAN INFORMATION TECH CO LTD

Full-band speech enhancement method based on earphone inertial sensor

The invention relates to the technical field of speech enhancement, in particular to a full-band speech enhancement method based on an earphone inertial sensor, which comprises the following steps: performing user registration and limited data acquisition; carrying out segmentation, denoising, dimension reduction and normalization technologies on the collected original data to obtain error-free data; converting inertial sensor data into high-dimensional pronunciation kinematics video features by using a GAN-based method; realizing cross-modal voice frequency domain information enhancement based on a double-flow codec structure; adopting an unsupervised field adaptive technology based on adversarial training to adapt the pre-training model to a new target user; and estimating voice phase information by using a phase denoising network, and reconstructing a time domain voice signal in combination with frequency spectrum information. Unified speech enhancement of full frequency bands is realized, so that an inertia channel and an acoustic channel are self-adaptively complementary in a dynamic environment, and meanwhile, low delay, low power consumption and high quality are taken into account.
Owner:NANJING UNIV OF POSTS & TELECOMM

Intelligent accompanying-oriented emotional anthropomorphic multi-mode voice interaction large model system

The invention provides an intelligent accompanying-oriented emotional personification multi-modal voice interaction large model system, which relates to the field of language processing and is characterized by comprising the following modules: a user interaction module, a multi-modal data fusion module, a database module, a data analysis module and a function processing module, the user interaction module comprises touch interaction, voice interaction, text interaction and local interaction. The intelligent accompanying toy has the advantages that by integrating core technologies such as multi-modal data fusion, hierarchical memory and RAG, dynamic emotion response and edge-cloud collaboration, the problems of emotion interaction templating, lack of role consistency, high response delay, insufficient utilization of complex modals, weak long-term memory ability and the like of an existing intelligent accompanying toy are systematically solved; and the reality sense, the individuation degree and the real-time performance of interaction are remarkably improved, so that high-quality accompanying experience with higher emotional value and immersion is provided for the user.
Owner:SUZHOU PINGPING MAINLAND TECHNOLOGY CO LTD

Multi-mode voice interaction method, device and equipment and readable storage medium

The invention discloses a multi-mode voice interaction method, device and equipment and a readable storage medium, and relates to the technical field of voice interaction. Comprising the following steps: firstly, acquiring reflected ultrasonic information reflected based on emitted ultrasonic information, and judging whether a target person exists in a target space or not according to the ultrasonic information; if the target person exists in the target space and a voice wake-up instruction is received, orientation information of the voice wake-up instruction is determined based on the microphone array, and a continuous motion track of the target person is determined according to the orientation information and the reflected ultrasonic information; and finally, noise reduction and enhancement are carried out on a voice instruction sent by the target person based on the continuous motion track, and voice interaction is carried out based on the voice instruction after noise reduction and enhancement. The distance information of the azimuth information of the target person is mastered through the reflection of the ultrasonic waves, and then the continuous motion trail of the target person is generated, so that the tracking noise reduction and enhancement of the voice instruction sent to the target person are realized, and the stability of voice interaction can be enhanced without a camera.
Owner:AISPEECH CO LTD

A cross-modal speech recognition method

The application relates to the technical field of speech recognition, and discloses a cross-modal speech recognition method. The method comprises the following steps: obtaining video information to be analyzed and extracting call audio and visual information therefrom; obtaining a speech frame sequence; performing decoding operation on the speech frame sequence by using a speech recognition model to obtain corresponding text information; performing multiple feature extraction operation on the visual information to obtain a feature sequence of the visual information; obtaining a face information sequence by using a preset visual information extraction model; performing extraction analysis on the face information sequence by using a preset target detection model, so that further optimization is performed; performing decoding operation on the optimized lip sequence by using a trained lip movement conversion model to obtain a candidate word set; and comparing and fusing by using a trained fusion neural model to output final text information. The application can solve the problem that current speech recognition is prone to interference.
Owner:HANGZHOU ZHIXIN TECH CO LTD

A voice interaction recognition enhancement method, device and storage medium

ActiveCN119207393BSpeech recognitionModal voiceLip feature
The application provides a speech interaction recognition enhancement method, which comprises the following steps: collecting a video of a speaker facing a camera for speech interaction, splitting the video into N segments of to-be-identified speech and N frames of to-be-identified images, and forming to-be-identified data; inputting the to-be-identified data into a preset speech interaction recognition enhancement model, wherein the model comprises a lip feature extraction network, a speech feature extraction network, a time feature extraction network and an activation network, the time feature extraction network is used for adding time sequence information to a lip feature matrix extracted by the lip feature extraction network and a speech feature matrix extracted by the speech feature extraction network, the activation network is used for simulating an activation-inhibition mechanism of visual information on an auditory nerve circuit in a biological brain, realizing the interaction of two modalities of vision and hearing, and obtaining a speech recognition result. The application realizes the accurate recognition of lip language and speech in a noisy environment by using the dual-modality speech recognition of vision and hearing, and the response capability and accuracy meet the requirements.
Owner:TSINGHUA UNIVERSITY

Voice code generation method and system for multi-modal identity authentication of television end

The application relates to the technical field of voice recognition, and discloses a voice verification code generation method and system for television end multi-modal identity authentication. The method comprises the following steps: obtaining a correction parameter by performing compensation processing on a living room environment acoustic signal, extracting a user voiceprint template and a spatial position feature vector to perform multi-modal identity authentication to obtain a user identity feature code, generating a user-specific voice verification code audio according to the identity feature code and the acoustic correction parameter, performing spectrum coding processing to obtain a television end multi-modal voice verification code, and finally associating and binding the television end multi-modal voice verification code with the user identity feature code to generate a television end identity authentication credential. The application realizes dynamic generation of a personalized voice verification code by fusing multi-modal features such as a user voiceprint and a spatial position, and improves the security and user experience of television end identity authentication.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1

Multi-modal speech language large model training method and device, equipment and medium

The invention provides a multi-modal speech language large model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence, in particular to the technical field of large models, speech data processing and data generation. According to the implementation scheme, first inquiry voice data is input into a multi-mode voice language large model, and first reply voice data generated by the multi-mode voice language large model is obtained; determining an inquiry text corresponding to the first inquiry voice data and a reply text corresponding to the first reply voice data; determining a first score based on the inquiry text and the reply text; determining a second score based on voice features of the first inquiry voice data and voice features of the first reply voice data, the voice features including at least one of voice sharpness, speed features, tone features, intonation features and emotion features; and adjusting parameters of the multi-modal speech language large model based on the first score and the second score.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Multi-scene adaptive AI speech recognition method and device for large-screen speech assistant

The invention provides a multi-scene adaptive AI voice recognition method and device for a large-screen voice assistant, and the method comprises the following steps: 1, constructing a large-screen special multi-mode voice instruction data set, and dividing the data set into a training set, a verification set and a test set according to a stratified sampling strategy; 2, training a lightweight real-time noise reduction model based on the training set, and outputting high-fidelity pure speech features; 3, based on the training set, training a large-screen instruction optimization recognition model fused with FunASR and Whisper features, taking the pure voice features as input, and outputting an initial text and a corresponding instruction intention; and 4, training a context semantic association model based on Transform, performing anaphora resolution and intention correction by taking the initial text as a current round of input and combining a historical dialogue context, and executing corresponding operation. According to the method, end-to-end high-robustness voice interaction is finally realized.
Owner:BEIJING MYSHER TECH

Method and apparatus for training multi-modal speech language large-scale model, device and medium

The present disclosure provides a method and apparatus for training a multi-modal speech language large scale model, a device and a medium.SOLUTION: A solution is to input first question speech data into a multi-modal large speech language model to obtain first answer speech data generated by the multi-modal large speech language model, determine a question text corresponding to the first question speech data and an answer text corresponding to the first answer speech data, and determine a first score according to the question text and the answer text. A second score is determined according to a speech feature of the first question speech data and a speech feature of the first answer speech data, where the speech feature includes at least one of a speech clarity, a speech rate feature, a timbre feature, an intonation feature, and an emotion feature.SELECTED DRAWING: Figure 2
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

A cross-modal speech and face association method based on a heterogeneous hash network

The application belongs to the technical field of communication, and discloses a cross-modal speech and face association method based on a heterogeneous hash network, and the specific technical scheme is as follows: a cross-modal deep residual alignment network based on hash learning is proposed, the network maps input features into hash codes with fixed length, effectively reduces the feature dimension and the use of memory, meanwhile, the hashization greatly improves the calculation efficiency, the deep network integrated with the residual structure has stronger feature representation capability and good generalization, by minimizing the loss function, the speech and face features of the same identity can be made closer in the Hamming distance in the hash space, and the speech and face features of different identities are farther away, so that the feature alignment of the speech and the face is achieved, and the cross-modal matching, verification and retrieval tasks are completed, and the accuracy of the model in the cross-modal matching, retrieval and verification tasks is obviously improved compared with the existing method.
Owner:XIAN UNIV OF POSTS & TELECOMM

Intelligent multi-mode voice conference interaction method and system

The invention relates to the field of voice interaction, in particular to an intelligent multi-mode voice conference interaction method and system. The method comprises the following steps: collecting historical voice signal data, carrying out feature extraction and analysis on original voice signal data, generating a voice acoustic vector, analyzing the voice acoustic vector, and constructing a voice baseline vector; collecting conference voice data, analyzing the conference voice data, and constructing a conference semantic state diagram; acquiring real-time voice signal data, and analyzing the real-time voice signal data based on the voice baseline vector to obtain a voice deviation parameter; and on the basis of the voice deviation parameter and the voice baseline vector, adjusting voice recognition adaptation, generating an adjustment strategy, stopping recording the conference semantic state diagram before conference interruption when the conference interruption is detected, and performing reconnection processing on the conference semantic state diagram after conference reconnection. According to the invention, the voice recognition suitability and the semantic understanding continuity can be improved.
Owner:JIAN XIANGE ACOUSTIC ELECTRONIC CO LTD

Speech enhancement method, neural network training method, device, equipment and medium

ActiveCN121122306ASpeech analysisBiological modelsPhonetic environmentEnvironmental noise
The invention provides a voice enhancement method, a neural network training method, devices, equipment and a medium, and relates to the technical field of voice processing, voice environment complexity analysis is carried out to determine a noise estimation value corresponding to voice data, and under the condition that the noise estimation value is greater than a target threshold value, image-assisted voice enhancement is adopted to improve the voice enhancement efficiency. Whether more information needs to be brought through image data assistance or not can be determined in a targeted mode according to the complexity of the voice environment, voice enhancement in the complex environment is effectively achieved, and therefore the method adapts to different environment complexities, meets different voice enhancement requirements, adopts image data to carry out auxiliary enhancement on voice data, and improves the voice enhancement efficiency. Compared with single-mode speech enhancement, image signals insensitive to noise are fused, the enhancement effect is better, the environmental noise tolerance is improved, and the enhanced speech data have higher robustness, so that the calculation power consumption and the calculation complexity of multi-mode fusion speech enhancement are reduced while the speech enhancement performance is guaranteed.
Owner:TSINGHUA UNIVERSITY

A multimodal speech recognition error correction method and system

The present invention discloses a multimodal speech recognition error correction method and system, comprising: obtaining original sample data from a corpus, generating error samples from the original samples using a fuzzy sound generator; annotating the initials and finals of the error sample text to construct fuzzy sound texts with different similarity levels; adjusting the parameters of the different similarity ratios of the fuzzy sound texts of the error sample data based on the annotated error sample data; constructing a speech and text fusion feature vector based on the original correct sample and error sample data, inputting the fusion feature vector into an error correction model for training, and outputting the word with the highest correct probability for each speech position through a fully connected layer and an activation function. The method and system utilize the multimodal fusion features of speech and text for training to obtain a model for customer service speech error correction. The error correction model based on the combination of text and speech can reduce the impact of dialects and environmental noise, thereby improving the accuracy of customer service speech quality inspection.
Owner:SUNYARD SYST ENG CO LTD

Multimodal speech synthesis detection method

This invention relates to the field of speech synthesis detection technology, and more particularly to a multimodal speech synthesis detection method, comprising: extracting key audio detection features and key video detection features of speech audio signals and video modal signals at different signal points; dividing the speech audio signals and video modal signals into multiple audio events and video events based on the key audio detection features and key video detection features, respectively; calculating the uncertainty between audio events and video events; constructing a cross-modal anchor point map and extracting the alignment relationship between anchor points; using a cross-modal speech synthesis detection model to detect the forgery probability of audio events in the anchor points and the forgery synthesis probability of the speech audio signals; thereby achieving joint identification of local audio event anomalies and global modal inconsistencies, thus accurately determining whether the speech signal is forged and synthesized, and improving the accuracy and robustness of speech synthesis detection.
Owner:XIAMEN UNIV