Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Speech reconstruction" patented technology

Speech reconstruction method and system based on entropy coding residual quantization and spectrum repair

ActiveCN121096348ASpeech recognitionFrequency spectrumSpeech reconstruction
The invention provides a voice reconstruction method and system based on entropy coding residual quantization and frequency spectrum restoration, and relates to the technical field of artificial intelligence voice signal processing, and the method comprises the steps: obtaining an original voice waveform, inputting the voice waveform into a neural voice coding and decoding model, firstly entering a coder to map the input voice waveform into acoustic potential representation, and then entering a frequency spectrum restoration model; performing residual quantization on the acoustic potential characterization layer by layer through a residual vector quantization module, introducing a gating-based dynamic layer number selection mechanism and entropy regularization constraint, enabling bits to be adaptively distributed among different voice segments, reconstructing reconstructed acoustic features of the potential characterization, inputting the reconstructed acoustic features into a decoder, restoring the reconstructed acoustic features into a time domain waveform, and outputting the time domain waveform. And mapping to a logarithmic magnitude spectrum domain through a spectrum repairing module, predicting a residual error in the logarithmic magnitude spectrum domain and performing confidence gating fusion to obtain a complex spectrum, and outputting after time domain synthesis to obtain reconstructed speech. According to the invention, high fidelity, intelligibility and transmission reliability of the voice can be considered at an extremely low bit rate.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Voice coding and decoding method, device, equipment and medium

ActiveCN121054006ASpeech analysisProduct quantizationDecoding methods
The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical science and technology, and discloses a voice coding and decoding method, device, equipment and medium. A continuous vector is obtained by encoding the data through an encoder trained by a first-stage mirror image architecture; segmenting a continuous vector into sub-vectors, matching the sub-vectors with corresponding sub-coding dictionaries through a quantizer to obtain sub-indexes, and combining the sub-indexes to generate an overall index; analyzing the overall index to obtain sub-indexes, calling sub-discrete vectors and splicing the sub-discrete vectors into a discrete vector; and reconstructing the discrete vector through a decoder trained by a second-stage non-mirror-image architecture to obtain a target speech spectrum feature, and converting the target speech spectrum feature into a target speech signal. According to the method, double-stage training is adopted, the first-stage mirror image architecture guarantees the coding stability, the second-stage non-mirror image architecture improves the decoding flexibility, and the product quantization technology is combined, so that the calculation efficiency and the storage overhead are balanced, and meanwhile, the voice reconstruction quality is improved.
Owner:平安科技(上海)有限公司

Anonymization privacy protection method and system for voice information retention

The embodiment of the invention provides an anonymization privacy protection method and system for voice information retention. The method comprises the following steps: extracting speaker embedding of an original audio, eliminating the tone of the speaker, and keeping semantic and rhythm speaker irrelevant features; embedding and inputting a speaker into a speaker anonymous module matched with a three-stage stream based on a U-Net architecture to obtain anonymous embedding; and combining irrelevant features of the speaker with anonymous embedding by using a pre-trained voice reconstruction model to generate anonymized voice with tone privacy. According to the embodiment of the invention, voice anonymization facing content privacy and tone privacy reserved by voice information is realized, the effectiveness of the anonymized voice generated by using the method in a downstream task is superior to that of a baseline model, meanwhile, the privacy of a speaker is also guaranteed, and safe use of data is realized.
Owner:SHANGHAI JIAOTONG UNIV

Speech signal processing method and device, electronic equipment and storage medium

The present disclosure relates to a speech signal processing method and device, electronic equipment and storage medium, and relates to the technical field of audio. The method comprises: acquiring a multi-channel far-field speech signal collected for a target sound source; inputting the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing according to a prediction result of a first neural network model for the multi-channel far-field speech signal; and inputting the target speech spectrum information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal. The technical scheme provided by the embodiments of the present disclosure can realize the transformation of speech hearing sensation and improve the quality and intelligibility of speech.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Lightweight semantic preserving coding and decoding and voice signal reconstruction method

PendingCN121905195ASpeech analysisNeural learning methodsNerve networkSpeech reconstruction
The invention discloses a lightweight semantic preserving coding and decoding and voice signal reconstruction method, which comprises the following steps of: acquiring an analog audio signal from an environment, inputting an original audio signal and an audio signal subjected to high-frequency filtering to two ends of a comparator, outputting an ultralow-bit-rate binary bit stream and transmitting the ultralow-bit-rate binary bit stream to a receiving end, and after the receiving end receives the signal, outputting the ultralow-bit-rate binary bit stream to the voice signal reconstruction end. The method comprises the following steps: inputting a voice signal to an end-to-end U-Net convolutional neural network of a bottleneck layer integrated LSTM module, outputting a reconstructed voice signal, training the network by using voice coding sample reconstruction loss and multi-scale short-time Fourier transform, and performing voice signal reconstruction on an ultra-low bit rate binary bit stream of a to-be-reconstructed audio after training. The technical contradiction that in resource-limited remote audio collection and distributed sensing tasks, node hardware resources are extremely limited, and ultra-low bit rate signal compression distortion causes that node microminiaturization deployment and high voice reconstruction quality cannot be considered at the same time is solved.
Owner:ZHEJIANG UNIV

Intelligent plotting method based on voice input

PendingCN121366571ASpeech recognitionIntelligibility (communication)Speech reconstruction
The invention relates to the technical field of intelligent plotting, and provides an intelligent plotting method based on voice input. The intelligent plotting method based on voice input comprises the following steps: S1, receiving an original audio signal containing a voice instruction input by a user in real time at a terminal equipment side; according to the intelligent plotting method based on voice input, the problem of communication interruption is solved, the problem of voice recognition under high-strength physical vibration is solved through an innovative voice reconstruction technology, full-domain and full-condition adaptation in the true sense is achieved, vibration distortion reconstruction is preferentially carried out on the voice signals, and therefore the voice recognition efficiency is improved. The method has the advantages that the jittering voice signals are corrected before the jittering voice signals are sent to the recognition engine, so that the voice definition and intelligibility are guaranteed from the information source, the accuracy of all subsequent processing links is greatly improved, and the method is particularly suitable for key instructions which are issued in high-pressure and high-risk environments and have success or failure of events and even life safety. And the reliability guarantee is provided.
Owner:BOZHI NETWORK SECURITY (NANJING) TECH CO LTD

Speech generation method and device based on semantic adjustment, equipment and medium

PendingCN122313946ASemantic representationSpeech reconstruction
This invention relates to the field of speech synthesis technology, and discloses a speech generation method, apparatus, device, and medium based on semantic adjustment. The method includes: encoding text content and style cues to obtain text semantic representation and style semantic representation; performing semantic consistency analysis based on the two to obtain semantic mismatch results; determining dynamic guidance strength based on the semantic mismatch results, and performing autoregressive generation processing in conjunction with alternative style cues to obtain an acoustically labeled sequence; and reconstructing the acoustically labeled sequence to obtain the target speech. This invention can be applied to business scenarios such as fintech and healthcare. By obtaining semantic mismatch results through semantic consistency analysis and determining dynamic guidance strength based on the semantic mismatch results, the speech generation process can be adjusted according to changes in semantic consistency, thereby reducing the deviation caused by inconsistencies between semantic and emotional expression and improving the expressive coordination of the target speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

Cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion perception

The application discloses a cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion sensing, and belongs to the technical field of human-computer interaction and wearable computing. The method uses a micro-pressure sensing unit placed in an in-ear earphone to collect a non-acoustic air pressure sequence caused by the movement of a sound-producing organ; through adaptive baseline drift suppression and rhythm perception data enhancement processing, a robust feature space is constructed; further, an end-to-end deep neural network containing domain adversarial adaptation, cross-modal semantic alignment, coarse-grained mel-spectrogram generation and residual detail correction is used to map the TPVS to a high-fidelity acoustic mel spectrum. The application effectively breaks through the technical bottleneck of the lack of high-frequency acoustic features in low-frequency mechanical signals, realizes high-precision silent speech command analysis in a mobile and noisy scene, and introduces a coupled quality evaluation gate and trigger-based start / stop control at the inference end to suppress invalid inference and reduce power consumption when wearing is poor or there is no trigger condition.
Owner:DONGHUA UNIV

Balanced speech reconstruction and noise suppression using dual-asymmetric loss

PendingUS20260253599A1Speech reconstructionNoise
An audio processor can include a circuit configured to provide a digital signal representative of a sound, and a digital signal processor configured to process the digital signal. The digital signal processor can be configured to process an enhancement model and including a loss function having a first penalty term selected to control sound reconstruction, and a second penalty term selected to control noise suppression. The first and second penalty terms can be weighted by an adjustable parameter to allow the enhancement model to be tuned for quality of sound or performance of noise suppression.
Owner:SKYWORKS SOLUTIONS INC

Voice conversion method, training method, device, equipment and medium

ActiveCN115565520BSpeech synthesisSpeech reconstructionVoice transformation
The voice conversion method, the training method, the device, the equipment and the medium provided by the embodiments of the present application obtain the source linear spectrum of the source speaker according to the source voice data of the source speaker; input the source linear spectrum into a pre-trained voice coding model to output corresponding spectral feature prediction data; input the spectral feature prediction data and target speaker feature data of a target speaker into a voice reconstruction model to output corresponding target voice data; in the above manner, the content information of the source voice data and the speaker feature are decoupled, and in the training stage and the voice conversion stage of the voice coding model, the input and the output of the voice coding model respectively only contain content information, the voice coding model is used to reconstruct voice data together with the content information and the speaker feature through the voice reconstruction model, which is beneficial to improve the training speed and the training effect of the voice coding model, and further improves the voice conversion effect.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech extraction methods, devices, equipment and media

ActiveCN119993130BSpeech recognitionSequence reconstructionSpeech reconstruction
This invention relates to the field of artificial intelligence technology and discloses a speech extraction method, apparatus, device, and medium. The method includes: first, acquiring reference speech of the target speaker and mixed speech of all speakers; preprocessing and encoding the reference speech and mixed speech to generate two discrete token sequences; fusing the two discrete token sequences to form a fused discrete token sequence; using a language model to predict the fused discrete token sequence to generate candidate discrete token sequences for the target speaker; calculating the probability distribution of the candidate token sequences using a linear classifier and selecting the sequence with the highest probability as the target discrete token sequence; and then reconstructing the target discrete token sequence into a speech waveform to obtain the speech of the target speaker. This invention transforms the complex audio generation problem into a classification problem, simplifying model training; and utilizes the sequence modeling capability of a language model to capture long-term dependencies between speech tokens, achieving high-quality speech reconstruction.
Owner:PING AN TECH (SHENZHEN) CO LTD

Millimeter wave radar voice reconstruction and recognition method based on physical guidance network

A millimeter-wave radar voice reconstruction and recognition method based on a physical guide network comprises the steps that a millimeter-wave radar is used for transmitting a radio-frequency signal to a to-be-detected target and receiving an echo signal, and meanwhile a reference audio signal is collected; extracting a steady-phase signal Mel spectrum according to the echo signal; generating a simulated radar Mel spectrum by performing an audio signal simulation on the reference audio signal and the common speech data set; synchronizing and standardizing the stable-phase signal Mel spectrum and the simulated radar Mel spectrum, and constructing a voice signal data set; constructing a multi-mode voice reconstruction network model; training a multi-modal voice reconstruction network model according to the voice signal data set; inputting a newly collected real millimeter wave radar signal into the trained multi-mode voice reconstruction network model for voice reconstruction, and outputting a non-contact voice Mel-frequency spectrogram; and inputting the voice Mel spectrogram into the constructed lightweight convolutional neural network classifier, and outputting an identity category label of the speaker.
Owner:HANGZHOU DIANZI UNIV

A method, apparatus, device and medium for training a speech generation model

The application belongs to the field of artificial intelligence, and relates to a training method of a speech generation model, comprising the following steps: obtaining reference timbre spectrum, phoneme information and speech spectrum of a target object; training a preset initial speech generation model based on the reference timbre spectrum, the phoneme information and the speech spectrum to obtain model parameters; and adjusting parameters of a multi-timbre feature extraction network, a phoneme feature extraction network, a prosody feature discretization network, a time sequence alignment module, an attention fusion module and a speech reconstruction decoding network of the initial speech generation model based on the model parameters to construct the speech generation model. The application also provides an apparatus, a device and a medium. In addition, the application also relates to blockchain technology, and speech training data and model parameters can be stored in a blockchain. The application can realize decoupling of timbre and prosody information, and flexibly adjust the timbre and prosody information to generate synthesized speech with diversity and flexibility.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech synthesis method, speech synthesis device, electronic equipment and storage medium

The invention provides a speech synthesis method, a speech synthesis device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, and is suitable for the financial field and the medical field. The method comprises the steps of obtaining a target text, and determining a target language to which the target text belongs; acquiring a sample voice set and a sample text belonging to the target language; performing semantic modeling on the sample text and the target language through an initial semantic modeling device to obtain a first sample semantic unit sequence; performing voice quantization on the sample voice to obtain a second sample semantic unit sequence; performing parameter adjustment on the initial semantic modeler according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain a target semantic modeler; performing semantic modeling on the target text and the target language through a target semantic modeling device to obtain a text semantic unit sequence; and performing voice reconstruction according to the text semantic unit sequence to obtain a target synthetic voice. According to the invention, the accuracy of cross-language speech synthesis can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation

PendingCN121354580ASpeech analysisHigh level techniquesComputation complexitySpeech reconstruction
The invention discloses a real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation, and relates to a real-time single-microphone voice noise reduction algorithm. The invention aims to solve the problem that conversation voice is buried by noise due to noisy and diverse background noise of an interphone in a special communication scene. Noise power spectrum estimation is optimized by fusing a continuous minimum value tracking algorithm, a speech enhancement residual error is introduced as a real noise approximate value to participate in a recursive average process, and the response speed and accuracy of noise estimation are improved; a gain function is calculated in combination with an optimal correction logarithm MMSE estimator, and a closed expression approximates exponential integration to reduce the calculation complexity. The algorithm specifically comprises the steps of preprocessing and framing, noise power spectrum estimation, gain function calculation, voice reconstruction and the like, the segmentation signal-to-noise ratio is remarkably improved on the premise that low voice distortion is guaranteed, and the real-time processing requirement is met. The invention belongs to the technical field of voice signal processing.
Owner:HARBIN INST OF TECH

Speech conversion model training methods, speech conversion methods, devices and media

This application relates to the field of speech conversion technology, and provides a speech conversion model training method, speech conversion method, apparatus, and medium. The method includes: extracting speech sample features from preset speech samples using an encoder; then decoupling the speech samples based on a preset masking strategy to obtain sample feature representations; inputting the sample feature representations into a generator to reconstruct the Mel spectrogram of the speech samples based on the sample feature representations to obtain the Mel spectrogram of the target sample; calculating the speech reconstruction loss of the speech conversion model based on the Mel spectrogram of the target sample and the original sample Mel spectrograms corresponding to the preset speech samples; and optimizing the parameters in the speech conversion model based on adversarial loss and speech reconstruction loss to obtain a trained speech conversion model. By decoupling the speech sample features through a preset masking strategy and a preset adversarial network, the robustness of the speech conversion model is improved, thereby increasing training efficiency.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice signal generation method and system based on millimeter wave radar

PendingCN121214958ASpeech analysisBiological modelsFrequency spectrumSpeech reconstruction
The invention discloses a voice signal generation method and system based on a millimeter wave radar, and belongs to the technical field of Internet of Things sensing and voice signal processing. The method comprises the following steps: designing a millimeter-wave radar fine-grained vibration signal acquisition model; obtaining noisy material vibration information based on the model, and constructing an end-to-end voice reconstruction model by taking a spectrogram of the noisy material as input; performing data enhancement by synthesizing the simulated noisy millimeter wave spectrogram; noisy vibration data captured by a radar is input, and a priori-free and unconstrained recognizable voice is reconstructed by using a deep neural network. The system correspondingly comprises a design module, a construction module, an enhancement module and an output module. The method breaks through the environmental limitation of traditional voice acquisition, improves the model robustness and the voice generation quality, and is suitable for the voice generation demand in a non-contact and complex interference scene.
Owner:XI AN JIAOTONG UNIV

Speech reconstruction method and device based on real-time training, computer device and medium

The application relates to the technical field of artificial intelligence, in particular to a speech reconstruction method and device based on real-time training, computer equipment and a medium. The method divides training speech into a first speech segment and a second speech segment, respectively inputs the first speech segment and the second speech segment into a feature encoder to obtain first speech features and second speech features, inputs the training speech into a filter encoder to obtain filter features, inputs the mean of the first speech features and the second speech features and the filter features into a decoder to obtain reconstructed speech, trains a speech reconstruction model according to the first speech features, the second speech features, the reconstructed speech and the training speech, reconstructs to-be-processed speech based on the trained speech reconstruction model, improves the accuracy of feature decoding, and further improves the accuracy of speech reconstruction, can improve the simulation of machine customer service speech under a financial service platform, and further improves the user experience of users in the financial service platform.
Owner:PING AN TECH (SHENZHEN) CO LTD

Low-complexity speech enhancement method based on parallel GRU-convolutional neural network

PendingCN121483274ASpeech analysisBiological modelsSpeech reconstructionNoise
The invention discloses a low-complexity speech enhancement method based on a parallel GRU-convolutional neural network, and the method comprises the steps: synthesizing pure speech and noise data, and generating a mixed speech data sample with noise; processing through a Mel filter to obtain Mel logarithmic energy spectrum data characteristics of the mixed voice data; inputting a parallel GRU-convolutional neural network for training to obtain an amplitude value mask of the mixed voice data with noise; performing amplitude estimation according to the amplitude value mask and the amplitude spectrum of the mixed voice data to obtain the amplitude value of the enhanced target voice data; and performing voice reconstruction according to the amplitude value and the phase angle of the enhanced target voice data to obtain enhanced voice data. The method has low complexity and certain speech enhancement performance, and can be deployed in most low-cost processing chips in the market.
Owner:ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE

Speech reconstruction method and device based on intracranial neural electrical signals and electronic equipment

PendingCN122392482ASpeech reconstructionEngineering
The present application relates to a kind of speech reconstruction methods, device and electronic equipment based on intracranial nerve electric signal, the method comprises: the neural feature time sequence stream corresponding to the intracranial nerve electric signal continuously collected by the language-related brain area of subject is acquired;Neural feature time sequence stream is input to a flow causal decoding model, to generate corresponding speech acoustic feature time sequence stream in real time;Wherein, flow causal decoding model is the causal deep neural network model trained based on sample data, and sample data includes sample neural signal feature sequence and its corresponding sample speech acoustic feature sequence;Speech acoustic feature time sequence stream is integrated into speech waveform stream and output.The present application realizes real-time speech synthesis of millisecond level delay by flow causal decoding model and dynamic block strategy, solves the high delay problem of existing scheme, simultaneously avoids auditory feedback interference by causal mask mechanism, can adapt to the application demand in silent scene.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV

Multi-object vibration fusion voice perception and reconstruction method based on millimeter wave radar

The invention relates to a multi-object vibration fusion voice perception and reconstruction method based on a millimeter-wave radar. The method comprises the following steps: S1, carrying out range frequency conversion on millimeter-wave radar receiving signals to obtain a plurality of discrete object vibration signals; s2, obtaining a frequency response function of each object by using the trained feature extraction network based on object vibration signals; s3, fusing the object vibration signal of each object and the corresponding frequency response function along a frequency axis to obtain a first spectrogram of each object; s4, performing cross-channel frequency attention matching on the first spectrograms of the objects to realize dynamic alignment and feature complementation so as to obtain a fused spectrogram; and S5, based on the fused spectrogram, obtaining a speech spectrum by using the trained speech reconstruction model. Compared with the prior art, the method has the advantages that the inherent frequency response difference of an object is utilized, the physical compensation of acoustic propagation characteristics is realized, and the voice reconstruction quality is improved.
Owner:SOUTHEAST UNIV

Laryngectomy postoperative speech reconstruction method and system based on latent space editing

PendingCN122511225ALaryngectomyPersonalization
The present application relates to a kind of laryngectomy post-voice reconstruction method and system based on hidden space editing, low-frequency body surface vibration signal and high-frequency air-conducted acoustic signal produced simultaneously when patient vocalization is carried out, and dual-channel heterogeneous signal acquisition and fusion are carried out;Based on soft feature mapping and hidden space editing, the signal after fusion is reconstructed to output correction text;Based on reference audio and correction text, personalized waveform file is synthesized and played back.By retaining soft features and intervening in hidden space, the probability distribution is retained, the decision-making power is moved backward, real-time error correction during inference can be achieved, and high-fidelity intent reconstruction can be achieved.Through dual-channel sampling, the low-frequency body surface vibration signal and high-frequency air-conducted acoustic signal collected are complementary in frequency domain, and the complete acoustic information is reconstructed by combining them, which improves the noise immunity.Based on reference audio, personalized speech files can be synthesized without samples, noise immunity and personalization can coexist, and important technical support is provided for the development of intelligent medical treatment.
Owner:PEKING UNIV

Speech coding method, apparatus, device, and medium

ActiveCN121054006BSpeech analysisProduct quantizationDecoding methods
This invention relates to the field of artificial intelligence technology and can be applied to fintech and medical technology. It discloses a speech encoding and decoding method, apparatus, device, and medium. The method includes: acquiring an original speech signal and extracting its spectral features; encoding the signal using an encoder trained with a first-stage mirror architecture to obtain a continuous vector; segmenting the continuous vector into sub-vectors, matching them with corresponding sub-encoding dictionaries using a quantizer to obtain sub-indices, and combining them to generate a global index; parsing the global index to obtain sub-indexes, retrieving sub-discrete vectors and concatenating them into a discrete vector; reconstructing the target speech spectral features from the discrete vectors using a decoder trained with a second-stage non-mirror architecture, and then converting the target speech signal. This invention employs a two-stage training process: the first-stage mirror architecture ensures encoding stability, while the second-stage non-mirror architecture enhances decoding flexibility. Combined with product quantization technology, it improves speech reconstruction quality while balancing computational efficiency and storage overhead.
Owner:平安科技(上海)有限公司

Speech reconstruction system for multimedia files

PendingEP4670153A1Speech recognitionSpeech synthesisSpeech reconstructionAcoustics
A speech recognition system may determine speech in the presence of multiple, different forms of corrupted audio. The system may obtain audio-visual data including visual data associated with a person and audio data associated with the person. The system may also determine, based on the visual data, pronunciation data associated with speech by the person. The system may also convert the speech to encoded data. The system may also synthesize, based on the encoded data, the speech to obtain synthesized speech.
Owner:META PLATFORMS TECHNOLOGIES LLC

Voice signal processing method and device, electronic equipment and storage medium

PendingCN121862131ATo achieve the purpose of voice attribute decompositionimprove intelligibilitySpeech synthesisSpeech reconstructionSpeech sound
The invention relates to a voice signal processing method and device, electronic equipment and a storage medium. The method comprises the steps that content coding features, rhythm coding features and timbre coding features are acquired, the content coding features are used for representing the content of a voice signal, the rhythm coding features are used for representing the rhythm of the voice signal, and the timbre coding features are used for representing the timbre of the voice signal; content context construction is carried out on the content coding features based on the tone coding features and the rhythm coding features, and content context features are obtained; and performing voice reconstruction according to the timbre coding feature, the rhythm coding feature and the content context feature to obtain a target reconstructed voice signal. According to the technical scheme provided by the invention, high-quality speech reconstruction can be provided at a low code rate.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Communication methods, devices and systems based on voiceprint recognition and speech reconstruction

This invention discloses a communication method, apparatus, and system based on voiceprint recognition and speech reconstruction, relating to the field of communication security technology. The communication apparatus based on voiceprint recognition and speech reconstruction includes a voice acquisition module, a storage module, a voice conversion module, a data transmission module, a data reception module, a speech synthesis module, and a voice playback module, all connected to a central processing module. The voice acquisition module is connected to the storage module, the voice conversion module is connected to the speech synthesis module and the data reception module, and the speech synthesis module is connected to the voice playback module. The communication method of this invention includes: S1. Voiceprint extraction, S2. Speech-to-text conversion, S3. Data transmission, S4. Data reception, and S5. Text-to-speech conversion and playback. This invention solves the problems of high risk of voiceprint biometric leakage and difficulty in balancing natural call quality and security in existing technologies. It achieves encryption while preventing voiceprint feature transmission, maintaining a natural call experience, and significantly reducing bandwidth requirements due to the small data volume.
Owner:深圳市南山区明涵软硬件技术服务工作室

Voice reconstruction method and system based on gating re-calibration and route weighting

The invention discloses a voice reconstruction method and system based on gating re-calibration and routing weighting, and the method comprises the steps: introducing an intra-group channel gating re-calibration module at a coding end of a neural vocoder model, carrying out the self-adaptive re-calibration of coding features through a channel grouping and gating mechanism, and improving the effectiveness and quantification efficiency of feature expression. Meanwhile, a routing network is introduced in a neural vocoder model training stage, probability distribution is output based on coding characteristics, and self-adaptive weighted reconstruction loss for a channel segment of the Mel filter bank is constructed, so that the model emphasizes on optimizing and sensing a reconstruction error of the channel segment of the Mel filter bank corresponding to a key frequency region under a low bit rate. The routing network only participates in loss calculation and parameter updating in a training stage, and is removed in a reasoning stage, so that additional routing reasoning calculation overhead is not introduced. The method can improve the perception quality and stability of voice reconstruction at an extremely low bit rate, and is suitable for bandwidth-limited scenes such as satellite communication and short-wave communication.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Helium voice key information extraction method based on scarce sample

PendingCN121725800ASpeech analysisBiological modelsFrequency spectrumIntelligibility (communication)
The invention provides a helium voice key information extraction method based on scarce samples, and belongs to the technical field of voice processing and artificial intelligence. The technical problems of frequency spectrum distortion, formant upward shift and tone distortion caused by sound velocity change of voice signals in saturated diving and other high-pressure helium-oxygen environments are solved. According to the technical scheme, the method comprises the following steps: S1, introducing a target tone enhancement and leakage prevention mechanism behind a content coding module; s2, a self-adaptive generation network structure suitable for the correction task is constructed in a waveform generation module based on EVA-GAN; and S3, based on a ContentVec fine tuning strategy of self-supervised learning, introducing a mask prediction task on unlabeled helium voice data. According to the method, high-fidelity voice reconstruction under a complex acoustic condition is realized, so that the intelligibility, naturalness and communication reliability of the voice are remarkably improved.
Owner:NANTONG UNIV

Primary-secondary network speech enhancement system with fusion attention mechanism

This invention belongs to the field of signal and information processing technology, specifically relating to a master-slave network speech enhancement system incorporating an attention mechanism. It includes a master network, a slave network, and a speech reconstruction module. The master network comprises a first feature extraction module and a second feature extraction module. The first feature extraction module performs convolution and feature mapping on the speech input signal to obtain a first feature. The second feature extraction module processes the input signal through a three-layer bidirectional gated recurrent unit (BiGRU) and a multi-head attention mechanism to obtain a second feature. The first and second features are concatenated to form the master network output feature. The slave network concatenates the master network output feature with the input signal to obtain a third feature, which is then sent to a first CNN network to calculate the slave network output feature. The speech reconstruction module adds the master network output feature and the slave network output feature and reconstructs the enhanced speech. This invention can improve speech enhancement capabilities.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Speech reconstruction method and system based on entropy coding residual quantization and spectral repair

ActiveCN121096348BSpeech recognitionFrequency spectrumSpeech reconstruction
The present disclosure provides a speech reconstruction method and system based on entropy coding residual quantization and spectrum repair, relating to the technical field of artificial intelligence speech signal processing, comprising: obtaining an original speech waveform, inputting the speech waveform into a neural speech coding and decoding model, first entering the encoder to map the input speech waveform into an acoustic latent representation, and then performing residual quantization layer by layer through a residual vector quantization module, introducing a dynamic layer number selection mechanism based on gating and an entropy regularization constraint, so that the bits are adaptively allocated between different speech segments, reconstructing the reconstructed acoustic features of the latent representation, inputting the reconstructed acoustic features into the decoder to restore the time domain waveform, and then mapping to the log amplitude spectrum domain through the spectrum repair module, predicting the residual in the log amplitude spectrum domain and performing confidence gating fusion to obtain a complex spectrum, and then outputting the reconstructed speech after time domain synthesis. The present disclosure can balance the high fidelity, intelligibility and reliability of transmission of speech at very low code rate.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1