Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

101 results about "Speech spectrum" patented technology

Speech spectrum - the average sound spectrum for the human voice. acoustic spectrum, sound spectrum - the distribution of energy as a function of frequency for a particular sound source.

Classroom behavior analysis system based on adaptive algorithm

The invention relates to the technical field of classroom behavior analysis, and discloses a classroom behavior analysis system based on an adaptive algorithm, which comprises a detection module, a parameter extraction module, a feature integration module and a decision generation module. The detection module collects classroom multi-modal behavior data and identifies active student terminals; the parameter extraction module obtains a real-time posture vector, a voice spectrum feature, an interaction response time delay and a teaching content semantic tag of a target student; the feature integration module performs cross-modal fusion on the data to generate attitude, voice, interactive fusion features and semantic association features; and the decision generation module performs hierarchical association modeling and dynamic priority ranking through a multi-head attention mechanism, generates a behavior matching degree and determines a target analysis object. The system realizes multi-modal data fusion by means of algorithms such as space-time decomposition and dynamic density clustering, improves the accuracy and real-time performance of classroom behavior analysis, and provides support for teaching optimization.
Owner:JIAN YINGJIA ELECTRONICS TECH

Speech recognition method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence, and particularly provides an artificial intelligence-based speech recognition method, which comprises the following steps of: acquiring multiple paths of speech signals through a microphone array; performing noise reduction processing on each path of voice signal, removing background noise, enhancing the voice signal by adopting an adaptive beam forming algorithm, and extracting a target voice signal; performing short-time Fourier transform on the preprocessed voice signals, extracting voice spectrum features, and extracting high-level semantic features of the voice through a deep learning model; inputting the extracted speech features into a speech recognition model based on an attention mechanism, generating text transcription of target speech, dynamically adjusting model parameters through an adaptive learning module, and optimizing a recognition result; according to the method, the recognition result is corrected according to the real-time feedback of the user, the corrected data is used for online updating of the model, the robustness and adaptability of the system are improved, and the method has the effects that training and evaluation are conducted through high-quality data, and delay is reduced in real-time processing.
Owner:BEIJING HURRICANE SOFTWARE CO LTD

Intelligent financial customer service interaction optimization method and system

The invention relates to the technical field of financial science and technology, and discloses an intelligent financial customer service interaction optimization method and system, and the method comprises the steps: collecting a real-time session data flow and a session path when a target customer carries out the session interaction with an intelligent customer service system, calculating a satisfaction degree attenuation value of the intelligent customer service system, and evaluating an interaction quality index of the session path; extracting speech spectrum features and text semantic vectors in the real-time session data stream, drawing an emotion fluctuation curve of the target customer in the session process, extracting a wave crest semantic fragment corresponding to an emotion wave crest in the emotion fluctuation curve, and calculating an emotion entropy value of the target customer; calculating the service integrating degree of the intelligent customer service system in the interaction channel; generating a session strategy optimization scheme of the target customer; and executing interaction processing about the target client, calculating a client journey conversion rate corresponding to the session strategy optimization scheme, and performing optimization adjustment on the session strategy optimization scheme to obtain a final session strategy. According to the invention, the accuracy of financial customer service interaction optimization can be improved.
Owner:TIBET DOLPHIN INFORMATION TECHNOLOGY CO LTD

Key information extraction method and system based on multi-modal large model

The invention discloses a key information extraction method and system based on a multi-modal large model, and belongs to the technical field of information processing.The key information extraction method and system based on the multi-modal large model comprise the following specific steps that firstly, input heterogeneous modal data are subjected to standardized preprocessing, and the preprocessed data are subjected to data preprocessing; comprising text vectorization, image feature coding and speech spectrum analysis; 2, establishing an inter-modal incidence matrix through a cross-modal attention mechanism, and dynamically adjusting the feature weight of each modal; and 3, sequentially executing context semantic understanding, entity relationship modeling and core information positioning by adopting a hierarchical feature extraction architecture. According to the method, intelligent selection of inter-modal features is realized through dynamic weight distribution, the key information recall rate is greatly improved in a noise interference scene, and after a feature distillation loss function is introduced and redundant features are eliminated, the error propagation rate is effectively reduced in a noisy voice and low-resolution image coexistence scene.
Owner:SHANGHAI XIAOGONGYI E-COMMERCE CO LTD

Information processing method and system for voice conversion

The invention relates to the technical field of voice signal processing and synthesis, and particularly discloses an information processing method and system for voice conversion, and the method comprises the steps: extracting a language feature vector of each word from an input text; calculating posterior probabilities of the words in different languages by combining a Bayesian reasoning mechanism, and generating language attribution confidence coefficient characteristic values; a Monte Carlo sampling method is adopted to carry out multiple times of context sensitive simulation, and pronunciation path selection probability distribution characteristic values are generated; further fusing the feature values into a multi-language pronunciation decision vector, inputting the multi-language pronunciation decision vector into a multi-language end-to-end speech synthesis model, calling a phoneme mapping rule of a corresponding language and an acoustic parameter prediction module, and generating a high-quality target speech spectrogram; and finally, dynamically adjusting Bayesian prior distribution and a language recognition threshold according to the output speech spectrogram.
Owner:SHANDONG POLYTECHNIC COLLEGE

Electric power emerging business risk keyword identification and early warning method, system and device based on NLP technology and medium

The invention discloses an electric power emerging business risk keyword identification and early warning method, system and device based on an NLP technology and a medium, and belongs to the technical field of electric power risk, and the method comprises the steps: obtaining multi-source customer service text data, and extracting risk keywords and corresponding voice spectrum data from the multi-source customer service text data; inputting the risk keyword into a power service knowledge graph, and generating a first risk research and judgment result through knowledge reasoning, the first risk research and judgment result comprising a risk traceability path; performing acoustic feature analysis on the speech spectrum data, extracting acoustic feature parameters, calculating an emotional excitation index, and generating a second risk identification result; performing weighted fusion on the first risk research and judgment result and the second risk identification result to generate a comprehensive risk score; and when the comprehensive risk score reaches a preset threshold value, generating a structured early warning work order, and carrying out data binding on the structured early warning work order, the original customer service work order and the call record. According to the invention, the efficiency of risk management and control is improved.
Owner:GUIZHOU POWER GRID CO LTD

A formant extraction method for continuous speech based on peak selection

ActiveCN115064180BSpeech analysisFrequency spectrumFormant
The present invention discloses a continuous speech formant extraction method based on peak selection, comprising: performing a preprocessing operation on a single frame of input speech; using a linear prediction method to preliminarily estimate the peak value in the spectral envelope of the speech frame; establishing a reference point and a formant trough, and then using a peak selection method to establish a mapping relationship between the peak value and the reference point; using the mapping relationship between the peak value and the reference point and the formant trough to determine the formant of the speech frame; and performing formant estimation on the continuous speech: dividing the continuous speech into frames according to different frame numbers, using the above algorithm to loop 100 times to obtain the formant parameters under different frame number tests, averaging the results after 100 loops, and obtaining the final result after smoothing. The method of the present invention can eliminate the influence of merged peaks and false peaks, and has a fast convergence speed and strong robustness.
Owner:NANJING UNIV OF POSTS & TELECOMM

Voice coding and decoding method, device, equipment and medium

ActiveCN121054006ASpeech analysisProduct quantizationDecoding methods
The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical science and technology, and discloses a voice coding and decoding method, device, equipment and medium. A continuous vector is obtained by encoding the data through an encoder trained by a first-stage mirror image architecture; segmenting a continuous vector into sub-vectors, matching the sub-vectors with corresponding sub-coding dictionaries through a quantizer to obtain sub-indexes, and combining the sub-indexes to generate an overall index; analyzing the overall index to obtain sub-indexes, calling sub-discrete vectors and splicing the sub-discrete vectors into a discrete vector; and reconstructing the discrete vector through a decoder trained by a second-stage non-mirror-image architecture to obtain a target speech spectrum feature, and converting the target speech spectrum feature into a target speech signal. According to the method, double-stage training is adopted, the first-stage mirror image architecture guarantees the coding stability, the second-stage non-mirror image architecture improves the decoding flexibility, and the product quantization technology is combined, so that the calculation efficiency and the storage overhead are balanced, and meanwhile, the voice reconstruction quality is improved.
Owner:平安科技(上海)有限公司

Voice enhancement method based on AI echo cancellation

The invention discloses a speech enhancement method based on AI echo cancellation, and relates to the technical field of speech processing. According to the voice enhancement method based on AI echo cancellation, voice data sets in different echo environments are obtained through a recording device, time-frequency conversion is carried out on voice signals, and voice spectrum features are extracted. Time-frequency features are extracted through a convolutional neural network, and an echo cancellation model is constructed by combining a time sequence dependency relationship of a long and short term memory network capture signal. And according to the echo type label and the echo cancellation model, target echo signal characteristics are extracted, and echo-removed signal estimation is generated in combination with an adaptive filter. According to the method, the echo cancellation effect index is calculated by comprehensively analyzing the signal quality and echo residual characteristics after the echo is removed, the filter parameters are dynamically adjusted, voice enhancement is further carried out on the signal after the echo is removed through the deep convolution auto-encoder, and dual improvement of echo cancellation and voice enhancement is achieved.
Owner:SHENZHEN ZHILIAN TECH CO LTD

Audio feature extraction method based on cranial nerve model

The invention discloses an audio feature extraction method based on a cranial nerve model, and the method comprises the steps: obtaining voice signal data, carrying out the preprocessing of a voice signal, and segmenting the voice signal into N voice segments; extracting acoustic features representing emotional conditions from each voice segment, normalizing the acoustic features, and synthesizing the acoustic features into a multi-channel voice spectrogram; and inputting the multi-channel speech spectrogram into the RBA-FE model, outputting a BDI score, carrying out depression degree prediction classification according to the BDI score, adjusting parameters of each network layer in the RBA-FE model according to the difference between prediction classification and actual classification, and obtaining an optimal RBA-FE model to carry out depression degree detection. According to the method, the cell selectivity in the auditory cortex of the human brain is simulated, the problems of overfitting and non-noise resistance of the LSTM are solved by adopting a self-adaptive activation mechanism, and precise recognition of the depression voice under different noise interferences is realized.
Owner:SOUTH CHINA UNIV OF TECH +1

Voice signal processing method and device, electronic equipment and readable storage medium

The invention relates to a voice signal processing method and device, electronic equipment and a readable storage medium. Under the condition that a part of sound signals of a user object are collected, corresponding intermediate feature representation is extracted from the part of sound signals, supplementary sound features corresponding to the user object are determined, synthesis processing is performed on the intermediate feature representation and the supplementary sound features, and a complete speech spectrum corresponding to the part of sound signals is obtained. And carrying out acoustic code conversion on the complete speech spectrum to obtain a converted audio signal. Since the intermediate feature representation is text content extracted from a part of sound signals, and the supplementary sound features describe the sound characteristics and styles of the user object, the intermediate feature representation of the part of sound signals of the user object and the supplementary sound features of the user object are synthesized, so that the sound characteristics and styles of the user object are obtained. Therefore, the finally obtained audio signal is not interfered by environmental noise, and the quality of the audio signal can be ensured.
Owner:LUXSHARE PRECISION TECH(NANJING) CO LTD

Speech spectrum reconstruction method and system combining time domain half-wave rectification and weighted Gaussian mixture model decoder, terminal and medium

The invention discloses a speech spectrum reconstruction method and system combining time-domain half-wave rectification and a weighted Gaussian mixture model decoder, a terminal and a medium, and relates to the technical field of speech signal processing, the method comprises the following steps: acquiring a low-resolution audio, executing time-domain half-wave rectification, and executing weighted Gaussian mixture model decoding; obtaining a mixed amplitude spectrum based on the amplitude spectrum of the low-resolution audio and the amplitude spectrum of the rectification signal; outputting GRU features based on an encoder, inputting the GRU features into a weighted Gaussian mixture model decoder, and generating a plurality of groups of Gaussian component parameters through three parallel linear layers and applying three constraint designs; and calculating a frame-level frequency distribution weight, combining the frame-level frequency distribution weight with the mixed amplitude spectrum to obtain a model output amplitude spectrum, inputting the model output amplitude spectrum and the amplitude spectrum of the low-resolution audio into a frequency band guide masking module to obtain a final amplitude spectrum, and reconstructing an audio time domain waveform based on the final amplitude spectrum and the phase of the low-resolution audio. According to the method, high-fidelity and high-frequency reconstruction can be realized while the complexity of the model is greatly reduced.
Owner:ELEVOC TECH CO LTD

Artificial intelligence psychological assessment method and device based on multiple modes

The invention discloses an artificial intelligence psychological assessment method and equipment based on multiple modes, particularly relates to the technical field of psychological assessment analysis, and aims to solve the problem of high detection error rate caused by biological characteristic distortion due to privacy desensitization in existing psychological assessment. The method comprises the following steps of: dynamically generating an evaluation instruction matched with a user behavior, collecting a facial micro-expression video, segmenting and quantifying a privacy interference level based on region pixels, and fuzzifying a non-sensitive region while keeping key micro-expression region characteristics; synchronously acquiring a voice spectrum waveform and reconstructing voiceprint tone modification and fundamental frequency waveform; extracting text sentiment keywords to construct a cross-stage sentiment association chain to repair semantic fracture caused by local desensitization; and according to the privacy interference level, dynamically adjusting the fusion weight of the video and the voice to generate a stage evaluation parameter and a time sequence label, and finally constructing a psychological state evolution graph and outputting an evaluation report, so that the emotional feature fidelity and the psychological state analysis precision are remarkably improved.
Owner:HANGZHOU XINWA TECHNOLOGY CO LTD

Bluetooth earphone call intelligent noise reduction method based on cloud collaboration

The invention discloses a Bluetooth earphone call intelligent noise reduction method based on cloud cooperation, and the method comprises the following steps: S1, obtaining a call voice signal and an environment state parameter, carrying out the preprocessing, and constructing a voice spectrum sequence and an environment feature sequence; s2, performing noise estimation and suppression processing on the speech spectrum sequence by using a DCCRN model to obtain a spectrum feature sequence; s3, constructing a cooperative processing sequence based on the environment feature sequence and the spectrum feature sequence; s4, using a Conv-TasNet model to execute speech enhancement on the co-processing sequence, and generating a target spectrum sequence; s5, performing phase reconstruction and inverse frequency spectrum transformation based on the target frequency spectrum sequence to generate a target voice signal; and S6, constructing a loss function according to the target voice signal, and updating the Conv-TasNet model. According to the method, the DCCRN model and the like are fused, and the method has the advantages of high noise reduction precision, high environment adaptability and sustainable optimization.
Owner:深圳市美迪声科技有限公司

Radio analog telephone communication encryption method, device, equipment and medium

The invention discloses a radio analog telephone communication encryption method, device and equipment and a medium, and the method comprises the steps: generating a frequency spectrum segment number and a pseudo-random sorting sequence through a preset encryption algorithm; and carrying out controllable segmentation on the total spectrum of the analog voice according to a continuous sequence from low frequency to high frequency, carrying out sequence recombination on each sub-spectrum according to a pseudo-random sequence, finally combining and outputting encrypted voice, and recovering the original voice by utilizing the same pseudo-random sequence reverse recombination at a receiving end. Therefore, the time-frequency continuity and statistical characteristics of the voice spectrum in the traditional radio analog communication are fundamentally changed. According to the method, frequency spectrums are recombined as a whole according to a pseudo-random rule, so that an eavesdropper can only obtain disordered voice signals after the frequency spectrums are rearranged even if the eavesdropper intercepts the signals, and the difficulty of passive cracking means such as frequency spectrum analysis, filtering reduction and voice recognition is remarkably increased; the problem that in the prior art, radio analog telephone communication is insufficient in encryption and easy to crack is solved.
Owner:SHENZHEN TIANHAI COMM CO LTD

An emotion-controllable joint coding VITS speech synthesis method and related device

The present invention provides an emotion-controllable joint encoding VITS speech synthesis method and related devices, which relate to the field of speech synthesis technology. The method extracts emotion features from input speech samples and generates an emotion feature vector; constructs a set of emotion category relative ranking functions to generate a relative attribute vector; concatenates the emotion feature vector with the relative attribute vector to obtain a fused emotion feature representation; performs weighted concatenation of the text feature vector and the fused emotion feature representation to generate a joint feature vector; performs joint encoding on the joint feature vector; converts the jointly encoded features into a speech spectrum, and dynamically controls the emotional expression of the synthesized speech by adjusting the weight ratio of each emotion category in the relative attribute vector, thereby achieving a technical effect of effectively, conveniently, and flexibly controlling the emotion of the synthesized speech.
Owner:XIANGJIANG LAB

Low-power communication method and apparatus for wireless audio device, soc chip, and storage medium

PCT designated stageWO2026175268A1Source encodingData stream
A low-power communication method and apparatus for a wireless audio device, an SoC chip, and a storage medium. The low-power communication method for a wireless audio device comprises: sampling acoustic features of heterogeneous audio data from multiple target audio devices, calculating speech spectra corresponding to device characteristics, extracting multi-device acoustic features, and performing multi-frequency sub-band speech enhancement, heterogeneous segmentation and grouping of the multiple target audio devices, check value calculation, and differentiated encoding on each extracted acoustic feature sequence to obtain a multi-source encoded speech segment; dividing the multi-source encoded speech segment into transmission units, and performing cooperative signal modulation, cooperative resource allocation, and data stream transmission on each transmission unit to obtain a transmitted speech data stream; and acquiring the transmitted speech data stream transmitted with low power, and performing audio recovery on the transmitted speech data stream to obtain a target audio data stream transmitted by the multiple target audio devices, thereby reducing the communication power consumption of the multiple target audio devices and ensuring the audio transmission quality.
Owner:SHENZHEN HESHENGCHENG TECHNOLOGY CO LTD

An audio feature extraction method based on brain neural model

ActiveCN120412657BSpeech recognitionNoiseAuditory cortex
The present invention discloses an audio feature extraction method based on a brain neural model, comprising the following steps: obtaining speech signal data, preprocessing the speech signal, and dividing it into N speech segments; extracting acoustic features representing the emotional state from each speech segment, normalizing the acoustic features and synthesizing a multi-channel speech spectrogram; inputting the multi-channel speech spectrogram into an RBA-FE model, outputting a BDI score, predicting and classifying the degree of depression based on the BDI score, adjusting the parameters of each network layer in the RBA-FE model based on the difference between the predicted classification and the actual classification, and obtaining an optimal RBA-FE model for depression detection. The present invention simulates the cell selectivity in the auditory cortex of the human brain and adopts an adaptive activation mechanism to solve the overfitting and noise resistance of LSTM, thereby achieving accurate recognition of depressive speech under different noise interferences.
Owner:SOUTH CHINA UNIV OF TECH +1

Speech synthesis model training method, speech synthesis method, device, and product

The application relates to a speech synthesis model training method, a speech synthesis method, equipment and products. The speech synthesis model training method comprises the following steps: obtaining training speech spectrum information corresponding to a training speech sample of a training speech object; inputting the training speech spectrum information into a first coding module in a speech synthesis model to be trained, obtaining a first coding vector corresponding to the training speech spectrum information through the first coding module, and determining training vector distribution parameters corresponding to the first coding vector; obtaining a training object timbre feature corresponding to the training speech object based on the training vector distribution parameters, obtaining first synthesized speech according to the training object timbre feature and training text information corresponding to the training speech sample; obtaining a first model loss value according to the difference between the first synthesized speech and the training speech sample; and adjusting model parameters according to the first model loss value to obtain a pre-trained speech synthesis model, which can effectively improve the similarity between the timbre of the synthesized speech and the timbre of the speaker.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Speech enhancement method and system based on cross-modal knowledge migration

The invention discloses a voice enhancement method and system based on cross-modal knowledge migration. The speech enhancement method adopted by the invention comprises the following steps: inputting a text corresponding to an audio into a large language model to obtain context-related language embedding features; performing cross-modal knowledge migration, and fusing the language embedding features into the voice embedding features to obtain cross-modal embedding features; training a speech enhancement model; calculating loss between an enhanced speech spectrum and a target speech spectrum by using a mean absolute error loss function; using a cosine similarity loss function to calculate the loss between the cross-modal embedded feature and the output of the large language model; integrating the two losses, and optimizing parameters of the speech enhancement model through a gradient descent algorithm; and in the reasoning stage, voice enhancement processing is carried out only by using the trained voice enhancement model. According to the method, the enhancement effect of the voice in the noisy environment can be effectively improved, and text data or a language model does not need to participate in the reasoning stage.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT +1

Multi-point recording interference method and system

The invention relates to the technical field of information security, and discloses a multi-point recording interference method and system, and the method comprises the steps: obtaining environment information, and determining a basic interference volume according to the environment information; acquiring a sound source-node distance vector and sound field uniformity, and converting the sound source-node distance vector and the sound field uniformity into a spatial distribution coefficient; acquiring a voice initial instantaneous slope and an interference-voice cross-correlation function peak value, and converting the slope and the peak value into a dynamic response coefficient; receiving a target voice, obtaining the spectral characteristics and the voice spectrum centroid of the target voice, and converting the spectral characteristics and the voice spectrum centroid into a spectral efficiency coefficient; reading the basic interference volume, the spatial distribution coefficient, the dynamic response coefficient and the spectral efficiency coefficient, and generating an adjustment instruction; according to the invention, through collaborative analysis of the environment reference coefficient, the spatial distribution coefficient, the dynamic response coefficient and the spectrum efficiency coefficient, multi-dimensional adaptive matching of the interference volume and the environment state, the spatial layout, the voice dynamic characteristic and the spectrum characteristic is realized, and the flexibility is extremely high.
Owner:BEIJING HUAZHONG CHUANGSHI TECH DEV CO LTD

Speech enhancement method based on dynamic convolution and narrowband conformer

The present application relates to the technical field of speech processing, and particularly relates to a speech enhancement method based on dynamic convolution and narrowband Conformer, which comprises a training stage and a test stage and can realize high-quality speech enhancement.The speech enhancement model proposed in the present application is composed of a generator and a discriminator, the narrowband Conformer network is used in the generator to improve the extraction ability of the model to speech spectrum information, and the dynamic convolution is further used to replace the traditional convolution, so that the parameter quantity and the calculation quantity of the model are greatly reduced, the noise reduction effect is improved, and the operation efficiency of the algorithm and the stability and reliability of the model are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method and electronic device for reducing echo residue

The present application discloses a method and an electronic device for reducing echo residue. The method for reducing echo residue can be applied to the electronic device, and comprises: performing echo cancellation on a voice input signal according to an echo reference signal to obtain an echo cancellation signal; converting the echo reference signal into a reference spectrum signal of each frame by fast Fourier transform; converting the echo cancellation signal into a voice spectrum signal of each frame by fast Fourier transform; obtaining an a priori signal-to-noise ratio of a current frame by using the reference spectrum signal of the current frame and the voice spectrum signal of the current frame according to an additive noise principle; filtering the voice spectrum signal of the current frame by a Wiener filter coefficient of the current frame determined by the a priori signal-to-noise ratio of the current frame to obtain a target spectrum signal of each frame; and converting the target spectrum signal of each frame by inverse fast Fourier transform to obtain a target voice signal. Therefore, the method for reducing echo residue can accurately filter out residual echo.
Owner:ALI CORP

Speech signal processing method and device, electronic equipment and storage medium

The present disclosure relates to a speech signal processing method and device, electronic equipment and storage medium, and relates to the technical field of audio. The method comprises: acquiring a multi-channel far-field speech signal collected for a target sound source; inputting the multi-channel far-field speech signal into a preset multi-channel filtering model to perform speech enhancement processing to obtain target speech spectrum information; the preset multi-channel filtering model is a filtering model obtained by adaptively updating an original multi-channel filtering model according to the multi-channel far-field speech signal, or is a filtering model obtained by adaptively constructing according to a prediction result of a first neural network model for the multi-channel far-field speech signal; and inputting the target speech spectrum information into a preset near-field speech generation model to perform speech reconstruction processing to obtain a near-field speech signal corresponding to the multi-channel far-field speech signal. The technical scheme provided by the embodiments of the present disclosure can realize the transformation of speech hearing sensation and improve the quality and intelligibility of speech.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Artificial intelligence-based audio generation method, apparatus, device, and storage medium

The application relates to the field of artificial intelligence and discloses an audio generation method, device and equipment based on artificial intelligence and a storage medium, the method comprising the following steps: obtaining to-be-converted audio data, and obtaining a target field identifier corresponding to expected sound quality and / or expected emotion; performing cepstrum feature conversion based on a speech spectrum on the to-be-converted audio data to obtain cepstrum feature data of the to-be-converted audio data; inputting the cepstrum feature data and the target field identifier into a speech conversion model to encode the cepstrum feature data at an encoding layer, randomly sample the encoded vector at a resampling layer, and decode and reconstruct the vector sampled at a decoding layer based on the target field identifier to obtain target audio data; and on the basis of effectively removing sound rhythm information of a sounder, the application effectively converts the audio data into audio data with expected rhythm style information, thereby improving an audio conversion effect.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multimodal-based speech synthesis method and device, equipment and storage medium

The application relates to the technical field of artificial intelligence, and discloses a speech synthesis method and device based on multiple modes, equipment and a storage medium, which comprises the following steps: preprocessing a text to be synthesized to obtain character sequence information, character-level graph sequence information and word-level graph sequence information as input sequence information; encoding the character sequence information to obtain a time domain coding vector; encoding the character-level graph sequence information and the word-level graph sequence information to obtain a first space domain coding vector and a second space domain coding vector; performing first cross-modal attention calculation on the time domain coding vector and the first space domain coding vector to obtain a first decoding vector; performing second cross-modal attention calculation on the first decoding vector and the second space domain coding vector to obtain a second decoding vector; and obtaining a speech spectrum graph according to the second decoding vector to generate synthesized speech. The application guarantees the prosody and accuracy of the synthesized speech and effectively improves the level of financial services.
Owner:PING AN TECH (SHENZHEN) CO LTD

Frequency spectrum shaping parameter derivation method in LC3 coding framework based on neural network

The invention relates to the technical field of voice signal processing and audio coding, and discloses a frequency spectrum shaping parameter derivation method in an LC3 coding framework based on a neural network. The method comprises the following steps: performing frequency band energy analysis on an input audio frame, calculating energy of each frequency band according to a preset frequency band division rule, and obtaining multi-frequency band energy information representing spectrum distribution characteristics of the current frame; based on the multi-band energy information, combining a frame-level coding parameter and a transient detection mark, constructing an input feature vector for neural network reasoning, and predicting a spectrum shaping parameter consistent with an LC3 standard by a neural network; and the prediction parameters are further quantized and coded according to the LC3 standard and are used for subsequent spectrum shaping processing, so that on the premise of not changing an LC3 standard bit stream structure and decoder end behaviors, the improvement of a spectrum shaping parameter derivation mode is realized, and the stability and adaptability of coded speech spectrum features are improved.
Owner:GUANGDONG UNIV OF TECH

A speech filtering method, device, storage medium and equipment

The present application discloses a speech filtering method, device, storage medium and equipment, which belongs to the field of speech coding and decoding technology. The method mainly includes: encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain speech spectrum coefficients; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filtering module, inputting the target spectrum coefficients into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target speech signal corresponding to the target spectrum coefficients. The present application omits the complex post-filtering operation in the Bluetooth encoding process, and only uses the pre-trained neural network model for filtering in the Bluetooth decoding process, so that it achieves a sound quality close to that of standard decoding.
Owner:BEIJING BAIRUI INTERNET TECH CO LTD

Speech-based linear prediction codec post-processing method, device, equipment and medium

The invention discloses a speech-based linear prediction codec post-processing method and device, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: obtaining original decoding parameters of all speech frames obtained by processing an original speech signal through a standard linear prediction codec; performing parameter enhancement on the original decoding parameters by using a preset parameter enhancement model corresponding to each original decoding parameter to obtain enhanced parameters; obtaining an original voice spectrum based on the original decoding parameter, obtaining an enhanced voice spectrum based on the enhanced parameter, and splicing the original voice spectrum and the enhanced voice spectrum to obtain a spliced voice spectrum; processing the spliced speech spectrum by using a preset spectrum enhancement network to obtain a depth filter coefficient; and performing depth filtering on the enhanced speech spectrum based on the depth filter coefficient to obtain a target speech spectrum, and obtaining a target speech signal based on the target speech spectrum. The voice quality can be improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice state intelligent classification method based on voice spectrum characteristics and reinforcement learning optimization mechanism

The invention belongs to the technical field of artificial intelligence and medical speech analysis, and discloses a voice state intelligent classification method based on speech spectrum features and a reinforcement learning optimization mechanism. The method comprises the following steps: carrying out data preprocessing, acoustic feature extraction and feature optimization on a tested voice sample, and mapping the preprocessed tested voice sample into a two-dimensional Mel spectrogram; realizing voice state classification by using a multi-scale feature extraction structure comprising a first convolution branch, a second convolution branch and a third convolution branch and a bidirectional time sequence feature learning structure; meanwhile, a reinforcement learning optimization mechanism is introduced, feature selection, model structure configuration, training hyper-parameters and an updating strategy are subjected to self-adaptive optimization, confidence coefficient calibration, uncertainty judgment and interpretability result generation are combined, and a final voice state judgment result, calibrated confidence coefficient and an acoustic attention area are output. The voice state classification accuracy, stability and interpretability can be improved.
Owner:DALIAN UNIV OF TECH