Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

190 results about "Acoustic model" patented technology

An acoustic model is used in automatic speech recognition to represent the relationship between an audio signal and the phonemes or other linguistic units that make up speech. The model is learned from a set of audio recordings and their corresponding transcripts. It is created by taking audio recordings of speech, and their text transcriptions, and using software to create statistical representations of the sounds that make up each word.

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Real-time speech recognition method based on Bluetooth audio stream

The invention relates to the technical field of speech recognition, and discloses a real-time speech recognition method based on a Bluetooth audio stream, which comprises the following steps: analyzing bit allocation parameters to calculate quantized bit distribution and generate a frequency domain confidence mask, monitoring a packet loss concealment state flag bit of a decoder, forcibly setting the mask as a blocking threshold when an algorithm is activated, and generating a real-time speech recognition result. According to the method, a cross-level feature purification mechanism based on protocol priori and link states is constructed, acoustic model illusion caused by forged waveforms is blocked through targeted arbitration while deterministic quantization noise is eliminated, and the acoustic model recognition accuracy is improved. And the identification accuracy under a severe channel is ensured.
Owner:SHENZHEN HUIJIEXIN TECH CO LTD

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Corpus expansion method, system and equipment based on speech synthesis and medium

The invention relates to the technical field of speech synthesis, in particular to a speech synthesis-based corpus expansion method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing including data annotation based on the collected audio and corresponding text of a target speaker; extracting acoustic features from the preprocessed audio; on the basis of a pre-trained acoustic model, performing personalized fine tuning by using the annotation data and the acoustic features, and training personalized acoustic models of a plurality of speakers at the same time through multi-thread parallel computing; calling the trained personalized acoustic model, and synthesizing a voice corpus of the target text in combination with a vocoder; and based on the trained personalized acoustic model, continuously expanding the corpus by changing the text. The personalized voice corpus is quickly generated through a small number of voice samples, the data acquisition cost is remarkably reduced, and the corpus construction efficiency is improved.
Owner:深圳市友杰智新科技有限公司

Method of recognizing speech, device, and medium

A method of recognizing a speech, a device, and a medium. The method includes: processing, by using an acoustic model, speech data to be recognized and a first text segment obtained by recognition to obtain respective acoustic probabilities of a plurality of candidate text segments; processing the first text segment by using a first language sub-model to obtain respective initial language probabilities of the plurality of candidate text segments; processing the first text segment by using a constraint sub-model to obtain extendibility relationships of the plurality of candidate text segments with respect to the first text segment; adjusting the initial language probabilities of the candidate text segments according to the extendibility relationships to obtain respective first language probabilities of the plurality of candidate text segments; and determining a target text segment from the plurality of candidate text segments according to the first language probabilities and the acoustic probabilities.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Vehicle-mounted voice interaction method and system and readable storage medium

The invention relates to the technical field of intelligent vehicle-mounted systems, and discloses a vehicle-mounted voice interaction method and system and a readable storage medium, and the method comprises the steps: synchronously collecting initial voice and video data in a vehicle-mounted environment; performing wake-up word detection through a local acoustic model, and based on the detection confidence, extracting a mouth shape visual feature sequence by using a mouth shape recognition model to perform mouth shape verification so as to obtain a wake-up state and sound source positioning information; activating an interaction module at a corresponding position, and performing semantic recognition on the collected interaction voice and video data through a local model and a cloud model respectively; and finally, carrying out fusion cross validation on the local semantic recognition result and the cloud semantic recognition result to generate a final semantic recognition instruction, and executing corresponding operation by the vehicle-mounted system. According to the method, the recognition accuracy, the response speed and the robustness of vehicle-mounted voice interaction in a complex environment are improved, the false wake-up rate is effectively reduced, and the user experience is optimized.
Owner:深圳海冰科技有限公司

Intelligent customer service automatic-to-manual switching method and system based on emotion recognition

The invention relates to the technical field of artificial intelligence, in particular to an intelligent customer service automatic-to-manual switching method and system based on emotion recognition, and the method comprises the steps: obtaining a voice signal when a user carries out a call through terminal equipment; and extracting acoustic characteristic parameters based on the voice signal, inputting the acoustic characteristic parameters into a pre-constructed emotion recognition model, outputting an emotion state level, correcting the primary emotion state level based on an interaction influence coefficient and a semantic emotion score change coefficient, and generating a target emotion level. According to the intelligent customer service automatic-to-manual switching method based on emotion recognition, a multi-dimensional correction and judgment mechanism is introduced on the basis of traditional emotion recognition, a more accurate and reasonable customer service switching decision is achieved, the primary emotion state is recognized based on acoustic features, and the customer service switching efficiency is improved. And the emotion level is corrected in combination with a semantic emotion score change coefficient and an interaction influence coefficient, so that identification deviation caused by a single acoustic model is avoided.
Owner:SHANGHAI XINDIAN PHOTOELETRON TECH CO LTD

Wind-noise-resistant self-adaptive volume adjusting method and system for riding earphone

The invention relates to the technical field of audio signal processing, and discloses a wind-noise-resistant self-adaptive volume adjustment method and system for a riding earphone, and the method comprises the following steps: S1, collecting environment sound and a source audio signal in parallel; s2, analyzing environment sound, and determining macroscopic adjustment intensity; s3, analyzing the source audio in parallel, and identifying a content type and a transient signal of the source audio; s4, based on the psychological acoustic model, generating a frequency-related target method tone quality compensation gain; s5, performing dynamic smoothing and frequency weighting on the target gain according to the content and the transient characteristics to generate an application gain; s6, fusing the macroscopic adjustment intensity and the application gain, and synthesizing a final gain parameter; and S7, applying the final gain to the source audio, and outputting the compensated audio. According to the method, the psychological acoustic model and audio content parallel analysis are fused, frequency-level accurate compensation is realized, the problems of traditional degradation and sudden hearing sense change are solved, and the audio definition and comfort are improved.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD

Speech recognition method and system based on multi-modal fusion and adaptive noise modeling, and storage medium

The invention discloses a speech recognition method and system based on multi-modal fusion and adaptive noise modeling, and a storage medium, and belongs to the technical field of speech recognition. The method comprises the following steps: S1, synchronously acquiring multi-modal data; s2, performing voice preprocessing, and extracting an acoustic feature sequence; s3, extracting a visual lip shape feature sequence; s4, generating a low-dimensional noise vector; s5, performing cross-modal time alignment to obtain a multi-modal feature sequence corresponding to time; s6, unifying the semantic embedding space; s7, dynamically calculating and fusing the multi-modal features in the unified embedding space through an adaptive weighted fusion network to obtain a fusion feature vector; s8, noise self-adaptive denoising is carried out; s9, acoustic modeling is carried out through the end-to-end acoustic model, and a probability sequence is output; s10, decoding the probability sequence by the language model to obtain an initial text sequence; and S11, performing post-processing on the initial text sequence. The method can be adaptive to environmental noise, dynamic fusion between modes is realized, and high robustness is kept.
Owner:GUANGZHOU MARITIME INST

English pronunciation error correction training method based on speech recognition

The invention relates to the technical field of speech recognition and processing, in particular to an English pronunciation error correction training method based on speech recognition, and the method comprises the following steps: S1, collecting a speech signal generated by a learner in a pronunciation training process, digitalizing the speech signal, associating the digitalized speech signal with a target standard text, and generating an original audio data record with a timestamp; according to the invention, phoneme-level decoding is carried out on the voice signal by using the recurrent neural network acoustic model, and accurate alignment of the pronunciation of the learner and the standard phoneme sequence is realized in combination with the dynamic time warping algorithm, so that pronunciation errors such as misreading, missed reading and increased reading can be accurately identified; meanwhile, acoustic features such as Mel frequency cepstrum coefficient, pitch and fundamental frequency are extracted to be quantitatively compared with a standard native language pronunciation database, multi-dimensional evaluation covering accuracy, integrity, fluency and rhythm is generated, and the accuracy and systematicness of oral English pronunciation error correction are remarkably improved.
Owner:吕丽沙

An acoustic model training method and apparatus

The application provides an acoustic model training method and device. The acoustic model training method provided by the application comprises: acquiring unannotated acoustic samples; pre-training the acoustic model based on a self-supervised method, the acoustic model being used to predict the category of an input acoustic signal, in the self-supervised method, respectively generating masks for the time domain feature and the frequency domain feature of the acoustic signal based on the time domain feature and the frequency domain feature of the acoustic signal, the mask position and the mask quantity of different samples being different; dividing the levels to which the network structures of the pre-trained acoustic model belong, respectively freezing the layers of different levels, and asynchronously fine-tuning the acoustic model, the parameter freezing time of the layers of different levels being different, and the learning rate of the layers of different levels being different. The acoustic model training method and device provided by the application not only reduce the dependence on large-scale artificial annotation data, but also improve the training efficiency and task performance, and better adapt to task requirements and complex application scenarios.
Owner:HANGZHOU XUNSHENG MEDICAL TECHNOLOGY CO LTD

Acoustic and natural language processing models for velocity-based screening and monitoring of behavioral health

PendingJP2026136266AAcoustic modelBehavioural health
This invention provides an acoustic natural language processing model for predicting whether a person has a behavioral or mental health condition based on input speech. [Solution] A method for detecting behavioral or mental health status using an acoustic model including an encoder and a classifier, comprising: (a) acquiring a speech sample comprising a plurality of speech segments; (b) processing the speech sample with an encoder to generate an abstract feature representation, wherein the encoder is pre-trained to perform a first task other than detecting behavioral or mental health status; and (c) processing the abstract feature representation with a classifier to generate an output indicating whether or not the person has a behavioral or mental health status, wherein the classifier is trained on a training dataset comprising a plurality of speech samples from a plurality of speakers, and the speech samples are labeled as originating from or not originating from a speaker with a health status.
Owner:ELLIPSIS HEALTH INC

Speech recognition method and server

The application relates to a speech recognition method and a server. The method comprises the following steps: obtaining a to-be-recognized speech signal; recognizing each frame of the to-be-recognized speech signal according to an acoustic model of each language, and respectively outputting corresponding language phonemes and prediction probabilities; wherein the acoustic model of each language is respectively constructed according to shared hidden layer training; sequentially traversing a sentence decoding graph and a multi-language slot decoding graph connected with each other to obtain a corresponding path; wherein the sentence decoding graph is used for decoding phonemes entering a non-slot, and the slot decoding graph is used for decoding phonemes entering a slot; when it is determined that the path passes through the multi-language slot decoding graph in the speech decoding graph, screening the path according to the prediction probabilities of the language phonemes corresponding to each language, and determining the text information corresponding to the target path as a speech recognition result. The scheme provided by the application can accurately recognize mixed multi-language speech information.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Speech recognition method and device, equipment and storage medium

The invention relates to the technical field of speech recognition, and discloses a speech recognition method and device, equipment and a storage medium. Aiming at the defects of fixed filter coefficient, static context modeling and difficulty in deep training and deployment of the existing deep feed-forward sequence memory network (DFSMN), the method comprises the following steps of: acquiring a frame-level acoustic feature sequence of input voice, and inputting an acoustic model containing a neural network unit; in the unit, the filter weight of the memory module is generated in real time through a parameter generation network based on the intermediate feature of the current frame; performing weighted aggregation on the historical and future frame features by using the weight to obtain a context enhancement feature; and outputting features based on the feature generation unit, and decoding to obtain an identification result. According to the method, context dynamic self-adaptive modeling is realized, the recognition precision and robustness are improved, the efficient reasoning characteristic of a pure feed-forward network is kept, the calculation overhead is not remarkably increased, and the method is adaptive to a server and terminal equipment and is suitable for multi-scene speech recognition of keywords, command words and the like.
Owner:WUXUE GUANGJI INTELLIGENT BODY SOFTWARE TECHNOLOGY CO LTD

Road totally-enclosed sound barrier sound field prediction method and system

The invention relates to the technical field of environmental acoustics, discloses a road totally-enclosed sound barrier sound field prediction method and system, belongs to the field of environmental acoustics, and has very high practicability in sound field prediction and noise pollution evaluation of public road traffic noise after a totally-enclosed sound barrier is implemented. On the basis of the acoustic propagation principle, a new prediction method is provided for a road traffic noise source and a totally-closed sound barrier, a sound field is divided into two subsystems including a barrier internal space and an external space, and equivalent acoustic models of the two subsystems and a calculation method and recommendation parameters of the noise source sound power level are given respectively. According to the method, the sound environment of the high-rise building close to the urban road can be accurately predicted and calculated, and a basis is provided for noise pollution evaluation and control. And meanwhile, the method can be matched with a national standard method or commercial acoustic software for use, and has relatively high applicability.
Owner:CHINA SHIPPING ENVIRONMENT SCI & TECH (SHANGHAI) CO LTD

Data processing method and device, equipment and computer readable storage medium

ActiveCN116469378BImprove translation qualityfast trainingNatural language translationBiological modelsPattern recognitionGoal recognition
The application discloses a data processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: obtaining sample voice information and sample text information; processing the sample voice information through an acoustic model in an initial recognition model to obtain acoustic feature information; processing the acoustic feature information and the sample text information through a translation model in the initial recognition model to obtain a first predicted translation result and a second predicted translation result respectively; training the initial recognition model based on the first predicted translation result and the second predicted translation result to obtain a target recognition model; and the K vector and the V vector in the acoustic model are spliced with a prefix vector and / or the K vector and the V vector in the translation model are spliced with a prefix vector. The training efficiency of the target recognition model is improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

System and method for neural network multilingual speech recognition

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Self-adaptive voice dialogue generation system of intelligent toy and use method of self-adaptive voice dialogue generation system

The invention relates to a self-adaptive voice dialogue generation system of an intelligent toy and a use method thereof, and belongs to the technical field of voice dialogues of intelligent toys, the self-adaptive voice dialogue generation system comprises a voice acquisition module, a voice recognition module, a user portrait module, a dialogue management module, a natural language generation module, a safety filtering module and a voice synthesis module, the voice acquisition module acquires user voice signals, the voice recognition module is connected with the voice acquisition module and converts the voice signals into text information, and the user portrait module establishes and dynamically updates personalized user portraits corresponding to users. The user portrait at least comprises a cognitive level quantized value, an interest preference vector and an emotion mode label which are calculated on the basis of interaction historical data. According to the method, a high-fault-tolerance voice processing chain is formed from directional pickup of a beam forming algorithm at a hardware end to a CNN-LSTM acoustic model optimized for children at a software end, so that the smoothness and the stability of an interaction process in a real children use scene are fundamentally guaranteed.
Owner:XUZHOU HUAPEI INTELLIGENT MFG TECH CO LTD

A meta-learning based adaptive text-to-speech method and related devices thereof

The application belongs to the field of artificial intelligence and relates to a self-adaptive text-to-speech method based on meta-learning, which comprises the following steps: pre-training according to a full data set to obtain an initial value of a preset acoustic model; sampling sound training sample data, performing feature training through the preset acoustic model to generate a mel spectrum, and generating style encoding through a preset style encoder; performing adaptive instance normalization processing on layer normalization of the preset acoustic model to obtain a target acoustic model comprising a target mel spectrum; and finally converting stranger sample data to output target speech data with style encoding. The application also provides a self-adaptive text-to-speech device based on meta-learning, a computer device and a storage medium. In addition, the application also relates to blockchain technology, and the data involved in the conversion process can be stored in the blockchain. The application can reduce the complexity of training, realize adaptive learning and conversion of small sample data.
Owner:PING AN BANK CO LTD

Camouflage voice voiceprint recognition method based on Transform model and mixed features

PendingCN121862122Aimprove performanceFitting feature distribution is goodSpeech recognitionFeature extractionGammatone filter
The invention relates to the technical field of speech processing, and particularly provides a disguise speech recognition method based on a Transform model and mixed features, which is carried out from two aspects of feature extraction and model establishment. A resonance peak parameter is calculated by adopting a cepstrum method, a cepstrum coefficient (GFCC) is obtained through a Gammatone filter bank, then the resonance peak, the GFCC and a difference coefficient of the GFCC are combined into a mixed characteristic parameter, and complementary correlation between mixed characteristics is mined. From the perspective of model establishment, the mixed features are used as the input of the model, and the Transform network model is used as the acoustic model of the voiceprint recognition system, so that the feature distribution is better fitted, the classification effect is remarkably improved, and the performance of the camouflage voice voiceprint recognition system is effectively improved. The problem of performance degradation caused by feature redundancy and modal noise in a traditional method is solved.
Owner:CHINA CRIMINAL POLICE UNIV

Systems and methods for predicting mental health conditions based on passive processing of conversational speech and language

PendingUS20260100259A1Semantic analysisMedical automated diagnosisMedicineMore language
Described herein are systems and methods for identifying the severity of a mental health condition or symptoms of same by listening to a human-to-human conversation by receiving conversation data, processing the conversation data to generate a language model output and / or an acoustic model output using one or more language models and / or acoustic models. Further described herein are systems and methods for automatically tracking and providing analytics on self-report questionnaires administered during the conversation.
Owner:ELLIPSIS HEALTH INC

An efficient training and high expressiveness speech conversion model based on an acoustic model and a vocoder decoupled architecture

This invention discloses an efficient training and high-performance speech conversion model based on a decoupled architecture of acoustic model and vocoder, comprising an acoustic model and a vocoder; the acoustic model includes a speaker encoder, a content encoder, a normalized stream, a posterior encoder, a Mel decoder, and a discriminator. Its advantages include significant technological breakthroughs in improving speech conversion model training efficiency, sound quality, emotional expression, and interactive control, providing a novel solution for high-quality, highly controllable speech synthesis systems, and possessing good practical value and promising industrial application prospects.
Owner:HAPPY ELEMENTS TECH (BEIJING) CO LTD

Method, device, and system for providing an AI solution that detects deep voices generated by generative AI and prevents related accidents

PendingKR1020260114008ASignal qualityAlgorithm
According to one embodiment of the present invention, a system for detecting deep voice generated by generative artificial intelligence is provided, comprising: a memory for storing instructions; and one or more processors for executing said instructions, wherein the processor extracts a feature vector including frequency characteristics and temporal change patterns from input voice data, inputs said extracted feature vectors to an ensemble artificial intelligence model including a first AI model which is an acoustic model, a second AI model which is a language model, and a third AI model which is a pattern recognition model to calculate each individual deep voice probability, measures the signal quality of said input voice data to calculate a signal-to-noise ratio (SNR), calculates a final deep voice probability through an adaptive weight allocation algorithm that varies the weights to be applied to said AI model in real time based on said calculated SNR value, and generates a deep voice suspicion alarm when said final deep voice probability exceeds a threshold value.
Owner:METACROWD CORP

Traditional Chinese medicine voice input method and system based on semantic perception and dynamic dialectical reasoning

This invention belongs to the field of natural language processing technology. To address the problem of low input accuracy in existing technologies, it discloses a method and system for inputting traditional Chinese medicine (TCM) speech based on semantic perception and dynamic dialectical reasoning. This invention identifies TCM diagnostic elements in electronic medical records and dynamically constructs a diagnostic medicine pool by combining it with a TCM knowledge graph. This allows speech recognition to move beyond mechanical reliance on acoustic models and integrate with clinical diagnosis. Simultaneously, this invention expands the acoustic candidate word sequence with TCM-related confusing words, generating a more comprehensive set of TCM candidates. This set is then combined with the diagnostic medicine pool to generate candidate TCMs. Next, the semantic matching score between each candidate TCM and the diagnostic elements is calculated, and the corresponding real TCM name is determined based on this score. Therefore, this invention avoids mechanical recognition detached from syndromes and treatment principles and can effectively distinguish TCMs that are confused by homophones or near-homophones, thus reducing the input error rate.
Owner:CHENGDU ZIJIELIU TECH CO LTD

Training method of music generation model, music generation method, equipment and medium

The invention provides a training method of a music generation model, a music generation method, equipment and a medium, and relates to the technical field of artificial intelligence. Performing first up-sampling training on the training audio coding information based on a first generator in a vocoder in a music generation model to obtain intermediate training audio data of a first sampling rate, the intermediate training audio data is subjected to first up-sampling training based on a first generator in the vocoder, second up-sampling training is carried out on the intermediate training audio data based on a second generator in the vocoder to obtain training audio data of a second sampling rate, the second sampling rate is larger than the first sampling rate, and generative adversarial training is carried out on the vocoder based on the training audio data to obtain a trained music generation model. By adopting the method and the device, the dependence of model training on high-sampling-rate audio data can be reduced while the music generation model is ensured to output high-fidelity songs, so that the model training cost is controlled.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

An end-to-end speech splicing synthesis method

The application provides an end-to-end voice splicing synthesis method, adopts a completely end-to-end acoustic model to model acoustic distribution of a splicing unit, the model directly takes a phoneme sequence as input, does not need expert knowledge to design an acoustic model, and greatly reduces the complexity of acoustic modeling. In addition, the end-to-end acoustic model has stronger sequence modeling capability than the traditional acoustic model, and can output smoother acoustic parameters. Further, by using a hybrid probability network as the output of the end-to-end acoustic model, the acoustic model can more accurately describe the distance of high-dimensional spectral features.
Owner:SUZHOU QIMENGZHE NETWORK TECH CO LTD