Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

65 results about "Voice Training" patented technology

A variety of techniques used to help individuals utilize their voice for various purposes and with minimal use of muscle energy.

Phoneme alignment model training and speech synthesis method and device, equipment and medium

ActiveCN120748366ASpeech synthesisSpeech trainingMedicine
The invention relates to the technical field of speech synthesis, in particular to a phoneme alignment model training and speech synthesis method and device, equipment and a medium. The method comprises the following steps: performing convolution attention alignment on spectrum feature information and text feature information obtained according to voice training data to obtain a first alignment matrix; monotonic alignment search is executed based on the first alignment matrix to generate a second alignment matrix, wherein the second alignment matrix is a binary hard attention matrix; calculating relative entropy loss according to the first alignment matrix and the second alignment matrix; expanding the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing linear transformation on the expanded text feature information to generate a predicted Mel spectrum; calculating Mel loss according to the spectrum feature information and the predicted Mel spectrum; and training the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met. By adopting the method, the phoneme alignment accuracy can be improved, and the speech synthesis accuracy is further improved.
Owner:SHANGHAI PAIDI INTELLIGENT TECH CO LTD

Endoscope with Voice Control

An endoscope with voice control. A data processor for the endoscope obtains voice training data for a specific surgeon, and / or information indicating spoken utterances associated with a surgical procedure for which the endoscope is to be used. Voice utterances signals are received from a surgeon during the surgery. Speech recognition technology decodes the voice utterance signals into commands to control the endoscope, based at least in part on the data obtained from the database, and control commands are issued to implement the decoded voice utterances. Concurrently with the accepting and decoding of voice utterances, the processor accepts and decodes control signals from at least one other input device. The processor issues control commands to components of the endoscope to implement the decoded other input control signals. The voice utterances are accepted, decoded, and executed concurrently with the other input control signals.
Owner:PSIP2 LLC

Voice training partner method and device, electronic equipment and storage medium

The embodiment of the invention discloses a voice partner training method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a universal cue word of voice partner training and a current cue word corresponding to a current round of dialogue in a simulation service scene in a current dial test; the current broadcast text is accurately and conveniently generated based on the general prompt word, the current prompt word and the preset text generation model, so that the generation efficiency of the current broadcast text is improved; if it is detected that the current broadcast text does not contain the last sentence identifier, determining a current broadcast voice corresponding to the current round of dialogue based on the current broadcast text, the current prompt word and a preset voice generation model; the current broadcast voice is the voice with emotions, so that the emotional effect of voice broadcast is realized; determining a current response text corresponding to the current round of dialogue based on the received current response voice of the user and a preset voice recognition model; and based on the current broadcast text, the current response text and the current cue word, determining a next cue word corresponding to the next round of dialogue.
Owner:AGRICULTURAL BANK OF CHINA

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Speech recognition model training method and device, equipment and readable storage medium

The invention discloses a speech recognition model training method and device, equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: firstly, acquiring voice training data and annotation data corresponding to the voice training data; the annotation data comprises text annotation data and intention annotation data; fuzzy processing is carried out on the text labeling data, and voice features of the voice training data are extracted; and training a speech recognition model based on the speech features, the text annotation data and the intention annotation data until the speech recognition model converges. According to the method, the text annotation data is fuzzified, some characters which are not concerned about intention classification are ignored, the speech recognition model is more focused on keywords, intention classification information is introduced in the training process, and the accuracy of the recognition result generated by the speech recognition model for intention classification is improved.
Owner:AISPEECH CO LTD

Vocalise bottle

1. Name of the designed product: voice training bottle (all-voice). 2. Use of the designed product: the designed product is all-voice voice training bottle. 3. Design points of the designed product: combination of shape and pattern. 4. Picture or photo best indicating the design points: perspective view.
Owner:张悦

Natural language processing systems and methods for intent classification of speech transcription

Aspects of the subject disclosure may include, for example, generating a natural language processing model by training an automatic speech recognition (ASR) encoder with manual transcription. The training is performed by correcting and adjusting relevant factors of the ASR encoder based on determined triplet loss, classification loss and Kullback-Leibler divergence loss. In response to an ASR utterance, the trained natural language processing model generates a predicted intent associated with the ASR utterance with improved accuracy. Other embodiments are disclosed.
Owner:JPMORGAN CHASE BANK NA

Custom voice instruction recognition method and device based on twin network, and electronic equipment

The invention discloses a custom voice instruction recognition method and device based on a twin network and electronic equipment. Preprocessing the collected voice instruction data to obtain a voice training data set; constructing a twin network architecture; performing end-to-end joint training on the identification model and the discrimination model; deploying the trained identification model and discrimination model to the electronic equipment; performing feature extraction on the voice data, inputting the extracted features into the recognition model, and calculating to obtain a reference space vector; collecting a to-be-recognized voice signal in real time, and inputting the extracted voice features into the recognition model to calculate and obtain a detection space vector; and jointly inputting the detection space vector and the reference space vector into a discrimination model, calculating to obtain a discrimination value, comparing the discrimination value with a preset threshold value, and judging whether a target voice instruction is recognized or not according to a comparison result. The user-defined voice recognition instruction can be newly added on the basis of the preset voice instruction, the voice recognition effect is optimized, and the user interaction experience is improved.
Owner:ZHUHAI SPACETOUCH LTD

Voice training noise adding system and method based on mixed noise generation model

The invention aims to provide a voice training noise adding system and method based on a mixed noise generation model. The system comprises an input module, a noise environment enhancement module, a voice noise enhancement module and an output module. Wherein the input module is used for acquiring noise environment simple description information and clean voice data to be enhanced; the noise environment enhancement module converts the noise environment simple description information into a structured noise event sequence with time sequence characteristics; a voice noise enhancement module generates multi-source mixed noise according to the noise event sequence, and adds the multi-source mixed noise to the clean voice data according to a preset rule to obtain noisy voice data; and the output module is used for outputting the noisy voice data for voice model training. According to the invention, interaction characteristics of superposition, offset, interference and the like of different noise sources in the time dimension are fully joined, and the technical bottleneck that only linear superposition can be realized in a traditional mixing mode is solved.
Owner:GUANGDONG UNIV OF TECH

Speech synthesis method and device, electronic equipment and storage medium

The present application relates to the technical field of audio processing, and provides a speech synthesis method and device, electronic equipment and storage medium. The target acoustic feature of the to-be-processed text is obtained by extracting the acoustic feature from the to-be-processed text using a preset acoustic model; the target fundamental frequency feature is obtained by extracting the fundamental frequency feature from the target acoustic feature using a pre-trained fundamental frequency predictor, and the target energy feature is obtained by extracting the energy feature from the target acoustic feature using a pre-trained energy predictor; the target acoustic feature, the target fundamental frequency feature and the target energy feature are input into a pre-trained general vocoder to generate the speech audio of the to-be-processed text; the general vocoder is trained based on the speech audio of multiple speakers. By taking the acoustic feature, the fundamental frequency feature and the energy feature as the input of the vocoder for speech synthesis, and by training the vocoder using the speech of multiple speakers, the vocoder has general applicability, the training time of the vocoder is reduced, and the effect of speech synthesis is ensured.
Owner:SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD

Sound information protection method and device, storage medium, and electronic device

Embodiments of the present application provide a sound information protection system and method, a storage medium and an electronic device. The method comprises: after determining that a current voice incoming call belongs to the target call and obtaining an instruction to perform sound conversion, converting the sound of the current user into the sound of the target speaker through a lightweight voice conversion model, and then calling the current incoming call party. The lightweight voice conversion model is a model trained using sample voice containing the voice of the target speaker, and comprises a posterior encoder, a prior encoder and a decoder. The problem of how to avoid the leakage of user voice features in the related art is solved.
Owner:NANJING SILICON INTELLIGENCE TECH CO LTD

A voice activity detection method, system, terminal and storage medium

The application provides a voice activity detection method, system, terminal and storage medium, relates to a voice data processing technical field, and in particular to a voice activity detection method, which comprises the following steps: obtaining a voice training sample carrying a voice label; performing feature extraction according to the voice training sample to obtain a root amplitude spectrum feature sample and an Fbank feature sample; and constructing a voice activity detection model, wherein the voice activity detection model is used for performing feature fusion on the root amplitude spectrum feature and the Fbank feature obtained by performing feature extraction on voice data to obtain fusion features, and outputting a probability value of each frame of the voice data existing a voice signal based on the fusion features; training the voice activity detection model by using the root amplitude spectrum feature sample and the Fbank feature sample to obtain a trained voice activity detection model; and performing voice activity detection by using the trained voice activity detection model. The application can improve the accuracy of voice activity detection.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Speaker recognition method and related apparatus, device and storage medium

The application discloses a speaker recognition method and related device, equipment and storage medium, wherein the speaker recognition method comprises: obtaining a first voice of a first speaker and a second voice of a second speaker; determining whether the first speaker and the second speaker are the same speaker based on a feature distance between a first voice feature of the first voice and a second voice feature of the second voice; wherein the voice feature is extracted by a feature extraction model, the feature extraction model is trained based on a sample voice set to minimize the training loss, the sample voice set contains sample voices labeled with the sample speaker, the training loss is positively correlated with a first feature distance between the sample voice feature of the reference voice and the positive example voice feature, and negatively correlated with a second feature distance between the sample voice feature of the reference voice and the negative example voice feature. The above scheme can improve the accuracy of the feature extraction model in extracting voice features, thereby improving the speaker recognition accuracy.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Zero-shot task expansion of ASR models using task vectors

A method includes training, using an un-supervised learning technique, an auxiliary ASR model based on a first set of un-transcribed source task speech utterances to determine a first task vector, training, using the un-supervised learning technique, the auxiliary ASR model based on a second set of un-transcribed speech utterances to determine a second task vector, and training, using the un-supervised learning technique, the auxiliary ASR model based on un-transcribed target task speech utterances to determine a target task vector. The method also includes determining a first correlation between the first and target task vectors, determining a second correlation between the second and target task vectors, and adapting parameters of a trained primary ASR model based on the first and second source task vectors and the first and second correlations to teach the primary ASR model to learn how to recognize speech associated with the target task.
Owner:GOOGLE LLC

Voice training device supporting human-computer interaction

The utility model relates to the technical field of voice training devices, in particular to a voice training device supporting human-computer interaction, which comprises a shell, a storage bin and a protective cover, the storage bin is fixedly connected with one side surface of the shell, the other side surface of the shell is provided with a mounting groove, the inner bottom surface of the mounting groove is slidably connected with a plurality of clamping jaws, and the protective cover is mounted on the outer side of the mounting groove. The inner side face of the containing bin is fixedly connected with an inflatable doll, the inflatable doll is folded and filled in the containing bin, the outer side face of the containing bin is provided with an air nozzle, the air nozzle is connected with the inflatable doll, the bottom of the clamping jaw is connected with sliding pieces, the sliding pieces are inserted into the bottom of the installation groove, and an elastic rope is arranged between every two opposite sliding pieces. The electronic device is used for voice interaction of a user on the back face of the device and is deflated and stored when carried, the doll is conveniently stored and taken through the storage bin, the structure of the device is simplified, meanwhile, the overall cost of the device is greatly reduced, and meanwhile the requirements of special crowds are met through the inflatable doll.
Owner:HUNAN ART VOCATIONAL COLLEGE

Multi-dimensional AI platform intelligent voice response system using voice synthesis technology

PendingCN120808785ASpeech recognitionSpeech synthesisSpeech trainingText entry
The invention relates to the technical field of voice synthesis, in particular to a multi-dimensional AI platform intelligent voice response system using the voice synthesis technology, which screens out matched historical user question voices according to text semantic similarity corresponding to current user question voices and historical user question voices, and sends the matched historical user question voices to a user terminal. Obtaining a voice training set and the weight of each element in the voice training set according to the voice feature similarity between the question voice of the current user and the question voice of the matched historical user in combination with the user score value of the manual reply voice, training an acoustic model, obtaining a reply text corresponding to the question voice of the current user, and obtaining the question voice of the current user; and inputting into the trained acoustic model, generating a reply voice signal, and then outputting to the current user. According to the method, the voice training set is screened and constructed from historical manual reply voices, so that an acoustic model can learn a more natural and smooth voice synthesis mode, and the reply voice signal contains rich voice features and expression modes.
Owner:ROPEOK TECHNOLOGY GROUP CO LTD

Synthetically generating inner speech training data

ActiveUS12586568B2SensorsDiagnostic recording/measuringSpeech trainingAcoustics
Methods and systems are disclosed for synthetically generating inner speech training data. The methods and systems access a collection of overt speech signals representing phonemes, phoneme sounds, words or phrases spoken at least partially using overt speech. The methods and systems transform the collection of overt speech signals into inner speech training data comprising electromyograph (EMG) data representing inner speech corresponding to the phonemes, phoneme sounds, words or phrases spoken at least partially using the overt speech. The methods and systems train a machine learning model to decode inner speech signals into a set of corresponding phonemes, phoneme sounds, words or phrases based on the inner speech training data.
Owner:SNAP INC

Voice recovery degree identification method and device, equipment and storage medium

The invention relates to a voice recovery degree identification method and device, equipment and a storage medium. The method comprises the following steps: collecting voice data of multiple voices of a user and target voice training data, and extracting voice feature information of each piece of voice data; obtaining voice evaluation index information of a plurality of evaluation types, and for each piece of voice data, dividing each piece of voice feature information of the voice data into voice feature information of different evaluation types; respectively identifying an evaluation index value corresponding to the voice feature information of each evaluation type so as to identify a target voice evaluation value of the user in each evaluation type; and based on the target voice evaluation value of the user in each evaluation type, identifying the initial voice recovery degree of the user, querying the voice recovery degree range corresponding to the target voice training data of the user, and then determining the target voice recovery degree of the user. By adopting the method, the recognition accuracy of the voice recovery degree of the user can be improved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Training method and device of heterogeneous language model, equipment and storage medium

Embodiments of the present application provide a method, device and storage medium for training a heterogeneous language model. The method comprises: obtaining a voice training sample set; training a first initial network model and a second initial network model using the voice training sample set to obtain at least two first network models and at least two second network models; the first network model and the second network model have different structures, the first network model is used to process an input pinyin sequence to obtain at least one character sequence corresponding to the pinyin sequence, and the second network model is used to determine a target character sequence corresponding to the pinyin sequence from the at least one character sequence; and determining the heterogeneous language model according to the at least two first network models and the at least two second network models. The method, device and storage medium for training the heterogeneous language model provided by the embodiments of the present application are used to improve the accuracy of the language model.
Owner:WEBANK (CHINA)

A speech enhancement method based on bidirectional gated neural network and reference noise

The application provides a speech enhancement method based on BGRU and reference noise, and contains the following network structure: during reasoning, two-way input is accepted: reference noise and noisy speech, the two-way noise is respectively input to an enhancement network module through a short-time Fourier transform encoder to obtain noise estimation and speech estimation, and the speech estimation is input to an inverse short-time Fourier transform to obtain the speech after noise reduction; during training, four-way input is accepted: reference noise, noisy speech, noise true value and speech true value, the four-way noise is respectively input to a short-time Fourier transform encoder, the encoded reference noise and noisy speech are calculated through the enhancement network module to obtain noise estimation and speech estimation; the encoded speech true value and noise true value, and the noise estimation and speech estimation are all input to a 3CL loss module, the loss is calculated and fed back to the enhancement network module to adjust the parameters and improve the performance of the enhancement network module. The application can improve the speech quality and intelligibility under the condition of extremely low signal-to-noise ratio.
Owner:TIANJIN UNIV

Lip shape driving model generation method and device, electronic equipment, and storage medium

The present disclosure provides a method and device for generating a lip driving model, electronic equipment and storage medium, relating to the technical field of artificial intelligence, in particular to the technical field of computer vision, augmented reality, virtual reality, deep learning and the like, and can be applied to the scene of meta universe, virtual digital person and the like. It comprises: inputting audio data, a mask image and a reference face image into an initial lip driving model to obtain a lip image; determining a first loss according to the difference between the lip image and a sample face image; inputting the audio data and the lip image into a plurality of synchronization networks generated based on different types of speech respectively to obtain a second loss output by each synchronization network, and correcting the initial lip driving model according to the first loss and the minimum value of the plurality of second losses to obtain a lip driving model. Thus, the generated lip driving model can have high accuracy in different types of speech scenarios.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

system

An object of the system according to the embodiment is to consistently perform lyrics writing / composition, voice training, and production based on personal information.SOLUTION: A system includes an information collection part, an analysis part, a lyric composition part, a voice training part, a production part, and a debut support part. The information collection unit collects information on at least one of a user's daily conversation, SNS, and diary. The analysis unit analyzes the information collected by the information collection unit. The lyrics composer performs lyrics composition based on the information analyzed by the analyzer. The voice training unit performs voice training based on the lyrics and the melody generated by the lyrics composing unit. The producing unit produces the performance of the user trained by the voice training unit. The debut supporting unit supports debut of the user produced by the producing unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Adaptive self-trained computer engines with associated databases and methods of use thereof

In some embodiments, the present invention provides for an exemplary computer system which includes at least the following components: an adaptive self-trained computer engine programmed, during a training stage, to electronically receive an initial speech audio data generated by a microphone of a computing device; dynamically segment the initial speech audio data and the corresponding initial text into a plurality of user phonemes; dynamically associate a plurality of first timestamps with the plurality of user-specific subject-specific phonemes; and, during a transcription stage, electronically receive to-be-transcribed speech audio data of at least one user; dynamically split the to-be transcribed speech audio data into a plurality of to-be-transcribed speech audio segments; dynamically assigning each timestamped to-be-transcribed speech audio segment to a particular core of the multi-core processor; and dynamically transcribing, in parallel, the plurality of timestamped to-be-transcribed speech audio segments based on the user-specific subject-specific speech training model.
Owner:VOXSMART LTD

Methods to assist verbal communication for both listeners and speakers

Methods implemented in a system utilizing computing programs for a speaker and a listener in conversation are provided. Aspects include (i) a reminder provisioner for a speaker which is triggered according to speed, pitch or volume of the speaker's speech, (ii) a speech training provisioner for a speaker, and (iii) an application which records and plays back difficult conversation to understand.
Owner:SATO HIROKI

Wakeword detection

Techniques for implementing multiple wakeword detectors on a single device are described. A digital signal processor (DSP) of the device may implement a wakeword detection component to detect when captured speech includes a wakeword. A companion application installed on the device may implement a wakeword detection component trained using speech of a user of the device. If the DSP's wakeword detection component detects a wakeword in speech, the companion application's wakeword detection component may be used to determine whether the wakeword was spoken by the user of the device. If the companion application's wakeword detection component determines the user spoke the wakeword, audio data representing the speech may be sent to at least one server(s) for processing.
Owner:AMAZON TECH INC

Speech training data generation method, device, equipment, medium and program product

PendingCN122511223Aachieve recognizabilityImplement labelingSpeech trainingTimestamp
This application relates to a method, apparatus, device, medium, and program product for generating speech training data. The method includes: acquiring initial training data, the initial training data comprising at least one audio-text pair; processing the audio and text in each audio-text pair to obtain first timestamp information corresponding to each word in the audio-text pair; based on the audio in each audio-text pair, obtaining each sub-language event and second timestamp information corresponding to the sub-language event; based on the first timestamp information corresponding to each word and the second timestamp information corresponding to the sub-language event, generating text insertion positions corresponding to each sub-language event; and based on the text insertion positions corresponding to each sub-language event and the initial training data, obtaining target training data. This method can reduce costs.
Owner:MOORE THREADS TECH CO LTD

An end-to-end model training method and device, computer equipment and storage medium

ActiveCN114882874BSpeech recognitionSpeech trainingEngineering
The embodiment of the application belongs to the technical field of speech recognition in artificial intelligence, and relates to an end-to-end model training method and device applied to speech recognition, computer equipment and a storage medium. The output of an acoustic model is taken as expanded text of audio training data, and the expanded text and audio annotation text are taken as language model input to train the speech recognition model, thereby effectively solving the problem of too limited annotation text content in a traditional speech training set, enabling the language model of the speech recognition model to learn more comprehensive information, thereby effectively improving the recognition accuracy of the speech recognition model, and to a certain extent, reducing the coupling degree of acoustic information and language information in the end-to-end model, improving the robustness of the entire model in different scenes, especially when recognizing speech in different fields, avoiding the problem of a large decrease in accuracy when changing application scenarios, and increasing the flexibility of the model in actual use and deployment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-dimensional voice training plan generation method and system

The invention discloses a method and a system for generating a multi-dimensional voice training plan. The method comprises the following steps: firstly, collecting age, gender and voice samples of a patient, and obtaining seven-dimensional evaluation parameters including sound pressure, amplitude perturbation, maximum sounding duration, fundamental frequency perturbation, vowel space area, intonation damage and speech understanding degree score; and then, according to a preset logic, sequentially performing judgment based on the parameters, and dynamically combining different training modules of loudness, breath, pitch, vowel, gliding, tone, consonant and the like into a personalized voice training plan. Through multi-dimensional evaluation and dynamic module matching, the problem of training scheme solidification in the prior art is solved, and the individuation degree and rehabilitation effect of voice training are remarkably improved.
Owner:BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)

Speech recognition method and system based on large model and speech synthesis engine

PendingCN121306105ASpeech recognitionSpeech synthesisSpeech trainingAutomatic speech
The invention relates to the technical field of speech recognition, provides a speech recognition method and system based on a large model and a speech synthesis engine, and greatly improves the recognition accuracy of ASR recognition by combining a large language model language with a new generation of speech synthesis engine. A large language model is utilized to generate a large batch of high-quality corpora, and a speech synthesis engine is utilized to obtain a large amount of high-quality speech training data for training an automatic speech recognition model and checking and correcting a recognition result at the same time. A training closed loop of generating high-quality text training data, synthesizing natural speech, recognizing natural speech, finding errors, correcting the errors and generating training data is realized, so that an automatic speech recognition model can continuously, professionally and automatically reinforce learning at low cost; and furthermore, a high-accuracy and high-quality speech recognition result is output under semantic errors and complex scenes.
Owner:GUANGZHOU AUSUN INFORMATION TECH CO LTD

Model training method, voice wake-up method, device, equipment and medium

This application discloses a model training method, a voice wake-up method, a device, an apparatus, and a medium. The model training method includes: extracting features from acquired voice training samples to obtain first audio features of the voice training samples; using the first audio features, training at least two wake-up sub-models, each including an encoding network and a decoding network; wherein the at least two encoding networks have different parameter counts, and all decoding networks have the same model structure and parameter counts; using the first audio features, jointly training the constructed initial wake-up model to obtain a voice wake-up model; wherein the initial wake-up model includes a joint decoding network and the encoding networks in the at least two wake-up sub-models; the initial model parameters of the joint decoding network are initialized based on the model parameters of the decoding networks in the at least two wake-up sub-models; the voice wake-up model is used to recognize user audio data to wake up electronic devices.
Owner:MOORE THREADS TECH CO LTD