Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Speech training" patented technology

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Adversarial training of keyword spotting to minimize TTS data overfitting

PendingUS20260051318A1Biological modelsSpeech recognitionSpeech trainingHidden layer
A method includes receiving training utterances that include non-synthetic speech training utterances and synthetic speech utterances. For each training utterance, the method includes processing, using a memorized neural network, a corresponding sequence of input audio frames to generate a hotword detection output indicating a likelihood the training utterance includes a hotword, determining a first loss based on the hotword detection output, obtaining a hidden layer feature vector for each corresponding input audio frame; processing, using a speech classification model, the hidden layer feature vectors to predict a classification output for the training utterance; and determining an adversarial loss based on the classification output predicted for the training utterance. The method also includes training the memorized neural network on the first losses and the adversarial losses to teach the memorized neural network to learn how to detect the hotword in audio and prevent overfitting of the synthetic speech training utterances.
Owner:GDM HOLDING LLC

Speech communication aid trainer

ActiveCN309803572SSpeech trainingAcoustics
1. The name of the design product: speech communication auxiliary training device. 2. The use of the design product: for speech training or auxiliary communication for people with language barriers. 3. The design points of the design product: in shape. 4. The picture or photo that best shows the design points: perspective view 1.
Owner:谭文斯

Speech recognition method based on bimodal mixed contrast enhancement, electronic device, chip, storage medium and program product

The application provides a speech recognition method based on bimodal mixed contrast enhancement, an electronic device, a chip, a storage medium and a program product; the method comprises the following steps: acquiring multi-modal speech training data; a bimodal mixed contrast learning model is constructed and trained, wherein samples with the same emotional label under the same mode are taken as positive sample pairs, samples with different emotional labels are taken as negative sample pairs, intra-modal contrast loss is calculated to enhance the ability of the bimodal mixed contrast learning model to distinguish intra-modal fine-grained emotional features; samples with the same emotional label under different modes are taken as positive sample pairs, samples with different emotional labels under different modes are taken as negative sample pairs, inter-modal contrast loss is calculated to realize the alignment and complementarity of different modal feature spaces; the intra-modal contrast loss and the inter-modal contrast loss are fused to obtain multi-modal mixed contrast loss, and the model parameters of the bimodal mixed contrast learning model are optimized; based on the features extracted by the bimodal mixed contrast learning model, a downstream speech recognition task is trained.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Verbal communication aid training device

ActiveCN309812665SSpeech trainingEngineering
1. Name of the product in this design: Speech Communication Assistive Training Device. 2. Purpose of this design: To provide speech training or assist communication for people with language impairments. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: 3D view 1.
Owner:谭文斯

Synthetically generating inner speech training data

ActiveUS12586568B2SensorsDiagnostic recording/measuringSpeech trainingAcoustics
Methods and systems are disclosed for synthetically generating inner speech training data. The methods and systems access a collection of overt speech signals representing phonemes, phoneme sounds, words or phrases spoken at least partially using overt speech. The methods and systems transform the collection of overt speech signals into inner speech training data comprising electromyograph (EMG) data representing inner speech corresponding to the phonemes, phoneme sounds, words or phrases spoken at least partially using the overt speech. The methods and systems train a machine learning model to decode inner speech signals into a set of corresponding phonemes, phoneme sounds, words or phrases based on the inner speech training data.
Owner:SNAP INC

Adaptive self-trained computer engines with associated databases and methods of use thereof

In some embodiments, the present invention provides for an exemplary computer system which includes at least the following components: an adaptive self-trained computer engine programmed, during a training stage, to electronically receive an initial speech audio data generated by a microphone of a computing device; dynamically segment the initial speech audio data and the corresponding initial text into a plurality of user phonemes; dynamically associate a plurality of first timestamps with the plurality of user-specific subject-specific phonemes; and, during a transcription stage, electronically receive to-be-transcribed speech audio data of at least one user; dynamically split the to-be transcribed speech audio data into a plurality of to-be-transcribed speech audio segments; dynamically assigning each timestamped to-be-transcribed speech audio segment to a particular core of the multi-core processor; and dynamically transcribing, in parallel, the plurality of timestamped to-be-transcribed speech audio segments based on the user-specific subject-specific speech training model.
Owner:VOXSMART LTD

Methods to assist verbal communication for both listeners and speakers

Methods implemented in a system utilizing computing programs for a speaker and a listener in conversation are provided. Aspects include (i) a reminder provisioner for a speaker which is triggered according to speed, pitch or volume of the speaker's speech, (ii) a speech training provisioner for a speaker, and (iii) an application which records and plays back difficult conversation to understand.
Owner:SATO HIROKI

Speech training data generation method, device, equipment, medium and program product

PendingCN122511223Aachieve recognizabilityImplement labelingSpeech trainingTimestamp
This application relates to a method, apparatus, device, medium, and program product for generating speech training data. The method includes: acquiring initial training data, the initial training data comprising at least one audio-text pair; processing the audio and text in each audio-text pair to obtain first timestamp information corresponding to each word in the audio-text pair; based on the audio in each audio-text pair, obtaining each sub-language event and second timestamp information corresponding to the sub-language event; based on the first timestamp information corresponding to each word and the second timestamp information corresponding to the sub-language event, generating text insertion positions corresponding to each sub-language event; and based on the text insertion positions corresponding to each sub-language event and the initial training data, obtaining target training data. This method can reduce costs.
Owner:MOORE THREADS TECH CO LTD

An end-to-end model training method and device, computer equipment and storage medium

ActiveCN114882874BSpeech recognitionSpeech trainingEngineering
The embodiment of the application belongs to the technical field of speech recognition in artificial intelligence, and relates to an end-to-end model training method and device applied to speech recognition, computer equipment and a storage medium. The output of an acoustic model is taken as expanded text of audio training data, and the expanded text and audio annotation text are taken as language model input to train the speech recognition model, thereby effectively solving the problem of too limited annotation text content in a traditional speech training set, enabling the language model of the speech recognition model to learn more comprehensive information, thereby effectively improving the recognition accuracy of the speech recognition model, and to a certain extent, reducing the coupling degree of acoustic information and language information in the end-to-end model, improving the robustness of the entire model in different scenes, especially when recognizing speech in different fields, avoiding the problem of a large decrease in accuracy when changing application scenarios, and increasing the flexibility of the model in actual use and deployment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition method and system based on large model and speech synthesis engine

PendingCN121306105ASpeech recognitionSpeech synthesisSpeech trainingAutomatic speech
The invention relates to the technical field of speech recognition, provides a speech recognition method and system based on a large model and a speech synthesis engine, and greatly improves the recognition accuracy of ASR recognition by combining a large language model language with a new generation of speech synthesis engine. A large language model is utilized to generate a large batch of high-quality corpora, and a speech synthesis engine is utilized to obtain a large amount of high-quality speech training data for training an automatic speech recognition model and checking and correcting a recognition result at the same time. A training closed loop of generating high-quality text training data, synthesizing natural speech, recognizing natural speech, finding errors, correcting the errors and generating training data is realized, so that an automatic speech recognition model can continuously, professionally and automatically reinforce learning at low cost; and furthermore, a high-accuracy and high-quality speech recognition result is output under semantic errors and complex scenes.
Owner:GUANGZHOU AUSUN INFORMATION TECH CO LTD

Speech processing model training method, speech processing method, and speech translation method

Embodiments of the present specification provide a speech processing model training method, a speech processing method and a speech translation method. The speech processing model training method comprises: determining first speech training data corresponding to a speech processing task and second speech training data corresponding to a speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task; training a speech processing network layer in an initial speech processing model according to the second speech training data to obtain a trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask; and performing model training on the trained initial speech processing model according to the first speech training data to obtain a target speech processing model.
Owner:ALIBABA (CHINA) CO LTD

Cognitive and perceptual training kit

ActiveCN309950983SSpeech trainingPhysical medicine and rehabilitation
1. Name of the designed product: cognitive and perceptual training box. 2. Use of the designed product: for sensory training, cognitive training and speech training. 3. Design points of the designed product: in shape. 4. Picture or photo best indicating the design points: perspective view 1.
Owner:CHINA REHABILITATION RES CENT

A Chinese speech recognition method, system, storage medium and electronic device

The present application relates to a kind of Chinese speech recognition method, system, storage medium and electronic equipment, comprising: based on multiple Chinese speech training samples, the original CTC coding network added with a fine-grained loss module and two intermediate layer loss modules is trained, obtains the first Chinese speech recognition model, and delete the fine-grained loss module and the two intermediate layer loss modules in the first Chinese speech recognition model, obtain target Chinese speech recognition model;The Chinese speech data to be identified is input into the target Chinese speech recognition model, and Chinese speech recognition result is obtained.The present application adds the loss calculation of multilevel multi-granularity, so that CTC coding network can extract more rich and varied speech feature information, while not affecting model inference speed and model complexity, improve the accuracy of Chinese speech recognition.
Owner:BEIJING SHUMEI SHIDAI TECH CO LTD +1

Speech training noise adding system and method based on hybrid noise generation model

ActiveCN121708909BNoise generationSpeech training
The purpose of this disclosure is to provide a speech training noise enhancement system and method based on a hybrid noise generation model, comprising: an input module, a noise environment enhancement module, a speech noise enhancement module, and an output module; wherein, the input module is used to acquire simplified description information of the noise environment and clean speech data to be enhanced; the noise environment enhancement module converts the simplified description information of the noise environment into a structured noise event sequence with temporal features; the speech noise enhancement module generates multi-source hybrid noise according to the noise event sequence and adds the multi-source hybrid noise to the clean speech data according to preset rules to obtain noisy speech data; the output module is used to output the noisy speech data for use in speech model training. This disclosure fully integrates the interactive features of superposition, cancellation, and interference of different noise sources in the time dimension, solving the technical bottleneck that traditional mixing methods can only achieve linear superposition.
Owner:GUANGDONG UNIV OF TECH

Method and apparatus for accelerating neural network model inference, electronic device and medium

Embodiments of the present disclosure disclose a method, device, electronic equipment and storage medium for accelerating inference of a neural network model, wherein the method comprises: obtaining image training data, text training data or speech training data; determining a first neural network model to be accelerated; converting a preset operation of a preset network layer in the first neural network model into a first operation to obtain a second neural network model, the first operation being used to simulate operation logic of a target operation; performing quantization-aware training on the second neural network model according to a preset bit width based on the image training data, the text training data or the speech training data to obtain a third neural network model after quantization; and converting the first operation of the third neural network model into the target operation to obtain an accelerated target neural network model corresponding to the first neural network model. Embodiments of the present disclosure simulate errors caused by simplified operations in the training process to ensure high model accuracy under the condition of reducing model calculation amount.
Owner:NANJING HORIZON INFORMATION TECHNOLOGY CO LTD

Speech training system for persons with speech disorders, speech training method for persons with speech disorders, communication support system for persons with speech disorders, communication support method for persons with speech disorders, analysis system for persons with speech disorders, analysis method for persons with speech disorders, program and recording medium

PendingJP2026084736AHealthcare managementReadingSpeech trainingSpeech rate
This system provides speech therapy for individuals with speech disorders, including children with disabilities, enabling them to easily and independently continue their training at home or elsewhere, without time or location constraints, even after discharge from the hospital or during outpatient visits. [Solution] The speech training support system for persons with speech disorders is configured to present a model voice converted to a slower speed using speech rate conversion technology to the person with a speech disorder who is the target of the speech training support, and then use speech rate conversion technology to convert the voice spoken by the person with a speech disorder, who imitates the model voice converted to a slower speed, back to the same speed as the model voice before conversion, and then present the high-speed converted voice to the person with a speech disorder or to the listener. By presenting the high-speed converted voice to the person with a speech disorder, training is conducted that focuses attention on the accuracy of articulation movements. The speed of the model voice converted to a slower speed is the fastest speed at which the person with a speech disorder can speak with accurate articulation without difficulty, and is between 1 / 2 and 1 times the speed of the model voice before conversion.
Owner:THE UNIV OF TOKYO +1

A speech recognition model training method, device, apparatus and storage medium

The application discloses a speech recognition model training method and device, equipment and a storage medium, relates to the technical field of speech recognition, and comprises the following steps: processing an initial audio signal by using a preset audio processing method, and obtaining initial speech training data based on an acoustic feature consistency constraint condition, an identification accuracy constraint condition and a teacher-student comparison constraint condition and a target feature sequence obtained by processing; processing the initial speech training data by using a speech synthesis center, processing obtained to-be-processed speech training data by using a preset acoustic scene simulation system, and obtaining to-be-enhanced speech training data; processing the to-be-enhanced speech training data based on a word level by using an attention enhancement mechanism to obtain word-enhanced speech training data, processing the word-enhanced speech training data based on a sentence level by using a comparison learning framework, training an initial speech recognition model by using obtained target speech training data, and obtaining a target speech recognition model. In this way, the efficiency of training the speech recognition model can be improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Synthetically generating internal speech training data

PendingCN122055777ASensorsDiagnostic recording/measuringSpeech trainingAcoustics
Methods and systems for synthetically generating internal speech training data are disclosed. The method and system access a set of dominant speech signals representing phonemes, phoneme sounds, words or phrases spoken at least in part using dominant speech. The methods and systems transform a set of dominant speech signals into internal speech training data including electromyogram (EMG) data representing internal speech corresponding to phonemes, phoneme sounds, words, or phrases spoken at least in part using dominant speech. The method and system train a machine learning model based on the internal speech training data to decode the internal speech signal into a corresponding phoneme, phoneme sound, word or phrase set.
Owner:SNAP INC

3D virtual image hearing and speech training method based on AI interaction

ActiveCN115910036BInput/output for user-computer interactionSpeech recognitionSpeech trainingHearing speech
The application discloses a 3D virtual image hearing speech training method based on AI interaction, which is a brand-new English pronunciation learning method combining hearing training and body action stimulation and oral (speaking) practice, and combines intonation hearing method and 3D virtual image, AI speech recognition and evaluation, can effectively improve the pronunciation level of English learners, make up for the shortage of pronunciation teaching in the classroom and the shortage of teachers' pronunciation conditions, solve the timeliness problem in classroom teaching, and more importantly, the method capable of operating on a mobile intelligent terminal is an active innovation for promoting education fairness and realizing common balanced development of education.
Owner:YUNNAN BEIFEI TECH CO LTD

Phoneme alignment model training and speech synthesis method and device, equipment and medium

ActiveCN120748366BSpeech synthesisSpeech trainingMedicine
The application relates to the technical field of speech synthesis, in particular to a phoneme alignment model training and speech synthesis method and device, equipment and a medium. The method comprises the following steps: performing convolution attention alignment on spectral feature information and text feature information obtained according to speech training data to obtain a first alignment matrix; performing monotone alignment search based on the first alignment matrix to generate a second alignment matrix, the second alignment matrix being a binary hard attention matrix; calculating a relative entropy loss according to the first alignment matrix and the second alignment matrix; extending the text feature information to the length of a mel spectrum frame according to the second alignment matrix, and performing linear transformation on the extended text feature information to generate predicted mel spectrum; calculating a mel loss according to the spectral feature information and the predicted mel spectrum; and training a phoneme alignment model according to the relative entropy loss and the mel loss until a preset convergence condition is met. The method can improve phoneme alignment accuracy and thus improve speech synthesis accuracy.
Owner:SHANGHAI PAIDI INTELLIGENT TECH CO LTD

A discrete speech representation system and method oriented to symbolic expression

ActiveCN121075304BSpeech recognitionSpeech synthesisSpeech trainingAcoustics
This invention relates to the field of discrete speech representation, and discloses a discrete speech representation system and method oriented towards symbolic expression. The system includes: a symbolic normalization module, a speech synthesis module, a speech discretization module, a repetition suppression module, a boundary anchoring module, a cross-modal generation module, and a controlled decoding module. This invention not only rapidly constructs large-scale, high-quality symbolic speech training corpora, but also significantly improves the reading accuracy and robustness of end-to-end speech systems in symbolic tasks through the trained model, providing directly deployable technical support for educational, scientific research, and engineering applications.
Owner:JINAN UNIVERSITY

Long speech recognition model training method, electronic device, and storage medium

ActiveCN115798460BSpeech recognitionPattern recognitionSpeech training
The application discloses a long speech recognition model training method, an electronic device and a storage medium, and the method comprises the following steps: obtaining long speech training data which is constructed, wherein the long speech training data comprises extracted acoustic input features, frame-level classification labels for training an endpoint detection model, and text labels for training a speech recognition model; and the endpoint detection model and the speech recognition model are jointly trained by using the long speech training data. According to the embodiment of the application, the endpoint detection model and the speech recognition model are jointly trained by obtaining the long speech training data which is constructed, and the related information provided by the recognition model is introduced to assist the training optimization of the endpoint detection model on the basis of optimizing the endpoint detection model, so that a complete joint optimization method is realized, and the recognition performance of the long speech link is effectively improved.
Owner:AISPEECH CO LTD

A method and system for generating a multi-dimensional speech training program

ActiveCN121687374BAchieve precise quantitative analysisImprove targetingSpeech trainingSpeech comprehension
The application discloses a kind of multi-dimension speech training plan generation method and system.The method first collects patient age, gender and speech sample, obtains seven-dimensional evaluation parameters including sound pressure, amplitude perturbation, maximum vocalization duration, fundamental frequency perturbation, vowel space area, tone impairment and speech comprehension score.Subsequently, according to the preset logic, these parameters are sequentially determined based on the parameters, and dynamically combine different training modules such as loudness, breath, pitch, vowel, glide, tone and consonant into personalized speech training plan.The application solves the problem of existing technology training scheme solidification through multi-dimensional evaluation and dynamic module matching, significantly improves the individualization degree and rehabilitation effect of speech training.
Owner:BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)

Speech processing method and speech processing model training method

Embodiments of the present specification provide a speech processing method and a speech processing model training method. The speech processing method comprises: determining a speech to be translated, wherein the speech to be translated is real-time determined speech data or speech segment data; inputting the speech to be translated into a target speech processing model to obtain a speech translation result corresponding to the speech to be translated, wherein the target speech processing model is obtained by model training of an initial speech processing model according to a plurality of speech segment training data, and the initial speech processing model is obtained by model training of a speech processing model according to speech training data; thereby using the target speech processing model to perform speech translation on the speech to be translated to obtain an accurate speech translation result, improving the accuracy of the speech processing result of the neural network model, and avoiding the problem that the speech translation result is inaccurate due to complex speech data.
Owner:ALIBABA (CHINA) CO LTD

Guangzhou-hybrid speech recognition method and device, computer equipment and storage medium

The embodiment of the application belongs to the technical field of artificial intelligence, is applied to the field of digital medical treatment, relates to a Cantonese-Mandarin mixed speech recognition method, and comprises the following steps: adding an identifier to Cantonese text in obtained Cantonese text data to obtain extended Cantonese text data; combining obtained Mandarin audio data, Mandarin text data, Cantonese audio data and the extended Cantonese text data into a mixed speech training set, inputting the mixed speech training set into a pre-constructed deep neural network model for training to obtain a mixed speech recognition model; obtaining to-be-recognized speech data, inputting the to-be-recognized speech data into the mixed speech recognition model, and outputting a speech recognition result. The application also provides a Cantonese-Mandarin mixed speech recognition device, a computer device and a storage medium. In addition, the application also relates to the technology of blockchains, and the mixed speech training set can be stored in a blockchain. The application can distinguish whether the currently recognized speech type is Cantonese or Mandarin.
Owner:PING AN TECH (SHENZHEN) CO LTD