Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

46 results about "Speech training" patented technology

Speech training method and system based on artificial intelligence and application

PendingCN120600053ASpeech recognitionTeaching apparatusSpeech trainingMultimodal data
The invention discloses a speech training method and system based on artificial intelligence and application, and relates to the field of speech training, and the method comprises the steps: obtaining multi-modal data in real time when a user reads each training statement in each speech training process; obtaining a real-time multi-dimensional speech evaluation result according to the multi-modal data, and judging whether a next training statement exists in a speech training scheme in each speech training process or not; if not, generating a multi-dimensional speech evaluation result of each speech training process; if yes, the difficulty of the next training statement in the speech training scheme is adaptively and dynamically adjusted according to the real-time multi-dimensional speech evaluation result; and outputting the next training statement according to the adjusted speech training scheme, and repeating the steps until the next training statement does not exist in each speech training. The speech training scheme can be adaptively adjusted in real time in the speech training process, and the user experience, the participation degree and the training effect are remarkably improved.
Owner:CHANGSHA LIANYU TECHNOLOGY CO LTD

Phoneme alignment model training and speech synthesis method and device, equipment and medium

ActiveCN120748366ASpeech synthesisSpeech trainingMedicine
The invention relates to the technical field of speech synthesis, in particular to a phoneme alignment model training and speech synthesis method and device, equipment and a medium. The method comprises the following steps: performing convolution attention alignment on spectrum feature information and text feature information obtained according to voice training data to obtain a first alignment matrix; monotonic alignment search is executed based on the first alignment matrix to generate a second alignment matrix, wherein the second alignment matrix is a binary hard attention matrix; calculating relative entropy loss according to the first alignment matrix and the second alignment matrix; expanding the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing linear transformation on the expanded text feature information to generate a predicted Mel spectrum; calculating Mel loss according to the spectrum feature information and the predicted Mel spectrum; and training the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met. By adopting the method, the phoneme alignment accuracy can be improved, and the speech synthesis accuracy is further improved.
Owner:SHANGHAI PAIDI INTELLIGENT TECH CO LTD

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Adversarial training of keyword spotting to minimize TTS data overfitting

PendingUS20260051318A1Biological modelsSpeech recognitionSpeech trainingHidden layer
A method includes receiving training utterances that include non-synthetic speech training utterances and synthetic speech utterances. For each training utterance, the method includes processing, using a memorized neural network, a corresponding sequence of input audio frames to generate a hotword detection output indicating a likelihood the training utterance includes a hotword, determining a first loss based on the hotword detection output, obtaining a hidden layer feature vector for each corresponding input audio frame; processing, using a speech classification model, the hidden layer feature vectors to predict a classification output for the training utterance; and determining an adversarial loss based on the classification output predicted for the training utterance. The method also includes training the memorized neural network on the first losses and the adversarial losses to teach the memorized neural network to learn how to detect the hotword in audio and prevent overfitting of the synthetic speech training utterances.
Owner:GDM HOLDING LLC

An active five-tone speech therapy system

This invention discloses an active five-tone speech therapy system, primarily targeting individuals with speech dysfunction after stroke. Its core active five-tone speech therapy has been proven in clinical trials to effectively improve spontaneous speech, auditory comprehension, repetition, and naming functions in patients with subacute and chronic aphasia following stroke. Building upon this therapy, this invention adds a module for collecting and differentiating the patient's five internal organ syndrome elements, creating a novel, online, therapist-free remote diagnosis and treatment software that integrates "patient symptom and sign information collection – five internal organ syndrome element assessment – ​​five-tone repertoire recommendation – active five-tone speech therapy – speech collection during training, accuracy and pronunciation standard assessment, training difficulty adjustment, scale evaluation, and training efficacy feedback." This invention can improve the applicability of Western classical melodic intonation therapy to Chinese aphasia patients, mobilize patients' subjective initiative, provide precise five-tone music recommendations and individualized treatment for patients with different syndrome types, enhance the rehabilitation effect and efficiency of existing speech training, reduce the workload of therapists, alleviate the medical and economic burden on patients, and lay the foundation for the development of subsequent active five-tone speech therapy remote diagnosis and treatment equipment.
Owner:FUJIAN UNIV OF TRADITIONAL CHINESE MEDICINE

Speech communication aid trainer

ActiveCN309803572SSpeech trainingAcoustics
1. The name of the design product: speech communication auxiliary training device. 2. The use of the design product: for speech training or auxiliary communication for people with language barriers. 3. The design points of the design product: in shape. 4. The picture or photo that best shows the design points: perspective view 1.
Owner:谭文斯

Speech recognition method based on bimodal mixed contrast enhancement, electronic device, chip, storage medium and program product

The application provides a speech recognition method based on bimodal mixed contrast enhancement, an electronic device, a chip, a storage medium and a program product; the method comprises the following steps: acquiring multi-modal speech training data; a bimodal mixed contrast learning model is constructed and trained, wherein samples with the same emotional label under the same mode are taken as positive sample pairs, samples with different emotional labels are taken as negative sample pairs, intra-modal contrast loss is calculated to enhance the ability of the bimodal mixed contrast learning model to distinguish intra-modal fine-grained emotional features; samples with the same emotional label under different modes are taken as positive sample pairs, samples with different emotional labels under different modes are taken as negative sample pairs, inter-modal contrast loss is calculated to realize the alignment and complementarity of different modal feature spaces; the intra-modal contrast loss and the inter-modal contrast loss are fused to obtain multi-modal mixed contrast loss, and the model parameters of the bimodal mixed contrast learning model are optimized; based on the features extracted by the bimodal mixed contrast learning model, a downstream speech recognition task is trained.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Multi-dimensional AI platform intelligent voice response system using voice synthesis technology

PendingCN120808785ASpeech recognitionSpeech synthesisSpeech trainingText entry
The invention relates to the technical field of voice synthesis, in particular to a multi-dimensional AI platform intelligent voice response system using the voice synthesis technology, which screens out matched historical user question voices according to text semantic similarity corresponding to current user question voices and historical user question voices, and sends the matched historical user question voices to a user terminal. Obtaining a voice training set and the weight of each element in the voice training set according to the voice feature similarity between the question voice of the current user and the question voice of the matched historical user in combination with the user score value of the manual reply voice, training an acoustic model, obtaining a reply text corresponding to the question voice of the current user, and obtaining the question voice of the current user; and inputting into the trained acoustic model, generating a reply voice signal, and then outputting to the current user. According to the method, the voice training set is screened and constructed from historical manual reply voices, so that an acoustic model can learn a more natural and smooth voice synthesis mode, and the reply voice signal contains rich voice features and expression modes.
Owner:ROPEOK TECHNOLOGY GROUP CO LTD

Verbal communication aid training device

ActiveCN309812665SSpeech trainingEngineering
1. Name of the product in this design: Speech Communication Assistive Training Device. 2. Purpose of this design: To provide speech training or assist communication for people with language impairments. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: 3D view 1.
Owner:谭文斯

A method, apparatus, device and medium for training a speech generation model

The application belongs to the field of artificial intelligence, and relates to a training method of a speech generation model, comprising the following steps: obtaining reference timbre spectrum, phoneme information and speech spectrum of a target object; training a preset initial speech generation model based on the reference timbre spectrum, the phoneme information and the speech spectrum to obtain model parameters; and adjusting parameters of a multi-timbre feature extraction network, a phoneme feature extraction network, a prosody feature discretization network, a time sequence alignment module, an attention fusion module and a speech reconstruction decoding network of the initial speech generation model based on the model parameters to construct the speech generation model. The application also provides an apparatus, a device and a medium. In addition, the application also relates to blockchain technology, and speech training data and model parameters can be stored in a blockchain. The application can realize decoupling of timbre and prosody information, and flexibly adjust the timbre and prosody information to generate synthesized speech with diversity and flexibility.
Owner:PING AN TECH (SHENZHEN) CO LTD

Synthetically generating inner speech training data

ActiveUS12586568B2SensorsDiagnostic recording/measuringSpeech trainingAcoustics
Methods and systems are disclosed for synthetically generating inner speech training data. The methods and systems access a collection of overt speech signals representing phonemes, phoneme sounds, words or phrases spoken at least partially using overt speech. The methods and systems transform the collection of overt speech signals into inner speech training data comprising electromyograph (EMG) data representing inner speech corresponding to the phonemes, phoneme sounds, words or phrases spoken at least partially using the overt speech. The methods and systems train a machine learning model to decode inner speech signals into a set of corresponding phonemes, phoneme sounds, words or phrases based on the inner speech training data.
Owner:SNAP INC

Speech enhancement method and device based on time-frequency domain feature fusion, and electronic equipment

PendingCN120636434ASpeech analysisSpeech trainingTime domain
The invention discloses a speech enhancement method and device based on time-frequency domain feature fusion and electronic equipment, and the method comprises the steps: carrying out the preprocessing of an initial speech training data set, and obtaining a target speech training data set; performing feature extraction on the target voice training data set to obtain candidate voice features; normalizing the candidate voice features to obtain target voice features; performing feature fusion on the target time domain feature and the target frequency domain feature to generate candidate fusion features; inputting the candidate fusion features into a multi-scale convolution feature enhancement module for feature enhancement, and generating to-be-trained fusion features; inputting the to-be-trained fusion features into the initial speech enhancement model for training to obtain a target speech enhancement model; and inputting the target fusion feature of the to-be-enhanced noisy speech into the target speech enhancement model to generate a target enhanced speech. The method can suppress noise interference, improves the speech enhancement effect in a complex noise scene, and can be widely applied to the technical field of speech enhancement.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Adaptive self-trained computer engines with associated databases and methods of use thereof

In some embodiments, the present invention provides for an exemplary computer system which includes at least the following components: an adaptive self-trained computer engine programmed, during a training stage, to electronically receive an initial speech audio data generated by a microphone of a computing device; dynamically segment the initial speech audio data and the corresponding initial text into a plurality of user phonemes; dynamically associate a plurality of first timestamps with the plurality of user-specific subject-specific phonemes; and, during a transcription stage, electronically receive to-be-transcribed speech audio data of at least one user; dynamically split the to-be transcribed speech audio data into a plurality of to-be-transcribed speech audio segments; dynamically assigning each timestamped to-be-transcribed speech audio segment to a particular core of the multi-core processor; and dynamically transcribing, in parallel, the plurality of timestamped to-be-transcribed speech audio segments based on the user-specific subject-specific speech training model.
Owner:VOXSMART LTD

Methods to assist verbal communication for both listeners and speakers

Methods implemented in a system utilizing computing programs for a speaker and a listener in conversation are provided. Aspects include (i) a reminder provisioner for a speaker which is triggered according to speed, pitch or volume of the speaker's speech, (ii) a speech training provisioner for a speaker, and (iii) an application which records and plays back difficult conversation to understand.
Owner:SATO HIROKI

Speech training data generation method, device, equipment, medium and program product

PendingCN122511223Aachieve recognizabilityImplement labelingSpeech trainingTimestamp
This application relates to a method, apparatus, device, medium, and program product for generating speech training data. The method includes: acquiring initial training data, the initial training data comprising at least one audio-text pair; processing the audio and text in each audio-text pair to obtain first timestamp information corresponding to each word in the audio-text pair; based on the audio in each audio-text pair, obtaining each sub-language event and second timestamp information corresponding to the sub-language event; based on the first timestamp information corresponding to each word and the second timestamp information corresponding to the sub-language event, generating text insertion positions corresponding to each sub-language event; and based on the text insertion positions corresponding to each sub-language event and the initial training data, obtaining target training data. This method can reduce costs.
Owner:MOORE THREADS TECH CO LTD

An end-to-end model training method and device, computer equipment and storage medium

ActiveCN114882874BSpeech recognitionSpeech trainingEngineering
The embodiment of the application belongs to the technical field of speech recognition in artificial intelligence, and relates to an end-to-end model training method and device applied to speech recognition, computer equipment and a storage medium. The output of an acoustic model is taken as expanded text of audio training data, and the expanded text and audio annotation text are taken as language model input to train the speech recognition model, thereby effectively solving the problem of too limited annotation text content in a traditional speech training set, enabling the language model of the speech recognition model to learn more comprehensive information, thereby effectively improving the recognition accuracy of the speech recognition model, and to a certain extent, reducing the coupling degree of acoustic information and language information in the end-to-end model, improving the robustness of the entire model in different scenes, especially when recognizing speech in different fields, avoiding the problem of a large decrease in accuracy when changing application scenarios, and increasing the flexibility of the model in actual use and deployment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition method and system based on large model and speech synthesis engine

PendingCN121306105ASpeech recognitionSpeech synthesisSpeech trainingAutomatic speech
The invention relates to the technical field of speech recognition, provides a speech recognition method and system based on a large model and a speech synthesis engine, and greatly improves the recognition accuracy of ASR recognition by combining a large language model language with a new generation of speech synthesis engine. A large language model is utilized to generate a large batch of high-quality corpora, and a speech synthesis engine is utilized to obtain a large amount of high-quality speech training data for training an automatic speech recognition model and checking and correcting a recognition result at the same time. A training closed loop of generating high-quality text training data, synthesizing natural speech, recognizing natural speech, finding errors, correcting the errors and generating training data is realized, so that an automatic speech recognition model can continuously, professionally and automatically reinforce learning at low cost; and furthermore, a high-accuracy and high-quality speech recognition result is output under semantic errors and complex scenes.
Owner:GUANGZHOU AUSUN INFORMATION TECH CO LTD

Speech processing model training method, speech processing method, and speech translation method

Embodiments of the present specification provide a speech processing model training method, a speech processing method and a speech translation method. The speech processing model training method comprises: determining first speech training data corresponding to a speech processing task and second speech training data corresponding to a speech processing subtask, wherein the speech processing subtask is a subtask of the speech processing task; training a speech processing network layer in an initial speech processing model according to the second speech training data to obtain a trained initial speech processing model, wherein the speech processing network layer is related to the speech processing subtask; and performing model training on the trained initial speech processing model according to the first speech training data to obtain a target speech processing model.
Owner:ALIBABA (CHINA) CO LTD

Cognitive and perceptual training kit

ActiveCN309950983SSpeech trainingPhysical medicine and rehabilitation
1. Name of the designed product: cognitive and perceptual training box. 2. Use of the designed product: for sensory training, cognitive training and speech training. 3. Design points of the designed product: in shape. 4. Picture or photo best indicating the design points: perspective view 1.
Owner:CHINA REHABILITATION RES CENT

A Chinese speech recognition method, system, storage medium and electronic device

The present application relates to a kind of Chinese speech recognition method, system, storage medium and electronic equipment, comprising: based on multiple Chinese speech training samples, the original CTC coding network added with a fine-grained loss module and two intermediate layer loss modules is trained, obtains the first Chinese speech recognition model, and delete the fine-grained loss module and the two intermediate layer loss modules in the first Chinese speech recognition model, obtain target Chinese speech recognition model;The Chinese speech data to be identified is input into the target Chinese speech recognition model, and Chinese speech recognition result is obtained.The present application adds the loss calculation of multilevel multi-granularity, so that CTC coding network can extract more rich and varied speech feature information, while not affecting model inference speed and model complexity, improve the accuracy of Chinese speech recognition.
Owner:BEIJING SHUMEI SHIDAI TECH CO LTD +1

Speech and singing voice synthesis method, training method, and related apparatus

PCT designated stageWO2025247230A1Speech synthesisSpeech trainingSynthesis methods
Disclosed in the present disclosure are a speech and singing voice synthesis method, a training method, and a related apparatus. The training method comprises: acquiring speech training data; acquiring singing voice training data; converting text data into text embedding, and converting speech data into a speech embedding; concatenating the text embedding and the speech embedding to form encoded speech training data; segmenting lyrics data into phrases and converting the phrases into phrase embeddings, segmenting singing voice data into singing voice segments corresponding to the phrases and converting the singing voice segments into singing voice segment embeddings, extracting, from musical score data, pitch sequences corresponding to the phrases and / or singing voice segments, and converting the pitch sequences into pitch embeddings; concatenating the phrase embeddings, the singing voice segment embeddings, and the pitch embeddings to form encoded singing voice training data; and inputting the encoded speech training data and the encoded singing voice training data into a speech and singing voice synthesis module of an initial model for training, to obtain a target model upon training.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

Emotional voice conversion model training method and emotional voice conversion method

PendingCN120612965AKernel methodsBiological modelsSpeech trainingAcoustics
The invention provides a training method of an emotional voice conversion model and an emotional voice conversion method, which are applied to the technical field of voice signal processing, and the training method comprises the following steps: obtaining an emotional voice training set which comprises a plurality of emotional voice samples; for each emotional voice sample, processing the emotional voice sample by using a content coding network to obtain a content coding feature, the content coding feature representing voice content corresponding to the emotional voice sample and key information related to the speaker; processing the emotional voice sample by using an emotional extraction network to obtain an emotional extraction feature, the emotional extraction feature representing an emotional corresponding to the emotional voice sample; processing the content coding features and the emotion extraction features by using a decoding vocoder to obtain predicted emotion voice information; and iteratively adjusting network parameters of the initial emotional voice conversion model according to the emotional voice samples, the content coding features and the emotional extraction features to obtain a trained target emotional voice conversion model.
Owner:TIANJIN UNIV +1

A speech synthesis method, apparatus and device for speech synthesis

ActiveCN113889070BSpeech synthesisSpeech trainingSynthesis methods
Embodiments of the present application provide a speech synthesis method and device and a device for speech synthesis, applied to a terminal device. The method comprises: training a multi-person acoustic model based on multi-person speech training data, the multi-person acoustic model comprising an encoder, a prosody prediction network, a duration prediction network, and a decoder; the acoustic features output by the decoder comprising fundamental frequency features and mel spectrum features; adaptively training the multi-person acoustic model based on single-person speech training data of a target speaker to obtain a single-person acoustic model of the target speaker; performing parameter quantization processing on the single-person acoustic model to obtain a target single-person acoustic model; and synthesizing audio data of acoustic features of the target speaker using the target single-person acoustic model and text to be synthesized. Embodiments of the present application ensure model effectiveness, and make the target single-person acoustic model obtained by training applicable to offline devices with limited computing power and storage space.
Owner:BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD

Speech training noise adding system and method based on hybrid noise generation model

ActiveCN121708909BNoise generationSpeech training
The purpose of this disclosure is to provide a speech training noise enhancement system and method based on a hybrid noise generation model, comprising: an input module, a noise environment enhancement module, a speech noise enhancement module, and an output module; wherein, the input module is used to acquire simplified description information of the noise environment and clean speech data to be enhanced; the noise environment enhancement module converts the simplified description information of the noise environment into a structured noise event sequence with temporal features; the speech noise enhancement module generates multi-source hybrid noise according to the noise event sequence and adds the multi-source hybrid noise to the clean speech data according to preset rules to obtain noisy speech data; the output module is used to output the noisy speech data for use in speech model training. This disclosure fully integrates the interactive features of superposition, cancellation, and interference of different noise sources in the time dimension, solving the technical bottleneck that traditional mixing methods can only achieve linear superposition.
Owner:GUANGDONG UNIV OF TECH

Prosody encoder training method, speech conversion method and related products

ActiveCN115294959BSpeech recognitionSpeech synthesisSpeech trainingAcoustics
An embodiment of the present invention provides a method for training a prosody encoder, comprising: obtaining multi-level coding features of speech training data from the training data based on a speech recognition model; extracting the multi-level coding features based on multiple coding output layers of the encoder of the speech recognition model; and training the prosody encoder using the multi-level coding features of the training data, so that the trained prosody encoder can be used to extract prosody information. The method of the present invention enables the prosody encoder to accurately extract prosody features, thereby significantly improving the speech conversion effect and providing a better user experience. Furthermore, embodiments of the present invention provide a speech conversion method, a prosody encoder training device, a speech conversion device, an apparatus, and a computer-readable storage medium.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Sub-band unvoiced and voiced sound parameter quantization method and system based on double-resolution codebook

InactiveCN121054007ASpeech analysisSpeech trainingAcoustics
The invention discloses a sub-band unvoiced and voiced sound parameter quantization method and system based on a dual-resolution codebook. The method comprises the following steps: constructing an unvoiced and voiced sound parameter quantization first-level codebook based on a voice training set, and constructing a local high-resolution second-level codebook by analyzing a cell vector frequency corresponding to a first-level codeword; counting a transfer relationship between the unvoiced and voiced sound modes of front and back frames, and generating a likelihood matrix; in an actual coding process, an optimal first-level codeword is selected as a rough matching result through a weighted mean square error criterion, maximum likelihood search is executed based on a previous frame coding result and a likelihood matrix, and an optimal codeword is selected from a current frame second-level codebook as a final quantization result. The system comprises a parameter extraction module, a codebook generation module, a likelihood modeling module, a real-time quantization coding module and a decoding reconstruction module, and can significantly improve the expression precision of sub-band unvoiced and voiced sound parameters and the naturalness of speech synthesis on the premise of not increasing the coding rate.
Owner:SUZHOU WUTONG MICRO ACOUSTIC TECH CO LTD

Method and apparatus for accelerating neural network model inference, electronic device and medium

Embodiments of the present disclosure disclose a method, device, electronic equipment and storage medium for accelerating inference of a neural network model, wherein the method comprises: obtaining image training data, text training data or speech training data; determining a first neural network model to be accelerated; converting a preset operation of a preset network layer in the first neural network model into a first operation to obtain a second neural network model, the first operation being used to simulate operation logic of a target operation; performing quantization-aware training on the second neural network model according to a preset bit width based on the image training data, the text training data or the speech training data to obtain a third neural network model after quantization; and converting the first operation of the third neural network model into the target operation to obtain an accelerated target neural network model corresponding to the first neural network model. Embodiments of the present disclosure simulate errors caused by simplified operations in the training process to ensure high model accuracy under the condition of reducing model calculation amount.
Owner:NANJING HORIZON INFORMATION TECHNOLOGY CO LTD

Parrot Cage

To provide a parrot cage which allows a player to reproduce a recording, realizes speech training for a parrot, and allows the parrot to learn language imitation at any time without the need for human cooperation. [Solution] The cage comprises a cage body having an accommodation cavity inside for the parrot to move around in, a partition plate, a first gripping rod with conductive parts on both opposing ends, and a training unit including the conductive part, wherein a player is provided on the partition plate, the conductive part includes an upright rod, a slide bush and a support rod, the first end of the upright rod is fixed to the bottom of the accommodation cavity, a first guide rod is provided on the second end of the upright rod, the slide bush is externally fitted onto the second end of the upright rod, a second guide rod is provided within the slide bush, the slide bush is connected to the first gripping rod via the support rod, and the slide bush is movable along a first preset path.
Owner:ゴン スーイン

Speech recognition in a patient monitor

PCT designated stageWO2025228870A1Speech recognitionSpeech trainingMedical terminology
A method of training and using a speech recognition model in a patient monitor. The speech recognition model is trained using integrated speech training data that comprises a combination of professional speech training data associated with medical terms and general speech training data. Voice inputs received by the patient monitor are processed using the trained speech recognition model to produce text outputs comprising estimates of the words in the voice input. Corresponding actions for the patient monitor are then determined responsive to the text outputs.
Owner:KONINKLIJKE PHILIPS NV +1

Method for training a speech recognition model and speech recognition method

ActiveCN115691475BBiological modelsSpeech recognitionSpeech trainingSpeech sound
The application relates to a method for training a speech recognition model, comprising: providing a speech training data set comprising a plurality of speech data and speech labels corresponding to each speech data; providing a speech recognition model to be trained, the speech recognition model comprising a convolutional neural network, a first fully connected network, a recurrent neural network and a second fully connected network coupled in cascade, wherein each network comprises one or more network layers with a parameter matrix; wherein the speech recognition model is used to process speech data to generate corresponding speech recognition results; and training the speech recognition model using the speech training data set, so that after training, the parameter matrices of at least two adjacent network layers in the speech recognition model satisfy a predetermined constraint condition; and so that the accuracy of the speech recognition results of the speech recognition model on the speech data calculated using at least one loss function satisfies a predetermined recognition target.
Owner:MONTAGE TECH CHENGDU CO LTD