Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Voice Training" patented technology

A variety of techniques used to help individuals utilize their voice for various purposes and with minimal use of muscle energy.

Voice training system and data training method after complete denture repair

The invention belongs to the technical field of electronic teaching aids, and discloses a voice training system and a data training method after complete denture repair, and the system comprises a feedback analysis and decision device which is used for carrying out the error classification of a used object, and obtaining the type of a high-frequency problem through a clustering algorithm; the dynamic content generation engine is used for recording a historical error pattern of the use object, presetting pronunciation training logic and generating personalized training content; the intelligent difficulty adjusting device is used for adopting a dynamic difficulty algorithm: scoring according to a result of a use object, and automatically adjusting the level difficulty; and the adaptive content generator is used for analyzing the pronunciation content of the use object, identifying wrong phonemes or grammar problems, and generating training content according with difficulty in real time according to the current level and error type of the use object. The method can help a user to recover clear voice as soon as possible, improves the use enthusiasm, and can solve the problems that personalized training is not available, real-time monitoring and feedback cannot be performed, and pronunciation cannot be corrected in time.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Phoneme alignment model training and speech synthesis method and device, equipment and medium

The invention relates to the technical field of speech synthesis, in particular to a phoneme alignment model training and speech synthesis method and device, equipment and a medium. The method comprises the following steps: performing convolution attention alignment on spectrum feature information and text feature information obtained according to voice training data to obtain a first alignment matrix; monotonic alignment search is executed based on the first alignment matrix to generate a second alignment matrix, wherein the second alignment matrix is a binary hard attention matrix; calculating relative entropy loss according to the first alignment matrix and the second alignment matrix; expanding the text feature information to a Mel spectrum frame length according to the second alignment matrix, and performing linear transformation on the expanded text feature information to generate a predicted Mel spectrum; calculating Mel loss according to the spectrum feature information and the predicted Mel spectrum; and training the phoneme alignment model according to the relative entropy loss and the Mel loss until a preset convergence condition is met. By adopting the method, the phoneme alignment accuracy can be improved, and the speech synthesis accuracy is further improved.
Owner:SHANGHAI PAIDI INTELLIGENT TECH CO LTD

Voice training scheme generation method and system

The invention provides a voice training scheme generation method and system. The method is applied to the technical field of voice training, and comprises the following steps: acquiring voice data, voice signals and physiological feature data related to voice generation, and performing preprocessing; extracting time domain and frequency domain features from the preprocessed data, wherein the time domain features comprise fundamental frequency, peak amplitude, signal energy and the like; the frequency domain characteristics comprise a harmonic noise ratio, a fast Fourier transform median, energy distribution of a voice signal in each frequency band, a frequency spectrum centroid, a width of the frequency spectrum centroid and the like; screening high-discrimination features from the initially extracted time domain features and frequency domain features through backward / forward sequence feature selection to construct an optimized feature set, and inputting the optimized feature set into an improved S4 neural network model for training; and analyzing the voice type based on the trained model, dynamically generating a personalized voice training scheme, and performing dynamic adjustment. According to the invention, the accuracy, adaptability and robustness of the voice training scheme are improved.
Owner:BEIJING FENGSHANGTIANCHENG CULTURAL DEV CO LTD

Controllable, natural paralinguistics for text to speech synthesis

A speech recognition module receives training data of speech and creates a representation for individual words, non-words, phonemes, and any combination. A set of speech processing detectors analyze the training data of speech from humans communicating. The set of speech processing detectors detect speech parameters that are indicative of paralinguistic effects on top of enunciated words, phonemes, and non-words in the audio stream. One or more machine learning models undergo supervised machine learning on their neural network to train on how to associate one or more mark-up markers with a textual representation, for each individual word, individual non-word, individual phoneme, and any combinations of these, that was enunciated with a particular paralinguistic effect. Each mark-up marker can correspond to its own paralinguistic effect.
Owner:SRI INTERNATIONAL

Endoscope with Voice Control

An endoscope with voice control. A data processor for the endoscope obtains voice training data for a specific surgeon, and / or information indicating spoken utterances associated with a surgical procedure for which the endoscope is to be used. Voice utterances signals are received from a surgeon during the surgery. Speech recognition technology decodes the voice utterance signals into commands to control the endoscope, based at least in part on the data obtained from the database, and control commands are issued to implement the decoded voice utterances. Concurrently with the accepting and decoding of voice utterances, the processor accepts and decodes control signals from at least one other input device. The processor issues control commands to components of the endoscope to implement the decoded other input control signals. The voice utterances are accepted, decoded, and executed concurrently with the other input control signals.
Owner:PSIP2 LLC

Speech recognition model acquisition method and device, computer equipment, readable storage medium and program product

The invention relates to a voice recognition model acquisition method and device, computer equipment, a computer readable storage medium and a computer program product, relates to the technical field of voice recognition, and can improve the training precision of a voice recognition model and the generalization ability of the model in an unlabeled application scene. The method comprises the following steps: acquiring unmarked voice training data; performing data cleaning on the voice training data to obtain cleaned voice training data; obtaining a pre-training voice recognition model, and performing pseudo-tag prediction on the cleaned voice training data through the model to obtain first voice training data with a tag; performing text error correction on a pseudo tag in the first voice training data with the tag, and performing voice correction on voice training data in the first voice training data with the tag to obtain second voice training data with the tag; and adjusting the pre-trained speech recognition model according to the second speech training data with the label to obtain a target speech recognition model.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Voice training partner method and device, electronic equipment and storage medium

The embodiment of the invention discloses a voice partner training method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a universal cue word of voice partner training and a current cue word corresponding to a current round of dialogue in a simulation service scene in a current dial test; the current broadcast text is accurately and conveniently generated based on the general prompt word, the current prompt word and the preset text generation model, so that the generation efficiency of the current broadcast text is improved; if it is detected that the current broadcast text does not contain the last sentence identifier, determining a current broadcast voice corresponding to the current round of dialogue based on the current broadcast text, the current prompt word and a preset voice generation model; the current broadcast voice is the voice with emotions, so that the emotional effect of voice broadcast is realized; determining a current response text corresponding to the current round of dialogue based on the received current response voice of the user and a preset voice recognition model; and based on the current broadcast text, the current response text and the current cue word, determining a next cue word corresponding to the next round of dialogue.
Owner:AGRICULTURAL BANK OF CHINA

Speaker verification based adaptive margin optimization method, system, and electronic device

This invention provides an adaptive margin optimization method, system, and electronic device based on speaker verification. The method includes: inputting speech training data including various speech durations into a speaker verification model; determining the loss function of the speaker verification model; adaptively optimizing the margin parameters of the loss function based on the speech durations in the speech training data and a preset target margin for each speech duration; and training the speaker verification model using the margin parameters of the adaptively optimized loss function to determine the acceptable training difficulty of the speaker verification model. This invention utilizes training speech of varying lengths to better simulate real-life scenarios. Through adaptive optimization and fine-tuning of the margin, adjusting the margin according to the duration and similarity of each speech, this method achieves good speaker verification performance for speech of different durations in real-world scenarios.
Owner:AISPEECH CO LTD

Speech recognition model training method and device, equipment and readable storage medium

The invention discloses a speech recognition model training method and device, equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: firstly, acquiring voice training data and annotation data corresponding to the voice training data; the annotation data comprises text annotation data and intention annotation data; fuzzy processing is carried out on the text labeling data, and voice features of the voice training data are extracted; and training a speech recognition model based on the speech features, the text annotation data and the intention annotation data until the speech recognition model converges. According to the method, the text annotation data is fuzzified, some characters which are not concerned about intention classification are ignored, the speech recognition model is more focused on keywords, intention classification information is introduced in the training process, and the accuracy of the recognition result generated by the speech recognition model for intention classification is improved.
Owner:AISPEECH CO LTD

Vocalise bottle

1. Name of the designed product: voice training bottle (all-voice). 2. Use of the designed product: the designed product is all-voice voice training bottle. 3. Design points of the designed product: combination of shape and pattern. 4. Picture or photo best indicating the design points: perspective view.
Owner:张悦

Natural language processing systems and methods for intent classification of speech transcription

Aspects of the subject disclosure may include, for example, generating a natural language processing model by training an automatic speech recognition (ASR) encoder with manual transcription. The training is performed by correcting and adjusting relevant factors of the ASR encoder based on determined triplet loss, classification loss and Kullback-Leibler divergence loss. In response to an ASR utterance, the trained natural language processing model generates a predicted intent associated with the ASR utterance with improved accuracy. Other embodiments are disclosed.
Owner:JPMORGAN CHASE BANK NA

Method for training a speech recognition model and method for speech recognition

This application relates to a method for training a speech recognition model comprising: providing a speech training data set comprising a plurality of speech data items and corresponding speech tags; providing a speech recognition model to be trained comprising a convolution neural network, a first fully connected network, a recurrent neural network and a second fully connected network which are cascade coupled together, wherein each of the networks comprises one or more network layers each having a parameter matrix; and the speech recognition model processing speech data items to generate corresponding speech recognition results; and using the speech training data set to train the speech recognition model such that the parameter matrices of at least two adjacent network layers satisfies a predetermined constraint condition; and the speech recognition model trained using at least one loss function can generate speech recognition results at an accuracy satisfying a predetermined recognition target.
Owner:MONTAGE TECH CHENGDU CO LTD

Graphical user interface for voice training of electronic devices

1. The name of this design product: Graphical user interface for voice training of electronic devices. 2. Purpose of this design product: for use in an electronic device. 3. The key design feature of this design product lies in the graphical user interface displayed on the electronic device. 4. The picture or photo that best illustrates the key points of the design: Main view of Design 1. 5. Designate Design 1 as the base design. 6. Purpose of the Graphical User Interface: This graphical user interface is used for voice training, and it can provide visual feedback of voice tone and emotions through virtual avatars. The main view of Design 1 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change the corresponding expression according to the tone and emotion of the uploaded voice, as shown in the interface change status diagram of Design 1; The main view of Design 2 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change the corresponding expression according to the tone and emotion of the uploaded voice, as shown in the interface change status diagram of Design 2; The main view of Design 3 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The interface The virtual avatar in the lower left corner will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change state diagram of Design 3; the main view of Design 4 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change state diagram of Design 4; the main view of Design 5 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change state diagram of Design 5 ; The main view of Design 6 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change the corresponding expression according to the tone and emotion of the uploaded voice, as shown in the design 6 interface change state diagram; The main view of Design 7 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change the corresponding expression according to the tone and emotion of the voice, as shown in the design 7 interface change state diagram; The main view of Design 8 is the initial interface of voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change the corresponding expression according to the tone and emotion of the voice, as shown in the design 7 interface change state diagram; The virtual avatar will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change status diagram of Design 8; the main view of Design 9 is the initial interface for voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change status diagram of Design 9; the main view of Design 10 is the initial interface for voice training. Users can upload voice by clicking the record button at the bottom of the interface. The virtual avatar in the lower left corner of the interface will change its expression according to the tone and emotion of the uploaded voice, as shown in the interface change status diagram of Design 10.
Owner:PICC INFORMATION TECH CO LTD

Custom voice instruction recognition method and device based on twin network, and electronic equipment

The invention discloses a custom voice instruction recognition method and device based on a twin network and electronic equipment. Preprocessing the collected voice instruction data to obtain a voice training data set; constructing a twin network architecture; performing end-to-end joint training on the identification model and the discrimination model; deploying the trained identification model and discrimination model to the electronic equipment; performing feature extraction on the voice data, inputting the extracted features into the recognition model, and calculating to obtain a reference space vector; collecting a to-be-recognized voice signal in real time, and inputting the extracted voice features into the recognition model to calculate and obtain a detection space vector; and jointly inputting the detection space vector and the reference space vector into a discrimination model, calculating to obtain a discrimination value, comparing the discrimination value with a preset threshold value, and judging whether a target voice instruction is recognized or not according to a comparison result. The user-defined voice recognition instruction can be newly added on the basis of the preset voice instruction, the voice recognition effect is optimized, and the user interaction experience is improved.
Owner:ZHUHAI SPACETOUCH LTD

Voice training noise adding system and method based on mixed noise generation model

The invention aims to provide a voice training noise adding system and method based on a mixed noise generation model. The system comprises an input module, a noise environment enhancement module, a voice noise enhancement module and an output module. Wherein the input module is used for acquiring noise environment simple description information and clean voice data to be enhanced; the noise environment enhancement module converts the noise environment simple description information into a structured noise event sequence with time sequence characteristics; a voice noise enhancement module generates multi-source mixed noise according to the noise event sequence, and adds the multi-source mixed noise to the clean voice data according to a preset rule to obtain noisy voice data; and the output module is used for outputting the noisy voice data for voice model training. According to the invention, interaction characteristics of superposition, offset, interference and the like of different noise sources in the time dimension are fully joined, and the technical bottleneck that only linear superposition can be realized in a traditional mixing mode is solved.
Owner:GUANGDONG UNIV OF TECH

Speech synthesis method and device, electronic equipment and storage medium

The present application relates to the technical field of audio processing, and provides a speech synthesis method and device, electronic equipment and storage medium. The target acoustic feature of the to-be-processed text is obtained by extracting the acoustic feature from the to-be-processed text using a preset acoustic model; the target fundamental frequency feature is obtained by extracting the fundamental frequency feature from the target acoustic feature using a pre-trained fundamental frequency predictor, and the target energy feature is obtained by extracting the energy feature from the target acoustic feature using a pre-trained energy predictor; the target acoustic feature, the target fundamental frequency feature and the target energy feature are input into a pre-trained general vocoder to generate the speech audio of the to-be-processed text; the general vocoder is trained based on the speech audio of multiple speakers. By taking the acoustic feature, the fundamental frequency feature and the energy feature as the input of the vocoder for speech synthesis, and by training the vocoder using the speech of multiple speakers, the vocoder has general applicability, the training time of the vocoder is reduced, and the effect of speech synthesis is ensured.
Owner:SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD

Sound information protection method and device, storage medium, and electronic device

Embodiments of the present application provide a sound information protection system and method, a storage medium and an electronic device. The method comprises: after determining that a current voice incoming call belongs to the target call and obtaining an instruction to perform sound conversion, converting the sound of the current user into the sound of the target speaker through a lightweight voice conversion model, and then calling the current incoming call party. The lightweight voice conversion model is a model trained using sample voice containing the voice of the target speaker, and comprises a posterior encoder, a prior encoder and a decoder. The problem of how to avoid the leakage of user voice features in the related art is solved.
Owner:NANJING SILICON INTELLIGENCE TECH CO LTD

A voice activity detection method, system, terminal and storage medium

The application provides a voice activity detection method, system, terminal and storage medium, relates to a voice data processing technical field, and in particular to a voice activity detection method, which comprises the following steps: obtaining a voice training sample carrying a voice label; performing feature extraction according to the voice training sample to obtain a root amplitude spectrum feature sample and an Fbank feature sample; and constructing a voice activity detection model, wherein the voice activity detection model is used for performing feature fusion on the root amplitude spectrum feature and the Fbank feature obtained by performing feature extraction on voice data to obtain fusion features, and outputting a probability value of each frame of the voice data existing a voice signal based on the fusion features; training the voice activity detection model by using the root amplitude spectrum feature sample and the Fbank feature sample to obtain a trained voice activity detection model; and performing voice activity detection by using the trained voice activity detection model. The application can improve the accuracy of voice activity detection.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Speaker recognition method and related apparatus, device and storage medium

The application discloses a speaker recognition method and related device, equipment and storage medium, wherein the speaker recognition method comprises: obtaining a first voice of a first speaker and a second voice of a second speaker; determining whether the first speaker and the second speaker are the same speaker based on a feature distance between a first voice feature of the first voice and a second voice feature of the second voice; wherein the voice feature is extracted by a feature extraction model, the feature extraction model is trained based on a sample voice set to minimize the training loss, the sample voice set contains sample voices labeled with the sample speaker, the training loss is positively correlated with a first feature distance between the sample voice feature of the reference voice and the positive example voice feature, and negatively correlated with a second feature distance between the sample voice feature of the reference voice and the negative example voice feature. The above scheme can improve the accuracy of the feature extraction model in extracting voice features, thereby improving the speaker recognition accuracy.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Zero-shot task expansion of ASR models using task vectors

A method includes training, using an un-supervised learning technique, an auxiliary ASR model based on a first set of un-transcribed source task speech utterances to determine a first task vector, training, using the un-supervised learning technique, the auxiliary ASR model based on a second set of un-transcribed speech utterances to determine a second task vector, and training, using the un-supervised learning technique, the auxiliary ASR model based on un-transcribed target task speech utterances to determine a target task vector. The method also includes determining a first correlation between the first and target task vectors, determining a second correlation between the second and target task vectors, and adapting parameters of a trained primary ASR model based on the first and second source task vectors and the first and second correlations to teach the primary ASR model to learn how to recognize speech associated with the target task.
Owner:GOOGLE LLC

Voice activity detection model training method, voice activity detection method and device

The present disclosure provides a training method for a voice activity detection model, a voice activity detection method, and a device. The training method includes obtaining a training set, wherein the training set includes multiple voice training samples; performing conversion processing on the voice training samples to obtain target logarithmic Mel-spectrogram features; processing the target logarithmic Mel-spectrogram features using a gated convolution layer and a maximum pooling layer to obtain an encoding result, wherein the convolutional encoding module includes a gated convolution layer and a maximum pooling layer; processing the encoding result using a first fully connected layer to obtain a predicted label, wherein the predicted label represents whether a voice signal exists in the voice training sample; processing the encoding result using a residual decoding module to obtain a predicted result; inputting the predicted label and the predicted result into a loss function to output a loss result; iteratively adjusting the network parameters of an initial voice detection model according to the loss result to obtain a trained voice activity detection model, wherein the initial voice detection model includes a convolutional encoding module, a first fully connected layer, and a residual decoding module.
Owner:UNIV OF SCI & TECH OF CHINA

Voice large model reasoning method and device for long voice

The invention provides a large model reasoning method and device for long voices, and the method comprises the steps: obtaining a voice training signal marked with a training label, carrying out the coding of the voice training signal through an information extraction module, obtaining the original voice representation of the voice training signal, and carrying out the coding of the voice training signal according to the text content and inter-frame similarity of the original voice representation. Performing compression combination on the original voice representation to obtain compressed voice representation; inputting the compressed voice representation into a large language model, executing a reasoning task to obtain a reasoning result corresponding to the voice training signal, and constructing a loss function training information extraction module according to the reasoning result and a training label; and inputting the long voice signal into the trained information extraction module to obtain a compressed voice representation of the long voice signal, and inputting the compressed voice representation into the large language model to obtain a reasoning result corresponding to the long voice signal. According to the method, the long speech understanding capability is enhanced, and the reasoning cost and the reasoning time are greatly reduced while the high generation quality is ensured.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Voice training device supporting human-computer interaction

The utility model relates to the technical field of voice training devices, in particular to a voice training device supporting human-computer interaction, which comprises a shell, a storage bin and a protective cover, the storage bin is fixedly connected with one side surface of the shell, the other side surface of the shell is provided with a mounting groove, the inner bottom surface of the mounting groove is slidably connected with a plurality of clamping jaws, and the protective cover is mounted on the outer side of the mounting groove. The inner side face of the containing bin is fixedly connected with an inflatable doll, the inflatable doll is folded and filled in the containing bin, the outer side face of the containing bin is provided with an air nozzle, the air nozzle is connected with the inflatable doll, the bottom of the clamping jaw is connected with sliding pieces, the sliding pieces are inserted into the bottom of the installation groove, and an elastic rope is arranged between every two opposite sliding pieces. The electronic device is used for voice interaction of a user on the back face of the device and is deflated and stored when carried, the doll is conveniently stored and taken through the storage bin, the structure of the device is simplified, meanwhile, the overall cost of the device is greatly reduced, and meanwhile the requirements of special crowds are met through the inflatable doll.
Owner:HUNAN ART VOCATIONAL COLLEGE

Multi-dimensional AI platform intelligent voice response system using voice synthesis technology

The invention relates to the technical field of voice synthesis, in particular to a multi-dimensional AI platform intelligent voice response system using the voice synthesis technology, which screens out matched historical user question voices according to text semantic similarity corresponding to current user question voices and historical user question voices, and sends the matched historical user question voices to a user terminal. Obtaining a voice training set and the weight of each element in the voice training set according to the voice feature similarity between the question voice of the current user and the question voice of the matched historical user in combination with the user score value of the manual reply voice, training an acoustic model, obtaining a reply text corresponding to the question voice of the current user, and obtaining the question voice of the current user; and inputting into the trained acoustic model, generating a reply voice signal, and then outputting to the current user. According to the method, the voice training set is screened and constructed from historical manual reply voices, so that an acoustic model can learn a more natural and smooth voice synthesis mode, and the reply voice signal contains rich voice features and expression modes.
Owner:ROPEOK TECHNOLOGY GROUP CO LTD

Multifunctional swallowing training and sound production auxiliary exerciser

The invention discloses a multifunctional swallowing training and sound production auxiliary exercising device, and belongs to the technical field of sound production training devices. Comprising an arc-shaped plate fixed to the lower jaw of a user, a supporting plate fixed to the chest of the user, a fixing column connecting the arc-shaped plate and the supporting plate and a rear neck adjusting belt, and a pickup, a displayer, a display lamp and a control terminal are integrated. The sound pick-up collects voice signals, the displayer displays volume, and the display lamp feeds back sound production quality through the control terminal. A piezoelectric sensor is embedded in the supporting plate to monitor the thoracic expansion amplitude; when chest breathing is detected, the clavicle heating module is linked to heat up, phrenic nerve reflex is activated through thermal stimulation to guide abdominal breathing, vocal cord closure is improved, and the risk of aspiration by mistake is reduced. The equipment is integrated with a rotating assembly to realize multi-angle adjustment and automatic switching of the display; a multi-frequency vibration module and a temperature control heating module are arranged in the arc-shaped plate, and targeted muscle stimulation is provided. Through mechanical linkage and multi-mode feedback, the breathing-sounding-swallowing cooperative training efficiency is optimized, and the user suitability and the rehabilitation effect are improved.
Owner:SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)

Synthetically generating inner speech training data

Methods and systems are disclosed for synthetically generating inner speech training data. The methods and systems access a collection of overt speech signals representing phonemes, phoneme sounds, words or phrases spoken at least partially using overt speech. The methods and systems transform the collection of overt speech signals into inner speech training data comprising electromyograph (EMG) data representing inner speech corresponding to the phonemes, phoneme sounds, words or phrases spoken at least partially using the overt speech. The methods and systems train a machine learning model to decode inner speech signals into a set of corresponding phonemes, phoneme sounds, words or phrases based on the inner speech training data.
Owner:SNAP INC

A data processing method, apparatus, device, storage medium, and program product

Embodiments of the present application disclose a data processing method, apparatus, device, storage medium, and program product. Since multiple sub-audio information segments in the first voice training sample correspond to users of the same accent type, the multiple sub-audio information segments have similar accent features. Based on this, based on artificial intelligence technology, a first loss function is determined according to the differences between the accent features of the multiple sub-audio information segments, and a second loss function is determined according to the differences between the to-be-determined accent type and the sample accent type. By adjusting the parameters of the initial accent classification model based on the first loss function and the second loss function, on the one hand, the accent type determined by the initial accent classification model can be made more accurate, and on the other hand, during the training of the model, the differences between the accent features of the determined sub-audio information segments can be controlled within a reasonable range, making the way of determining the accent features more conform to the real accent situation, and improving the accuracy and rationality of the training of the accent classification model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Intelligent English teaching system for English teaching

The invention relates to the technical field of English teaching, in particular to an intelligent English teaching system for English teaching, which comprises the following modules: an intelligent lesson preparation module, an interactive teaching module, an intelligent evaluation module, a voice training module, a learning resource library module, a learning condition analysis module, a personalized learning module and a home-school communication module. The intelligent lesson preparation module is used for helping teachers to quickly integrate teaching resources and generate personalized teaching plans according to teaching themes and learning conditions; the intelligent English teaching system enriches teaching resources, motivates learning interest, integrates massive teaching resources and covers various forms and themes, teachers can easily obtain and apply to teaching, through the interactive teaching module, various interactive tools such as preemptive answering and group discussion areas are provided, the teachers can conveniently organize classroom interaction and know learning conditions of students in real time, and the teaching efficiency is improved. The classroom atmosphere is activated, the participation degree and enthusiasm of students are improved, and learning condition analysis and personalized learning modules are used.
Owner:周春霞

Voice noise reduction method and device and storage medium

The invention relates to a voice noise reduction method and device and a storage medium. The voice noise reduction method comprises the following steps: acquiring a voice signal to be subjected to noise reduction, which is acquired by the electronic equipment in an actual use environment; based on the to-be-denoised voice signal and a denoising model, denoising the to-be-denoised voice signal to obtain a target voice; wherein the noise reduction model is obtained by training based on voice training data acquired by the electronic equipment in a plurality of different use environments, and the plurality of different use environments at least comprise the first use environment. The electronic equipment collects data in an actual use environment, so that the consistency of the model training data and actual application data is ensured, and the user experience is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Speech enhancement method and device based on time-frequency domain feature fusion, and electronic equipment

The invention discloses a speech enhancement method and device based on time-frequency domain feature fusion and electronic equipment, and the method comprises the steps: carrying out the preprocessing of an initial speech training data set, and obtaining a target speech training data set; performing feature extraction on the target voice training data set to obtain candidate voice features; normalizing the candidate voice features to obtain target voice features; performing feature fusion on the target time domain feature and the target frequency domain feature to generate candidate fusion features; inputting the candidate fusion features into a multi-scale convolution feature enhancement module for feature enhancement, and generating to-be-trained fusion features; inputting the to-be-trained fusion features into the initial speech enhancement model for training to obtain a target speech enhancement model; and inputting the target fusion feature of the to-be-enhanced noisy speech into the target speech enhancement model to generate a target enhanced speech. The method can suppress noise interference, improves the speech enhancement effect in a complex noise scene, and can be widely applied to the technical field of speech enhancement.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV