Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

376 results about "Acoustic model" patented technology

An acoustic model is used in automatic speech recognition to represent the relationship between an audio signal and the phonemes or other linguistic units that make up speech. The model is learned from a set of audio recordings and their corresponding transcripts. It is created by taking audio recordings of speech, and their text transcriptions, and using software to create statistical representations of the sounds that make up each word.

Conference summary processing method and system using AI

The invention relates to the technical field of intelligent conference processing, and relates to a conference summary processing method and system using AI, and the method comprises the steps: carrying out the real-time noise suppression of a collected conference audio stream and associated text data through a noise suppression algorithm, and carrying out the cross-modal alignment of the denoised data through a cross-modal alignment algorithm; a domain-specific attention head is inserted into an attention layer of the pre-trained Transform model, a domain-enhanced speech recognition model is constructed, and audio is converted into a text sequence with a speaker tag; adopting a heterogeneous graph neural network to construct a structured topic evolution graph; key decision nodes in the structured topic evolution graph are extracted based on a reinforcement learning strategy, and a final conference summary document is generated. In the decoding stage, the fusion proportion of the acoustic model and the language model is dynamically adjusted based on the real-time acoustic confidence coefficient, the recognition rate of the vocabularies in the professional field is increased, and the problems of frequent term transcription errors and poor semantic coherence in the professional conference are effectively solved.
Owner:GUANGZHOU DAZZLE VIEW INTELLIGENT TECH CO LTD

Detecting and using non-textual information in human speech

PCT designated stage expiredWO2025141559A1Speech recognitionMachine learningEncoder decoderAcoustic model
Automatic recognition of non-verbal messages in speech, and in particular to detection or analysis of prosodic multilayered analysis of intonation units such as prosodic unit prototypes and their multi-labeled variations, may form a hierarchical classification for the analysis of non¬ verbal information or cues in speech. A speech captured by a microphone is fed to a weakly- supervised deep learning acoustic model for speech recognition and transcription, that may be based on encoder-decoder Transformer architecture, such as Whisper by OpenAI. The model is trained to output multiple words form the text in the captured speech, to identify Intonation Units (IUs) that include one or more words, and associate non-verbal labels to each of the IUs. The labels may indicate a prototype, a discourse function (such as a conversation action), an emotion, an emphasis, or an attitude, as well as a genre of a part of, or whole of, the entire captured speech.
Owner:YEDA RES & DEV CO LTD

Voice command word recognition post-processing method, system and equipment and storage medium

The invention relates to the technical field of voice decoding, in particular to a voice command word recognition post-processing method, system and device and a storage medium, and the method comprises the steps: obtaining a phoneme probability matrix and a command word path score outputted by an acoustic model; preliminarily judging whether misrecognition, mixed recognition or out-of-set word recognition exists or not based on the phoneme probability matrix and the command word path score; if the preliminary judgment result shows that misrecognition, mixed recognition or out-of-set word recognition exists, secondary confirmation is carried out based on misrecognition, mixed recognition or out-of-set word recognition, and a final recognition result is output based on a confirmation result. According to the method, the model does not need to be retrained, three types of recognition problems can be efficiently processed by utilizing a post-processing mechanism, and the recognition sensitivity and the error recognition rate can be balanced in a scene that end-side offline resources are limited.
Owner:深圳市友杰智新科技有限公司

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Cross-language voice migration synthesis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a cross-language voice migration synthesis method, device, equipment and medium. And training an acoustic model by adopting a hierarchical self-adaptive fine tuning strategy, fusing the phoneme sequence for reasoning and the tone mark to generate a representation sequence, and finally synthesizing a target language speech signal. According to the method, a sharing and separation parallel cross-language modeling structure is constructed, the accuracy of phoneme and tone modeling in a low-resource language is effectively improved, the target language acoustic model has higher generalization ability and migration efficiency in combination with a multi-stage self-adaptive fine tuning and training data enhancement strategy, and the method is suitable for being applied to the field of multi-language modeling. And finally, synchronous improvement of the voice naturalness and the tone fidelity is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio and video recording-based ASR identification enhancement method

The invention discloses an ASR identification enhancement method based on audio and video recording. According to the method, the accuracy and compliance of voice recognition in the financial service interaction process are improved by fusing the audio and environment feature information in the banking business double-recording scene. The method comprises the following steps: firstly, constructing an acoustic model for a bank outlet environment, and extracting audio features and interaction scene information of conversation between a client and a worker; and then, designing a vocabulary recognition module special for the financial field, dynamically adjusting language model parameters according to professional term libraries and utterance modes of different business types, and effectively coping with key links such as financial product introduction, risk prompt and customer confirmation. Compared with a traditional ASR system, the voice recognition accuracy in the banking business handling process is remarkably improved, particularly, key term recognition and important information extraction are prominent, and more reliable technical support is provided for financial service standardized management and double-recording quality inspection.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

Self-adaptive collaborative acoustic environment active treatment method and system

The invention provides a self-adaptive collaborative acoustic environment active treatment system and method. The system comprises a multi-dimensional environment sensing network module, an intelligent voiceprint recognition and sound field prediction module, an active noise control module, a central intelligent collaborative regulation, diagnosis and self-learning module and a passive acoustic intervention module. The multi-dimensional environment sensing network module collects multi-dimensional sensing data in real time; the intelligent voiceprint recognition and sound field prediction module outputs a voiceprint recognition result and a noise source space coordinate and predicts a sound field evolution trend; the active noise control module outputs a residual noise signal; the central intelligent cooperative regulation, diagnosis and self-learning module outputs an active and passive cooperative control instruction set; the passive acoustic intervention module controls broadband noise to block or change its propagation path. According to the invention, through a dual-channel self-learning mechanism of the deviation diagnosis capability, continuous evolution of system performance can be realized according to a deviation autonomous optimization acoustic model and a control strategy, and accurate, efficient and prospective active treatment is carried out on a complex noise environment.
Owner:CHINA FIRST METALLURGICAL GROUP

Real-time speech recognition method based on Bluetooth audio stream

The invention relates to the technical field of speech recognition, and discloses a real-time speech recognition method based on a Bluetooth audio stream, which comprises the following steps: analyzing bit allocation parameters to calculate quantized bit distribution and generate a frequency domain confidence mask, monitoring a packet loss concealment state flag bit of a decoder, forcibly setting the mask as a blocking threshold when an algorithm is activated, and generating a real-time speech recognition result. According to the method, a cross-level feature purification mechanism based on protocol priori and link states is constructed, acoustic model illusion caused by forged waveforms is blocked through targeted arbitration while deterministic quantization noise is eliminated, and the acoustic model recognition accuracy is improved. And the identification accuracy under a severe channel is ensured.
Owner:SHENZHEN HUIJIEXIN TECH CO LTD

Bluetooth earphone AI voice control method and system

The invention relates to the technical field of voice recognition, in particular to a Bluetooth headset AI voice control method and system, and the method comprises the following steps: collecting a voice sample, carrying out the feature extraction of a voice signal through employing a Mel-frequency cepstrum coefficient, and generating voice feature data; and inputting the voice feature data into an acoustic model, and improving the recognition rate of the key instruction words by the acoustic model through learning features to obtain an optimized acoustic model. According to the method, acoustic feature capture is realized through voice signal feature extraction, the recognition accuracy of the key instruction is improved in cooperation with deep learning of acoustic features, the influence of different user speech speeds on the recognition effect is overcome by applying the time alignment technology of voice input, and the stability of the control instruction in different use scenes and speech speed changes is ensured. In addition, the recognition capability of a specific command is further mined and optimized by means of statistical characteristic analysis of voice, and efficient and automatic recognition and extraction of key control commands in continuous speech streams are completed.
Owner:JIANGXI CHANGRONG TECHNOLOGY CO LTD

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Voice generation method and device, equipment and storage medium thereof

The invention belongs to the technical field of voice generation, and relates to a voice generation method and device, equipment and a storage medium thereof. Obtaining a visual feature vector sequence of the target video frame sequence; inputting the text feature vector sequence and the visual feature vector sequence into a cross-modal fusion layer to obtain a feature fusion representation output by the cross-modal fusion layer; obtaining acoustic features output by the variant acoustic model according to the feature fusion representation; and inputting the acoustic features into a preset vocoder, and generating a target voice waveform through multi-scale convolution and up-sampling processing. According to the method, not only text features but also video features are introduced during voice generation, and voice generation is realized through cross-modal fusion features. The method is applied to financial or medical service intelligent customer service answering or service product marketing introduction scenes, and more natural and real voice is generated in combination with visual context information.
Owner:SHENZHEN PINGAN COMM TECH CO LTD

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice control ultrasonic wave adjustment setting method based on CTC-Attention mixed architecture

The invention relates to a voice control ultrasonic wave adjustment setting method based on a CTC-Attention mixed architecture, and the method comprises the following steps: S1, receiving a voice instruction of a doctor through a voice receiving device, and carrying out the preprocessing of an original voice signal; s2, constructing an acoustic model of a CTC-Attention hybrid architecture, and training the acoustic model to obtain a trained acoustic model; s3, inputting the preprocessed voice signal into the trained acoustic model, and performing voice recognition through CTC path probability alignment and Attention context dependence joint optimization; s4, outputting an identification result and decoding the identification result into a text instruction; and S5, analyzing the operation intention according to the text instruction, and controlling the ultrasonic instrument to execute a corresponding state adjustment operation. According to the method, the voice signal of the operator is received, the voice signal is recognized and analyzed, the obtained result is optimized and then converted into the character content with the meaning, then the ultrasonic energy with the intensity corresponding to the character content is generated by the ultrasonic instrument based on the character content, and the ultrasonic instrument is effectively used for medical operation.
Owner:SHUGUANG HOSPITAL AFFILIATED WITH SHANGHAI UNIV OF T C M

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Configuration method, system and equipment of multi-room audio system and storage medium

The invention relates to the technical field of audio processing, and discloses a configuration method, system and device of a multi-room audio system and a storage medium, which are used for improving the accuracy of parameter adjustment of each audio device in the audio system. The configuration method of the multi-room audio system comprises the following steps: acquiring spatial layout information of each room and physical position information and performance parameters of each audio device; based on the spatial layout information of each room and the physical position information and performance parameters of each audio device, constructing an acoustic model comprising the plurality of rooms and the audio devices; analyzing an acoustic coupling relationship between different rooms and signal transmission characteristics of each audio device based on an acoustic model to obtain an acoustic characteristic parameter set; according to the acoustic characteristic parameter set, sound channel distribution and volume adjustment of each audio device are optimized, and an optimized audio configuration parameter set is obtained; and according to the optimized audio configuration parameter set, automatically configuring sound channel distribution, volume adjustment and audio signal transmission paths of the multi-room audio system.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Suzhou dialect medical voice electronic medical record conversion system and method

The invention provides a Suzhou dialect medical voice electronic medical record conversion system and method, and the system comprises an acoustic feature extraction module which is used for extracting the acoustic features of an input voice signal; the dialect tone recognition module is used for recognizing a multi-tone system of Suzhou dialects; the voiced sound processing module is used for detecting voiced sound initial consonants in Suzhou dialects and performing acoustic feature mapping; the medical term mapping module comprises a corresponding relation library of dialect medical vocabularies and standard medical terms and a context-based ambiguity resolution unit; the speech recognition engine comprises an acoustic model, a pronunciation dictionary and a language model; and the medical record generation module is used for converting the identification result into a structured electronic medical record. According to the method, a sliding window processing strategy is adopted, continuous voice input of a doctor can be effectively processed, a long-time voice input scene is supported, and various requirements in actual clinical application are met.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Self-adaptive acoustic analysis method and system

The invention relates to the technical field of acoustic analysis, in particular to a self-adaptive acoustic analysis method and system, and the method comprises the following steps: obtaining a sound pressure amplitude screening channel, extracting a main sound point time index, judging a fluctuation direction to generate a path identifier, and completing channel alignment and path binding to generate an analysis result. According to the invention, by acquiring the maximum sound pressure amplitude of each channel and dynamically screening the effective channel, the capability of capturing significant sound in a complex sound source environment is improved, and by combining directional extraction and fluctuation trend judgment of main sound point time index differences in a continuous period, dynamic perception and track identification of sound source feature changes are realized. The stability of path identification is enhanced by using direction consistency and time sequence alignment operation, accurate binding and real-time adaptive sound source tracking of multiple channels are completed, the identification accuracy and path continuity in a sound source dynamic change scene are effectively improved, meanwhile, redundant channel interference is avoided, and the accuracy of path identification is improved. And the processing efficiency and the adaptive capacity of the acoustic model are enhanced.
Owner:XIAN INT UNIV

Method, device, storage medium and electronic device for generating virtual voice

The present invention discloses a method, device, storage medium, and electronic device for generating virtual speech. The method comprises: obtaining multiple different speech text samples and speech attribute information, wherein each speech text sample in multiple different language speech text samples corresponds to a language and an object; inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; and training a preset speech acoustic model based on generative adversarial network modeling using the text features and speech features to obtain a target acoustic model for generating virtual speech. The present invention can support cross-language data training and the generation of cross-language speakers. The multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

Speech recognition system based on improved Transform architecture

The invention belongs to the field of artificial intelligence and voice recognition, and particularly relates to a voice recognition system based on an improved Transform architecture, which comprises a self-positioning module used for receiving an original audio signal, outputting a self-supervised voice feature vector and a traditional audio feature vector in parallel, and sending the self-supervised voice feature vector and the traditional audio feature vector to a feature normalization conversion module; the feature normalization conversion module is used for mapping the self-supervised voice feature vector and the traditional audio feature vector to a standard speaker feature space and outputting a normalized feature; the perception modeling module performs multi-scale time sequence coding through an improved Transform structure, and outputs a voice semantic probability distribution sequence; the CTC loss module is used for optimizing the acoustic model according to the voice semantic probability distribution sequence; the collaboration unit is used for receiving multiple paths of original audio features, screening credible channels from the obtained synchronization features, and outputting corrected features; and the fusion filtering module is used for receiving the local features and the corrected features, generating global probability distribution through attention weight fusion, and decoding the global probability distribution into a final text sequence.
Owner:SHENYANG LIGONG UNIV

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice recognition method based on acoustic model, computer equipment and storage medium

The invention belongs to the field of voice recognition, and discloses a voice recognition method based on an acoustic model, computer equipment and a storage medium. The method comprises the following steps: acquiring voice features of to-be-recognized voice; inputting the voice features into an acoustic model, and outputting a recognition result by the model; wherein the time sequence processing network layer firstly determines the ratio of current input future frames needing to be pre-watched to context information through a pre-trained gating fusion unit, then calculates the number of the future frames needing to be pre-watched based on the ratio, obtains the corresponding future frames, calculates long-time context representation in combination with the future frames, processes the long-time context representation and outputs the long-time context representation to the next layer of network. According to the method and the device, the problem of static binding of delay and accuracy in the prior art is solved by dynamically adjusting the number of the future frames to be pre-watched, low-delay response to simple command words is realized, the recognition accuracy is improved through multiple future frames to be pre-watched for easily-confused instructions, the balance of the delay and the accuracy is realized, and the performance of a voice recognition system and the user experience are improved.
Owner:深圳市友杰智新科技有限公司

Mediator timbre cloning method and system, electronic equipment and storage medium

The invention provides a mediator timbre cloning method and system, electronic equipment and a storage medium. The method comprises the following steps: selecting a tone of a mediator in response to a selection instruction input by a user; obtaining an input text; predicting a target audio feature vector of the input text by using an autoregression model; a target clustering center matched with the target audio feature vector is searched in an audio dictionary, the audio dictionary comprises a plurality of clustering clusters, and each clustering cluster comprises a plurality of audio feature vectors; and inputting the input text phonemes of the input text, the frequency characteristics of a reference audio and the target clustering center into the trained acoustic model to obtain a target output audio, the reference audio being a mediator audio corresponding to the tone of the mediator. According to the method, the mediation efficiency can be improved, the service consistency is ensured, the user experience is improved, and the method has better practicability and market competitiveness.
Owner:SHANGHAI JINQIAO YIFA INFORMATION TECH CO LTD

Voice recognition method and system based on AI large model

The invention provides a voice recognition method and system based on an AI large model, and belongs to the technical field of voice recognizing.The method comprises the steps that firstly, noise reduction and frequency band enhancement in a noise environment are achieved through a pre-trained anti-noise suppression network, and a basis is provided for follow-up processing in combination with multi-dimensional spectrum quality scores; based on similarity matching of element feature vectors and a pre-constructed dialect thermodynamic diagram library and matching weight adjustment based on frequency spectrum quality scores, accurate modeling of specific dialect pronunciation deviation is achieved, moreover, through an acoustic adaptation matrix and a language model adaptation matrix generated through a super network, the adaptability of the model to different dialects is improved, and in addition, the accuracy of the model is improved. According to the method, a fusion thermodynamic diagram and an acoustic adaptation matrix are jointly injected into a pre-trained acoustic model, the recognition accuracy of dialect phonemes is improved through multi-level attention correction, and finally, accurate speech recognition is achieved by adopting a thermodynamic diagram guided cluster search algorithm and combining verification of an adversarial discrimination network.
Owner:GUANGZHOU WEIJIE INTELLIGENT TECHNOLOGY CO LTD

Voice interaction method and system based on intelligent doll

The invention discloses a voice interaction method and system based on an intelligent doll, and relates to the technical field of intelligent acoustic interaction, and the method comprises the steps: carrying out the time-frequency analysis processing of a vibration feature matrix, and generating a modal parameter set of vocal cord vibration through biomechanical modeling in combination with the human neck tissue density features; inputting the modal parameter set of vocal cord vibration into the vibration-acoustic model to generate a sound source excitation field, performing environmental noise compensation on the sound source excitation field by using the acoustic characteristic matrix, and generating a time domain pure voice signal through an acoustic wave equation; and collecting real-time position coordinates of the user, calculating and driving a piezoelectric loudspeaker array of the intelligent doll to adjust the phase according to the frequency spectrum characteristics of the time domain pure voice signal, forming a directional focusing sound field, recording user physiological feedback, and generating a user physiological feedback matrix. According to the method, the vocal cord displacement modal function vector is solved through the nonlinear integral and augmented Lagrange algorithm, so that the acquisition precision and the anti-noise capability of the voice signal are improved from the source.
Owner:XINGFUQUAN (BEIJING) INTERNATIONAL ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Corpus expansion method, system and equipment based on speech synthesis and medium

The invention relates to the technical field of speech synthesis, in particular to a speech synthesis-based corpus expansion method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing including data annotation based on the collected audio and corresponding text of a target speaker; extracting acoustic features from the preprocessed audio; on the basis of a pre-trained acoustic model, performing personalized fine tuning by using the annotation data and the acoustic features, and training personalized acoustic models of a plurality of speakers at the same time through multi-thread parallel computing; calling the trained personalized acoustic model, and synthesizing a voice corpus of the target text in combination with a vocoder; and based on the trained personalized acoustic model, continuously expanding the corpus by changing the text. The personalized voice corpus is quickly generated through a small number of voice samples, the data acquisition cost is remarkably reduced, and the corpus construction efficiency is improved.
Owner:深圳市友杰智新科技有限公司

Term speech recognition system and recognition method for coal mine

The invention discloses a coal mine term speech recognition system and method, and belongs to the technical field of coal mine management, and the system comprises a data collection and processing module which is used for simulating an underground operation environment to carry out multi-scene recording collection; performing synthetic data enhancement in a mode of mixing real noise and synthetic speech, and constructing field customized data; according to the invention, through customized data acquisition, acoustic model optimization and language model adaptation, the recognition accuracy of the coal mine terminology is effectively improved; the system integrates an anti-noise algorithm, supports dialect recognition, can stably operate in complex noise and different dialect environments, and is high in environmental adaptability; the voice recognition processing module and the edge calculation module work cooperatively, so that rapid information transmission and abnormity monitoring are realized, and the working efficiency is improved; meanwhile, accurate instruction recognition and real-time abnormity monitoring are achieved, equipment misoperation is effectively avoided, safety accident risks are reduced, and coal mine production safety is guaranteed.
Owner:XUZHOU MINING BUSINESS GROUP +3

Training system and method for acoustic model

An acoustic model training system includes a first device that is connectable to a network and that is used by a first user, and a server that is connectable to the network. The first device, under control by the first user, is configured to upload a plurality of sound waveforms to the server, select, as a first waveform set, one or more sound waveforms from the plurality of sound waveforms after or before updating the plurality of sound waveforms, and transmit to the server a first execution instruction for a first training job for an acoustic model configured to generate acoustic features. The server is configured to, based on the first execution instruction from the first device, start execution of the first training job using the first waveform set, and provide, to the first device, a trained acoustic model trained by the first training job.
Owner:YAMAHA CORP

Voice cloning method, training method, device and medium

An embodiment of the present invention provides a voice cloning method, a training method, a device and a medium. The voice cloning method specifically includes: receiving text and the original audio of a cloning object; determining the voiceprint feature corresponding to the original audio; inputting the text and the voiceprint feature into an acoustic model to obtain a corresponding acoustic feature, where the acoustic model is obtained according to the voiceprint features corresponding to training samples; and determining a corresponding target audio according to the acoustic feature. The embodiment of the present invention can reduce the audio data volume of the cloning object, and can improve the processing efficiency and application scope of voice cloning.
Owner:BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD

Method of recognizing speech, device, and medium

A method of recognizing a speech, a device, and a medium. The method includes: processing, by using an acoustic model, speech data to be recognized and a first text segment obtained by recognition to obtain respective acoustic probabilities of a plurality of candidate text segments; processing the first text segment by using a first language sub-model to obtain respective initial language probabilities of the plurality of candidate text segments; processing the first text segment by using a constraint sub-model to obtain extendibility relationships of the plurality of candidate text segments with respect to the first text segment; adjusting the initial language probabilities of the candidate text segments according to the extendibility relationships to obtain respective first language probabilities of the plurality of candidate text segments; and determining a target text segment from the plurality of candidate text segments according to the first language probabilities and the acoustic probabilities.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Vehicle-mounted voice interaction method and system and readable storage medium

The invention relates to the technical field of intelligent vehicle-mounted systems, and discloses a vehicle-mounted voice interaction method and system and a readable storage medium, and the method comprises the steps: synchronously collecting initial voice and video data in a vehicle-mounted environment; performing wake-up word detection through a local acoustic model, and based on the detection confidence, extracting a mouth shape visual feature sequence by using a mouth shape recognition model to perform mouth shape verification so as to obtain a wake-up state and sound source positioning information; activating an interaction module at a corresponding position, and performing semantic recognition on the collected interaction voice and video data through a local model and a cloud model respectively; and finally, carrying out fusion cross validation on the local semantic recognition result and the cloud semantic recognition result to generate a final semantic recognition instruction, and executing corresponding operation by the vehicle-mounted system. According to the method, the recognition accuracy, the response speed and the robustness of vehicle-mounted voice interaction in a complex environment are improved, the false wake-up rate is effectively reduced, and the user experience is optimized.
Owner:深圳海冰科技有限公司