Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

265 results about "Acoustic model" patented technology

An acoustic model is used in automatic speech recognition to represent the relationship between an audio signal and the phonemes or other linguistic units that make up speech. The model is learned from a set of audio recordings and their corresponding transcripts. It is created by taking audio recordings of speech, and their text transcriptions, and using software to create statistical representations of the sounds that make up each word.

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio and video recording-based ASR identification enhancement method

The invention discloses an ASR identification enhancement method based on audio and video recording. According to the method, the accuracy and compliance of voice recognition in the financial service interaction process are improved by fusing the audio and environment feature information in the banking business double-recording scene. The method comprises the following steps: firstly, constructing an acoustic model for a bank outlet environment, and extracting audio features and interaction scene information of conversation between a client and a worker; and then, designing a vocabulary recognition module special for the financial field, dynamically adjusting language model parameters according to professional term libraries and utterance modes of different business types, and effectively coping with key links such as financial product introduction, risk prompt and customer confirmation. Compared with a traditional ASR system, the voice recognition accuracy in the banking business handling process is remarkably improved, particularly, key term recognition and important information extraction are prominent, and more reliable technical support is provided for financial service standardized management and double-recording quality inspection.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

Real-time speech recognition method based on Bluetooth audio stream

The invention relates to the technical field of speech recognition, and discloses a real-time speech recognition method based on a Bluetooth audio stream, which comprises the following steps: analyzing bit allocation parameters to calculate quantized bit distribution and generate a frequency domain confidence mask, monitoring a packet loss concealment state flag bit of a decoder, forcibly setting the mask as a blocking threshold when an algorithm is activated, and generating a real-time speech recognition result. According to the method, a cross-level feature purification mechanism based on protocol priori and link states is constructed, acoustic model illusion caused by forged waveforms is blocked through targeted arbitration while deterministic quantization noise is eliminated, and the acoustic model recognition accuracy is improved. And the identification accuracy under a severe channel is ensured.
Owner:SHENZHEN HUIJIEXIN TECH CO LTD

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech enhancement and high-precision recognition method and system in complex environment

PendingCN121641016ASpeech recognitionSpectral density estimationNerve network
The invention provides a voice enhancement and high-precision recognition method and system in a complex environment, and relates to the technical field of voice processing, and the method comprises the steps: collecting a time domain signal in an off-road parking sentry box environment for preprocessing, detecting a mute segment signal in a standard time domain signal for noise power spectral density estimation, and obtaining a noise power spectral density value; a reverberation parameter is obtained by combining voice onset information and noise spatial correlation estimation, prediction is performed by using a deep neural network model, voice masking is applied to microphone array signals to perform enhancement processing, adaptive feature extraction is performed on time domain enhanced voice signals, and a voice signal is obtained. And performing high-precision recognition on the voice adaptive feature sequence based on an acoustic model and a language model, and outputting a target recognition text. The technical problems of poor voice signal quality and low recognition accuracy in a complex noise environment in the prior art are solved. The technical effects of improving the voice signal quality and the recognition accuracy and realizing clear, accurate and real-time voice interaction are achieved.
Owner:INTELLIGENT INTER CONNECTION TECH CO LTD

Suzhou dialect medical voice electronic medical record conversion system and method

The invention provides a Suzhou dialect medical voice electronic medical record conversion system and method, and the system comprises an acoustic feature extraction module which is used for extracting the acoustic features of an input voice signal; the dialect tone recognition module is used for recognizing a multi-tone system of Suzhou dialects; the voiced sound processing module is used for detecting voiced sound initial consonants in Suzhou dialects and performing acoustic feature mapping; the medical term mapping module comprises a corresponding relation library of dialect medical vocabularies and standard medical terms and a context-based ambiguity resolution unit; the speech recognition engine comprises an acoustic model, a pronunciation dictionary and a language model; and the medical record generation module is used for converting the identification result into a structured electronic medical record. According to the method, a sliding window processing strategy is adopted, continuous voice input of a doctor can be effectively processed, a long-time voice input scene is supported, and various requirements in actual clinical application are met.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Self-adaptive acoustic analysis method and system

The invention relates to the technical field of acoustic analysis, in particular to a self-adaptive acoustic analysis method and system, and the method comprises the following steps: obtaining a sound pressure amplitude screening channel, extracting a main sound point time index, judging a fluctuation direction to generate a path identifier, and completing channel alignment and path binding to generate an analysis result. According to the invention, by acquiring the maximum sound pressure amplitude of each channel and dynamically screening the effective channel, the capability of capturing significant sound in a complex sound source environment is improved, and by combining directional extraction and fluctuation trend judgment of main sound point time index differences in a continuous period, dynamic perception and track identification of sound source feature changes are realized. The stability of path identification is enhanced by using direction consistency and time sequence alignment operation, accurate binding and real-time adaptive sound source tracking of multiple channels are completed, the identification accuracy and path continuity in a sound source dynamic change scene are effectively improved, meanwhile, redundant channel interference is avoided, and the accuracy of path identification is improved. And the processing efficiency and the adaptive capacity of the acoustic model are enhanced.
Owner:XIAN INT UNIV

Speech recognition system based on improved Transform architecture

The invention belongs to the field of artificial intelligence and voice recognition, and particularly relates to a voice recognition system based on an improved Transform architecture, which comprises a self-positioning module used for receiving an original audio signal, outputting a self-supervised voice feature vector and a traditional audio feature vector in parallel, and sending the self-supervised voice feature vector and the traditional audio feature vector to a feature normalization conversion module; the feature normalization conversion module is used for mapping the self-supervised voice feature vector and the traditional audio feature vector to a standard speaker feature space and outputting a normalized feature; the perception modeling module performs multi-scale time sequence coding through an improved Transform structure, and outputs a voice semantic probability distribution sequence; the CTC loss module is used for optimizing the acoustic model according to the voice semantic probability distribution sequence; the collaboration unit is used for receiving multiple paths of original audio features, screening credible channels from the obtained synchronization features, and outputting corrected features; and the fusion filtering module is used for receiving the local features and the corrected features, generating global probability distribution through attention weight fusion, and decoding the global probability distribution into a final text sequence.
Owner:SHENYANG LIGONG UNIV

Speech synthesis method and device, computer equipment and storage medium

The invention discloses a speech synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-mode background sound condition input data; performing modal integrity detection on the multi-modal background sound condition input data to obtain a detection result; generating an environment background sound feature embedding vector according to a detection result; obtaining to-be-synthesized text data and speaker reference audio data, and performing feature extraction to obtain text semantic features and speaker timbre features; inputting the environment background sound feature embedded vector, the text semantic feature and the speaker timbre feature into an acoustic model to generate a Mel spectrum; and converting the Mel spectrum into a target voice waveform to obtain synthetic voice data. By implementing the method, scene requirements can be deeply matched, diversified scene types can be covered, accurate matching of background sounds and voice semantics is realized, and the technical scheme can be applied to the fields of finance and medical health.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice recognition method based on acoustic model, computer equipment and storage medium

The invention belongs to the field of voice recognition, and discloses a voice recognition method based on an acoustic model, computer equipment and a storage medium. The method comprises the following steps: acquiring voice features of to-be-recognized voice; inputting the voice features into an acoustic model, and outputting a recognition result by the model; wherein the time sequence processing network layer firstly determines the ratio of current input future frames needing to be pre-watched to context information through a pre-trained gating fusion unit, then calculates the number of the future frames needing to be pre-watched based on the ratio, obtains the corresponding future frames, calculates long-time context representation in combination with the future frames, processes the long-time context representation and outputs the long-time context representation to the next layer of network. According to the method and the device, the problem of static binding of delay and accuracy in the prior art is solved by dynamically adjusting the number of the future frames to be pre-watched, low-delay response to simple command words is realized, the recognition accuracy is improved through multiple future frames to be pre-watched for easily-confused instructions, the balance of the delay and the accuracy is realized, and the performance of a voice recognition system and the user experience are improved.
Owner:深圳市友杰智新科技有限公司

Corpus expansion method, system and equipment based on speech synthesis and medium

The invention relates to the technical field of speech synthesis, in particular to a speech synthesis-based corpus expansion method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing including data annotation based on the collected audio and corresponding text of a target speaker; extracting acoustic features from the preprocessed audio; on the basis of a pre-trained acoustic model, performing personalized fine tuning by using the annotation data and the acoustic features, and training personalized acoustic models of a plurality of speakers at the same time through multi-thread parallel computing; calling the trained personalized acoustic model, and synthesizing a voice corpus of the target text in combination with a vocoder; and based on the trained personalized acoustic model, continuously expanding the corpus by changing the text. The personalized voice corpus is quickly generated through a small number of voice samples, the data acquisition cost is remarkably reduced, and the corpus construction efficiency is improved.
Owner:深圳市友杰智新科技有限公司

Method of recognizing speech, device, and medium

A method of recognizing a speech, a device, and a medium. The method includes: processing, by using an acoustic model, speech data to be recognized and a first text segment obtained by recognition to obtain respective acoustic probabilities of a plurality of candidate text segments; processing the first text segment by using a first language sub-model to obtain respective initial language probabilities of the plurality of candidate text segments; processing the first text segment by using a constraint sub-model to obtain extendibility relationships of the plurality of candidate text segments with respect to the first text segment; adjusting the initial language probabilities of the candidate text segments according to the extendibility relationships to obtain respective first language probabilities of the plurality of candidate text segments; and determining a target text segment from the plurality of candidate text segments according to the first language probabilities and the acoustic probabilities.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Vehicle-mounted voice interaction method and system and readable storage medium

The invention relates to the technical field of intelligent vehicle-mounted systems, and discloses a vehicle-mounted voice interaction method and system and a readable storage medium, and the method comprises the steps: synchronously collecting initial voice and video data in a vehicle-mounted environment; performing wake-up word detection through a local acoustic model, and based on the detection confidence, extracting a mouth shape visual feature sequence by using a mouth shape recognition model to perform mouth shape verification so as to obtain a wake-up state and sound source positioning information; activating an interaction module at a corresponding position, and performing semantic recognition on the collected interaction voice and video data through a local model and a cloud model respectively; and finally, carrying out fusion cross validation on the local semantic recognition result and the cloud semantic recognition result to generate a final semantic recognition instruction, and executing corresponding operation by the vehicle-mounted system. According to the method, the recognition accuracy, the response speed and the robustness of vehicle-mounted voice interaction in a complex environment are improved, the false wake-up rate is effectively reduced, and the user experience is optimized.
Owner:深圳海冰科技有限公司

Intelligent quality inspection system for high-noise environment language service quality

The invention provides a high-noise environment language service quality intelligent quality inspection system. The intelligent quality inspection system for the language service quality in the high-noise environment comprises a noise sensing module for extracting noise features by performing short-time Fourier transform and spectrogram analysis on an input voice signal; the self-adaptive speech enhancement module is used for dynamically combining the subunits based on the noise characteristics so as to enhance human voice signals; and the noise conditional automatic voice recognition module is used for establishing an acoustic model based on the noise features and the enhanced human voice signals. The intelligent quality inspection system for the language service quality in the high-noise environment not only improves the reliability of recognition, but also can reduce recognition errors caused by noise, lays a solid foundation for multi-dimensional quality inspection, has the advantages of real-time performance, multiple dimensions, automation and self-adaption, and is suitable for popularization and application. And the monitoring efficiency and accuracy of the voice service quality in the high-noise environment are remarkably improved.
Owner:ZHONGYI CLOUD (BEIJING) INTERNET OF THINGS TECH CO LTD

System and method for improving speech synthesis speed of diffusion model

The invention provides a system and method for improving the speech synthesis speed of a diffusion model, and the system comprises an acoustic model which accelerates the speech synthesis process based on condition flow matching, and a data-driven and trainable vocoder, the acoustic model is in signal connection with the vocoder, and the sampling rate outputted by the vocoder is matched with the acoustic model; a main body of the acoustic model comprises a diffusion model used for generating high-quality voice through forward diffusion noise adding and backward diffusion noise removing processes. According to the method, the acoustic features can be generated with fewer iterations, the acoustic features are transmitted to the trained or fine-tuned vocoder to synthesize the voice signals, the voice synthesis speed is increased, the iterations of an acoustic model in voice synthesis can be reduced to one or two, and high-quality voice signals can be generated; the speech synthesis quality is not reduced, the contradiction that the speech synthesis speed and the speech synthesis quality cannot be achieved at the same time in the prior art is effectively solved, and the speech synthesis efficiency is improved.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Intelligent customer service automatic-to-manual switching method and system based on emotion recognition

The invention relates to the technical field of artificial intelligence, in particular to an intelligent customer service automatic-to-manual switching method and system based on emotion recognition, and the method comprises the steps: obtaining a voice signal when a user carries out a call through terminal equipment; and extracting acoustic characteristic parameters based on the voice signal, inputting the acoustic characteristic parameters into a pre-constructed emotion recognition model, outputting an emotion state level, correcting the primary emotion state level based on an interaction influence coefficient and a semantic emotion score change coefficient, and generating a target emotion level. According to the intelligent customer service automatic-to-manual switching method based on emotion recognition, a multi-dimensional correction and judgment mechanism is introduced on the basis of traditional emotion recognition, a more accurate and reasonable customer service switching decision is achieved, the primary emotion state is recognized based on acoustic features, and the customer service switching efficiency is improved. And the emotion level is corrected in combination with a semantic emotion score change coefficient and an interaction influence coefficient, so that identification deviation caused by a single acoustic model is avoided.
Owner:SHANGHAI XINDIAN PHOTOELETRON TECH CO LTD

Wind-noise-resistant self-adaptive volume adjusting method and system for riding earphone

The invention relates to the technical field of audio signal processing, and discloses a wind-noise-resistant self-adaptive volume adjustment method and system for a riding earphone, and the method comprises the following steps: S1, collecting environment sound and a source audio signal in parallel; s2, analyzing environment sound, and determining macroscopic adjustment intensity; s3, analyzing the source audio in parallel, and identifying a content type and a transient signal of the source audio; s4, based on the psychological acoustic model, generating a frequency-related target method tone quality compensation gain; s5, performing dynamic smoothing and frequency weighting on the target gain according to the content and the transient characteristics to generate an application gain; s6, fusing the macroscopic adjustment intensity and the application gain, and synthesizing a final gain parameter; and S7, applying the final gain to the source audio, and outputting the compensated audio. According to the method, the psychological acoustic model and audio content parallel analysis are fused, frequency-level accurate compensation is realized, the problems of traditional degradation and sudden hearing sense change are solved, and the audio definition and comfort are improved.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD

Speech recognition method and system based on multi-modal fusion and adaptive noise modeling, and storage medium

The invention discloses a speech recognition method and system based on multi-modal fusion and adaptive noise modeling, and a storage medium, and belongs to the technical field of speech recognition. The method comprises the following steps: S1, synchronously acquiring multi-modal data; s2, performing voice preprocessing, and extracting an acoustic feature sequence; s3, extracting a visual lip shape feature sequence; s4, generating a low-dimensional noise vector; s5, performing cross-modal time alignment to obtain a multi-modal feature sequence corresponding to time; s6, unifying the semantic embedding space; s7, dynamically calculating and fusing the multi-modal features in the unified embedding space through an adaptive weighted fusion network to obtain a fusion feature vector; s8, noise self-adaptive denoising is carried out; s9, acoustic modeling is carried out through the end-to-end acoustic model, and a probability sequence is output; s10, decoding the probability sequence by the language model to obtain an initial text sequence; and S11, performing post-processing on the initial text sequence. The method can be adaptive to environmental noise, dynamic fusion between modes is realized, and high robustness is kept.
Owner:GUANGZHOU MARITIME INST

English pronunciation error correction training method based on speech recognition

The invention relates to the technical field of speech recognition and processing, in particular to an English pronunciation error correction training method based on speech recognition, and the method comprises the following steps: S1, collecting a speech signal generated by a learner in a pronunciation training process, digitalizing the speech signal, associating the digitalized speech signal with a target standard text, and generating an original audio data record with a timestamp; according to the invention, phoneme-level decoding is carried out on the voice signal by using the recurrent neural network acoustic model, and accurate alignment of the pronunciation of the learner and the standard phoneme sequence is realized in combination with the dynamic time warping algorithm, so that pronunciation errors such as misreading, missed reading and increased reading can be accurately identified; meanwhile, acoustic features such as Mel frequency cepstrum coefficient, pitch and fundamental frequency are extracted to be quantitatively compared with a standard native language pronunciation database, multi-dimensional evaluation covering accuracy, integrity, fluency and rhythm is generated, and the accuracy and systematicness of oral English pronunciation error correction are remarkably improved.
Owner:吕丽沙

A pronunciation evaluation method and system, and an electronic device

The application provides a pronunciation evaluation method and system and an electronic device. The method comprises the following steps: obtaining to-be-evaluated audio and to-be-evaluated text corresponding to the to-be-evaluated audio; obtaining a word-phoneme sequence, wherein the word-phoneme sequence comprises a plurality of words, a phoneme sequence of an American phonetic alphabet sequence corresponding to each word under a unified phoneme rule, and a phoneme sequence of a British phonetic alphabet sequence corresponding to each word under the unified phoneme rule; the unified phoneme rule comprises a one-to-one correspondence relationship between the British phonetic alphabet and the phoneme and a one-to-one correspondence relationship between the American phonetic alphabet and the phoneme; and in the unified phoneme rule, at least one phoneme corresponds to one British phonetic alphabet and one American phonetic alphabet; constructing a decoding network based on the word-phoneme sequence and the to-be-evaluated text; obtaining an acoustic model; and outputting evaluation information based on the to-be-evaluated audio, the acoustic model and the decoding network. The method does not need to use an American pronunciation evaluation model and a British pronunciation evaluation model for evaluation, thereby reducing the complexity of the evaluation mode.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

An acoustic model training method and apparatus

The application provides an acoustic model training method and device. The acoustic model training method provided by the application comprises: acquiring unannotated acoustic samples; pre-training the acoustic model based on a self-supervised method, the acoustic model being used to predict the category of an input acoustic signal, in the self-supervised method, respectively generating masks for the time domain feature and the frequency domain feature of the acoustic signal based on the time domain feature and the frequency domain feature of the acoustic signal, the mask position and the mask quantity of different samples being different; dividing the levels to which the network structures of the pre-trained acoustic model belong, respectively freezing the layers of different levels, and asynchronously fine-tuning the acoustic model, the parameter freezing time of the layers of different levels being different, and the learning rate of the layers of different levels being different. The acoustic model training method and device provided by the application not only reduce the dependence on large-scale artificial annotation data, but also improve the training efficiency and task performance, and better adapt to task requirements and complex application scenarios.
Owner:HANGZHOU XUNSHENG MEDICAL TECHNOLOGY CO LTD

Acoustic and natural language processing models for velocity-based screening and monitoring of behavioral health

PendingJP2026136266AAcoustic modelBehavioural health
This invention provides an acoustic natural language processing model for predicting whether a person has a behavioral or mental health condition based on input speech. [Solution] A method for detecting behavioral or mental health status using an acoustic model including an encoder and a classifier, comprising: (a) acquiring a speech sample comprising a plurality of speech segments; (b) processing the speech sample with an encoder to generate an abstract feature representation, wherein the encoder is pre-trained to perform a first task other than detecting behavioral or mental health status; and (c) processing the abstract feature representation with a classifier to generate an output indicating whether or not the person has a behavioral or mental health status, wherein the classifier is trained on a training dataset comprising a plurality of speech samples from a plurality of speakers, and the speech samples are labeled as originating from or not originating from a speaker with a health status.
Owner:ELLIPSIS HEALTH INC

Speech recognition method and server

The application relates to a speech recognition method and a server. The method comprises the following steps: obtaining a to-be-recognized speech signal; recognizing each frame of the to-be-recognized speech signal according to an acoustic model of each language, and respectively outputting corresponding language phonemes and prediction probabilities; wherein the acoustic model of each language is respectively constructed according to shared hidden layer training; sequentially traversing a sentence decoding graph and a multi-language slot decoding graph connected with each other to obtain a corresponding path; wherein the sentence decoding graph is used for decoding phonemes entering a non-slot, and the slot decoding graph is used for decoding phonemes entering a slot; when it is determined that the path passes through the multi-language slot decoding graph in the speech decoding graph, screening the path according to the prediction probabilities of the language phonemes corresponding to each language, and determining the text information corresponding to the target path as a speech recognition result. The scheme provided by the application can accurately recognize mixed multi-language speech information.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Speech recognition method and device, equipment and storage medium

The invention relates to the technical field of speech recognition, and discloses a speech recognition method and device, equipment and a storage medium. Aiming at the defects of fixed filter coefficient, static context modeling and difficulty in deep training and deployment of the existing deep feed-forward sequence memory network (DFSMN), the method comprises the following steps of: acquiring a frame-level acoustic feature sequence of input voice, and inputting an acoustic model containing a neural network unit; in the unit, the filter weight of the memory module is generated in real time through a parameter generation network based on the intermediate feature of the current frame; performing weighted aggregation on the historical and future frame features by using the weight to obtain a context enhancement feature; and outputting features based on the feature generation unit, and decoding to obtain an identification result. According to the method, context dynamic self-adaptive modeling is realized, the recognition precision and robustness are improved, the efficient reasoning characteristic of a pure feed-forward network is kept, the calculation overhead is not remarkably increased, and the method is adaptive to a server and terminal equipment and is suitable for multi-scene speech recognition of keywords, command words and the like.
Owner:WUXUE GUANGJI INTELLIGENT BODY SOFTWARE TECHNOLOGY CO LTD

Road totally-enclosed sound barrier sound field prediction method and system

The invention relates to the technical field of environmental acoustics, discloses a road totally-enclosed sound barrier sound field prediction method and system, belongs to the field of environmental acoustics, and has very high practicability in sound field prediction and noise pollution evaluation of public road traffic noise after a totally-enclosed sound barrier is implemented. On the basis of the acoustic propagation principle, a new prediction method is provided for a road traffic noise source and a totally-closed sound barrier, a sound field is divided into two subsystems including a barrier internal space and an external space, and equivalent acoustic models of the two subsystems and a calculation method and recommendation parameters of the noise source sound power level are given respectively. According to the method, the sound environment of the high-rise building close to the urban road can be accurately predicted and calculated, and a basis is provided for noise pollution evaluation and control. And meanwhile, the method can be matched with a national standard method or commercial acoustic software for use, and has relatively high applicability.
Owner:CHINA SHIPPING ENVIRONMENT SCI & TECH (SHANGHAI) CO LTD

Speech signal recognition methods, devices, electronic equipment and storage media

This application proposes a speech signal recognition method, apparatus, electronic device, and storage medium. The method includes: acquiring first speech signals from multiple channels, wherein the first speech signals from each channel are raw speech signals synchronously acquired within a set time period; inputting the first speech signals from multiple channels into a trained first acoustic model to obtain a corresponding first phoneme sequence; and recognizing the first phoneme sequence to obtain speech content. This method realizes recognition based on global information of the first speech signals from multiple channels to obtain speech content, achieving low signal distortion and high purity, and improving the quality of speech content.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Data processing method and device, equipment and computer readable storage medium

ActiveCN116469378BImprove translation qualityfast trainingNatural language translationBiological modelsPattern recognitionGoal recognition
The application discloses a data processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: obtaining sample voice information and sample text information; processing the sample voice information through an acoustic model in an initial recognition model to obtain acoustic feature information; processing the acoustic feature information and the sample text information through a translation model in the initial recognition model to obtain a first predicted translation result and a second predicted translation result respectively; training the initial recognition model based on the first predicted translation result and the second predicted translation result to obtain a target recognition model; and the K vector and the V vector in the acoustic model are spliced with a prefix vector and / or the K vector and the V vector in the translation model are spliced with a prefix vector. The training efficiency of the target recognition model is improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

System and method for neural network multilingual speech recognition

Systems, methods, and computer-readable storage devices are disclosed for improved recognition of multiple languages in audio data. One method including: receiving a trained split head multilingual neural network model, the trained split head multilingual neural network model including shared acoustic model layers and a plurality of projection layers, each projection layer of the plurality of projection layers corresponding to a language that the trained split head multilingual neural network model recognizes; receiving audio data, the audio data including speech in a plurality of languages in the audio data, the speech in the plurality of languages corresponding the language recognized by a projection layer of the plurality of projection layers of the trained split head multilingual neural network model; and classifying one or more languages of the speech of the audio data using the trained split head multilingual neural network model.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Self-adaptive voice dialogue generation system of intelligent toy and use method of self-adaptive voice dialogue generation system

The invention relates to a self-adaptive voice dialogue generation system of an intelligent toy and a use method thereof, and belongs to the technical field of voice dialogues of intelligent toys, the self-adaptive voice dialogue generation system comprises a voice acquisition module, a voice recognition module, a user portrait module, a dialogue management module, a natural language generation module, a safety filtering module and a voice synthesis module, the voice acquisition module acquires user voice signals, the voice recognition module is connected with the voice acquisition module and converts the voice signals into text information, and the user portrait module establishes and dynamically updates personalized user portraits corresponding to users. The user portrait at least comprises a cognitive level quantized value, an interest preference vector and an emotion mode label which are calculated on the basis of interaction historical data. According to the method, a high-fault-tolerance voice processing chain is formed from directional pickup of a beam forming algorithm at a hardware end to a CNN-LSTM acoustic model optimized for children at a software end, so that the smoothness and the stability of an interaction process in a real children use scene are fundamentally guaranteed.
Owner:XUZHOU HUAPEI INTELLIGENT MFG TECH CO LTD