Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Word error rate" patented technology

Word error rate (WER) is a common metric of the performance of a speech recognition or machine translation system. The general difficulty of measuring performance lies in the fact that the recognized word sequence can have a different length from the reference word sequence (supposedly the correct one). The WER is derived from the Levenshtein distance, working at the word level instead of the phoneme level.

Multi-mode voice interaction system and method for intelligent cockpit

PendingCN121862092ASpeech recognitionSpeech synthesisModal voiceEngineering
The invention discloses a multi-mode voice interaction system and method for an intelligent cabin. Firstly, various vehicle working condition parameters in a CAN bus are read in real time, a wavelet domain Wiener filtering and judgment statistic enhancement strategy is dynamically adjusted, and self-adaptive voice purification based on vehicle state prior is achieved. Secondly, a bidirectional long-short-term memory network is adopted as a semantic understanding core, semantic feature vectors are extracted from enhanced signals through segmented maximum pooling operation, and the problem of semantic confusion caused by manifold collapse in a traditional method is effectively solved. And the feature vector and a safety state signal are synchronously input into an improved LLM-CoT engine and a dynamic arbiter, so that context accurate understanding and safety-up priority response are realized. According to the method, vehicle working condition perception and biLSTM depth time sequence modeling are deeply adapted to a vehicle-mounted complex acoustic scene, the error rate of semantic understanding is remarkably reduced in actual measurement under multi-person dialect dialogue, and the accuracy, robustness and safety of intelligent cabin interaction are systematically improved.
Owner:SOUTHEAST UNIV

Speech recognition optimization method and system based on adaptive dynamic programming

The invention provides a speech recognition optimization method and system based on adaptive dynamic programming, and relates to the technical field of speech recognition. The method comprises the following steps: firstly, acquiring voice data, preprocessing the acquired voice data, and constructing an acoustic feature sequence; secondly, constructing an acoustic state transition model, modeling a speech recognition process as an optimal control problem of a finite time domain, and determining an optimization target and strategy distribution; then, through two-stage training of a conform speech recognition model, strategy distribution is learned and optimized; in the first stage, a connection time sequence classification CTC framework is adopted for pre-training, and in the second stage, an adaptive dynamic programming ADP algorithm and a self-evaluation sequence training method are introduced for training; and finally, performing voice recognition by adopting the trained conform voice recognition model, and outputting a recognition result. The method is easy to implement and high in compatibility with an existing ASR architecture, and the word error rate is remarkably reduced while stability is guaranteed.
Owner:NORTHEASTERN UNIV CHINA

Self-adaptive voice semantic communication method based on hierarchical time sequence importance

The invention relates to a self-adaptive voice semantic communication method based on hierarchy time sequence importance, which belongs to the field of voice analysis, and comprises the following steps: a sending end performs priority division on a discrete feature matrix according to hierarchy and time sequence attributes of voice features, and screens the features according to importance scores in a feature matrix packaging and selecting link; the method comprises the following steps: interacting with wireless communication through a channel adaptive scheduling module, introducing a channel feedback mechanism, and dynamically adjusting transmission power and a resource allocation strategy by sensing current channel state information; and after completing signal demodulation, a receiving end inputs the acquired sparse feature flow into a voice restoration module, and globally reconstructs the received features by using voice priori knowledge of deep pre-training. According to the method, a hierarchical speech feature extraction technology based on a discrete codebook and a generative semantic repair technology are combined, so that the speech word error rate in a severe channel environment is reduced to the maximum extent, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Detecting unintended memorization in language-model-fused ASR systems

A method includes inserting a set of canary text samples into a corpus of training text samples and training an external language model on the corpus of training text samples and the set of canary text samples inserted into the corpus of training text samples. For each canary text sample, the method also includes generating a corresponding synthetic speech utterance and generating an initial transcription for the corresponding synthetic speech utterance. The method also includes rescoring the initial transcription generated for each corresponding synthetic speech utterance using the external language model. The method also includes determining a word error rate (WER) of the external language model based on the rescored initial transcriptions and the canary text samples and detecting memorization of the canary text samples by the external language model based on the WER of the external language model.
Owner:GOOGLE LLC

Training for long-form speech recognition

A method includes obtaining a set of training samples, wherein each training sample includes a corresponding sequence of speech segments corresponding to a training utterance and a corresponding sequence of ground-truth transcriptions for the sequence of speech segments, and wherein each ground-truth transcription includes a start time and an end time of a corresponding speech segment. For each training sample in the set of training samples, the method includes processing, using a speech recognition model, the corresponding sequence of speech segments to obtain one or more speech recognition hypotheses for the training utterance; and, for each speech recognition hypothesis obtained for the training utterance, identifying a respective number of word errors relative to the corresponding sequence of ground-truth transcriptions. The method trains the speech recognition model to minimize word error rate based on the respective number of word errors identified for each speech recognition hypothesis obtained for the training utterance.
Owner:GOOGLE LLC

Medical science popularization live broadcast dialogue transcription method and device based on large language model

The invention relates to the technical field of medical information technology and natural language processing, in particular to a medical science popularization live broadcast dialogue transcription method and device based on a large language model. And the final accuracy and availability of the medical science popularization live broadcast dialogue transcription can be obviously improved. According to the method, the inherent defects of a traditional ASR system in terms recognition and speaker separation in the medical industry are effectively overcome, and the word error rate and the medical concept error rate are greatly reduced, so that the medical knowledge spreading risk and the high manual calibration cost caused by transcription errors are reduced. Meanwhile, according to the scheme, complex structure modification or additional data training does not need to be carried out on a large language model, the inherent text understanding and generating capacity of the large language model is fully utilized, and efficient production and propagation of high-quality medical science popularization content are greatly promoted.
Owner:BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD

A performance analysis method of a graph-based Transformer automatic speech recognition model

The application relates to a performance analysis method of a graph-based Transformer automatic speech recognition model, and belongs to the field of artificial intelligence and speech recognition. The method comprises the following steps: obtaining a Transformer automatic speech recognition model; obtaining audio data; inputting the audio data into the Transformer automatic speech recognition model, obtaining weight matrices of multiple attention heads of each layer in the model through forward propagation, and extracting word texts output by the model; performing average processing on the weight matrices of each attention head within a given time to obtain artificial neural activities of the attention head; performing correlation calculation on the artificial neural activities of the attention head by using a Pearson correlation coefficient to obtain a correlation coefficient, constructing a functional connection matrix based on the correlation coefficient; calculating graph theory parameters of the functional connection matrix; calculating a word error rate of the output word texts; and analyzing the performance of the Transformer automatic speech recognition model based on the graph theory parameters and the word error rate. The method provides a basis for analyzing the performance of the Transformer automatic speech recognition model.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

A method and device for generating an ASR audio corpus based on a multi-modal large model

The application discloses a method and device for generating ASR audio corpus based on a multimodal large model, and relates to the field of audio corpus. In the method, a semantic vector and a conditional vector are spliced into a joint vector to generate first speech; target noise is selected from a preset noise library according to a scene label, and the target noise is superimposed on the first speech to generate noise-bearing speech, and an adversarial noise is injected to generate second speech; the second speech is subjected to noise labeling, text labeling, emotion labeling and speaker labeling, and is aligned to generate a multimodal labeling file; according to the scene label, the noise type and the speaker information of the multimodal labeling file, a word error rate threshold and a semantic similarity threshold are set, and target corpus is selected from the multimodal labeling file according to the word error rate threshold and the semantic similarity threshold. By implementing the technical scheme provided in the application, high-quality audio corpus that meets specific requirements and is effectively screened can be generated.
Owner:ZHIMING RIXIN (NANJING) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Speech synthesis method and device, electronic equipment and storage medium

The invention relates to a speech synthesis method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-synthesized text and prompt speech; performing duration prediction based on a phoneme sequence of the to-be-synthesized text and the prompt voice to obtain playing duration information of a target synthesized voice; adjusting the sequence length of the phoneme sequence based on the playing time length information to obtain a target phoneme sequence; the sequence length of the target phoneme sequence is matched with the playing duration information; performing feature extraction on the target phoneme sequence, and generating target semantic features based on a feature extraction result and prior distribution; target acoustic features are generated based on the target semantic features and the prior distribution, and the target synthetic speech is generated based on the target acoustic features. The speech synthesis speed is improved, the naturalness of the synthetic speech is high, the word error rate is low, and the quality of the synthetic speech is greatly improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Off-line embedded speech recognition system and method based on Zipform model and application

The invention belongs to the technical field of deep learning, edge calculation and speech recognition crossing, and particularly discloses an off-line embedded speech recognition system and method based on a Zipform model and application. According to the off-line embedded speech recognition system and method based on the Zipform model, the long-term and short-term dependency modeling capacity is enhanced through a Zipform-GRU fusion architecture, and the non-linear attention mechanism, INT8 quantization and other technologies are combined, so that the speech recognition efficiency is improved. The problems of high calculation complexity and low identification efficiency of embedded equipment are solved. The system performs joint decoding through a finite state converter (FST) and a language model, the low word error rate of 4.26% is achieved on an Aishell-1 data set, the inference delay of an embedded platform is smaller than or equal to 500 ms, and the real-time requirement is met. The scheme of the invention can be widely applied to embedded equipment scenes with limited resources, such as aeronautical communication, intelligent terminals and the like, improves the energy efficiency and reliability of edge equipment voice interaction, and has remarkable commercial value and industrial upgrading promotion effect.
Owner:SHAANXI FENGHUO ELECTRONICS

Continuous sign language recognition method and system based on hierarchical attention feature fusion

This invention proposes a continuous sign language recognition method and system based on hierarchical attention feature fusion. The method involves processing sign language videos, outputting multi-level, multi-scale feature maps through a visual encoder. These maps are then processed by a cross-scale aggregation module containing a hierarchical attention network before being fed into a classifier for word classification. The CTC decoder outputs a word sequence. Simultaneously, a two-layer fully connected network performs binary classification prediction to determine whether each frame is a boundary frame for a sign language word. An online self-distillation mechanism based on exponential moving average (EMA) is established, using an exponential moving average version of the model parameters as the EMA teacher model. The student model continuously learns to obtain a more stable word sequence output. A multi-level regularization framework is constructed, coordinating optimization at four levels: label space, temporal space, feature space, and parameter space. This invention effectively improves the model's generalization ability under small sample conditions, thereby reducing the word error rate (WER) in continuous sign language recognition.
Owner:TIANJIN POLYTECHNIC UNIV

An adaptive speech semantic communication method based on hierarchical temporal importance

The application relates to an adaptive voice semantic communication method based on hierarchical time sequence importance, and belongs to the voice analysis field, which comprises the following steps: a sending end divides a discrete feature matrix according to the priority of voice features and time sequence attributes, and filters the features according to the importance score in the feature matrix encapsulation and selection link; a channel adaptive scheduling module is interacted with wireless communication, a channel feedback mechanism is introduced, the current channel state information is perceived, and the transmission power and resource allocation strategy are dynamically adjusted; after signal demodulation is completed, a receiving end inputs the obtained sparse feature stream to a voice repair module, and global reconstruction is performed on the received features by using deep pre-trained voice prior knowledge. The application combines a hierarchical voice feature extraction technology based on a discrete codebook and a generative semantic repair technology, so that the voice word error rate in a severe channel environment is maximally reduced, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

A multi-sound source separation system and method based on a generative adversarial network

The application discloses a multi-sound source separation system and method based on a generative adversarial network, which comprises a sound signal acquisition and processing all-in-one machine, a server platform, and a multi-sound source separation client. The sound signal acquisition and processing all-in-one machine is used for collecting original audio data emitted by multiple sound sources and processing the original audio data to determine the number of sound sources. The server platform determines whether to process the original audio data according to the determination result of the number of sound sources. If the number of sound sources is one, the original audio data is not processed. If the number of sound sources is multiple, the original audio data is subjected to cross-correlation to determine the position of each sound source, and then the multiple sound source signals are separated through adaptive beamforming. The speech signals after beamforming are optimized by using a generative adversarial network, and the output image data of the generative adversarial network is restored to audio data. The multi-sound source separation client acquires the separation result of the server platform on the multiple sound source signals. The application can effectively separate the speech signals of multiple sound sources, and the speech signals can obtain a higher speech recognition accuracy and a lower word error rate after being recognized.
Owner:NANJING UNIV

Method and system for evaluating voice recognition and directional sound transmission performance of open wireless earphone

PendingCN122266395AElectrical apparatusSpeech analysisNoiseSpeech recognition performance
The application discloses an open wireless earphone voice recognition and directional sound transmission performance evaluation method and system, relates to the technical field of earphone equipment, and comprises the following steps: voice recognition test preparation is carried out through a convolutional neural network noise suppression and an adaptive audio feedback mechanism, and a dynamic voice signal library is constructed to collect earphone voice signals; a word error rate is calculated based on the collected earphone voice signals, voice recognition evaluation is carried out in combination with short-time objective voice intelligibility and syllable rate estimation, and sound field directional sound transmission testing is carried out in combination with a directional gain formula. The application accurately evaluates the similarity between enhanced voice and real voice through multi-band filtering, syllable rate analysis and short-time objective voice intelligibility scoring, quantifies the voice recognition effect in combination with the word error rate, establishes the relationship between the signal-to-noise ratio and the voice recognition performance, and provides a comprehensive and dynamic performance evaluation system.
Owner:GUANGZHOU VIKEN COMM TECH CO LTD +1

Energy storage operation and maintenance multi-modal dialogue system and method based on large language model

The invention relates to the technical field of energy storage power station operation and maintenance, large language model reasoning and multi-mode man-machine interaction, in particular to an energy storage operation and maintenance multi-mode dialogue system and method based on a large language model, and the system comprises five layers of architectures: a data collection and multi-mode perception layer obtains BMS data, images and voice; a semantic fusion and dynamic Prompt layer constructs a Prompt containing a real-time working condition; an edge reasoning layer quantifies a 7B parameter LLM by using Q4 and is combined with LoRA fine tuning, and runs in a 10W Edge GPU; the security layer performs confidence evaluation, white list and intention classification filtering; the human-computer interaction layer supports local ASR and self-adaptive TTS, according to the energy storage operation and maintenance multi-mode dialogue system and method based on the large language model, the first response is smaller than or equal to 1.9 s, the Top-1 accuracy is larger than or equal to 92.4%, the voice bandwidth is smaller than or equal to 128 kb / s, and the word error rate is 1t; and off-line use can be realized, an unauthorized instruction is blocked, and the operation and maintenance efficiency of the energy storage power station is remarkably improved.
Owner:ALPHA ESS CO LTD

Evaluation platform and evaluation method for Chinese reading and spoken language annotation

The invention discloses a Chinese reading and spoken language annotation evaluation platform and evaluation method, and belongs to the technical field of intelligent education evaluation. The platform adopts a five-layer architecture, wherein a client layer comprises a teacher terminal, a student terminal, an access layer, a service layer, a data processing layer and a storage layer. The AI automatically generates a TextGrid format three-layer annotation file conforming to the Praat standard, namely a phoneme layer, a tone layer and an error layer, and forms a closed-loop optimization mechanism with manual review of a teacher. The platform aligns the voice features with the text through a forced alignment algorithm, and automatically marks pronunciation errors; and after the teacher corrects in the Praat, the difference comparison module calculates an error and triggers model incremental learning, and the recognition accuracy is continuously optimized. The platform solves the problems that in the prior art, non-native language speech recognition robustness is poor, automatic scoring and manual review are separated, and fine-grained feedback is lacked, marking efficiency is remarkably improved, the workload of manual review is reduced by about 70%, and the word error rate of non-standard accent can be reduced by 15%-25% through iterative optimization.
Owner:EAST CHINA NORMAL UNIV

Encoding and decoding method, encoder and decoder

Embodiments of the present application provide a method and device for encoding and decoding. The method comprises: obtaining a first code block (CB) and a second CB, wherein the first CB and the second CB are encoded using the same scheme, the first CB and the second CB correspond to different MIMO antennas respectively, the first CB comprises m information bits, and the m information bits are mapped to the second CB; and transmitting the first CB and the second CB. According to the present application, the technology of using related bits between multiple CBs in a transmission block is applied to a MIMO transmission system, wherein the multiple CBs are from different MIMO antennas respectively, which helps to improve the accuracy of CB transmission in the MIMO transmission system and reduce the word error rate (WER) of the CB.
Owner:HUAWEI TECH CO LTD

Method and device for evaluating speech recognition accuracy

The invention discloses a method and device for evaluating speech recognition accuracy, and relates to the technical field of smart home / smart home, and the method comprises the steps: obtaining a corpus test set and a standard corpus text set table which is manually labeled based on the corpus test set, calling a language-to-text model to recognize and convert the voices in the corpus test set to obtain a first corpus text set; evaluating the first corpus text set based on a second corpus text set in the standard corpus text set table to obtain an evaluation result, and adjusting contents in the standard corpus text set table according to the evaluation result to generate a corpus evaluation table; wherein the corpus evaluation table comprises a word error rate and a sentence error rate of the first corpus text set; the word error rate and the sentence error rate of the first corpus text set are used for evaluating the recognition conversion effect of the language-to-character model. Through the method provided by the invention, the accuracy of speech recognition is automatically evaluated.
Owner:QINGDAO HAIER TECH +2

Electronic scale with functions of healthy weight management, diabetes risk assessment and AI voice interaction

The invention provides an electronic scale with functions of healthy weight management, diabetes risk assessment and AI voice interaction, and the electronic scale has the following technical advantages: (1) multi-modal health assessment: fusing weight, body fat, waistline (laser automatic measurement, error < = 0.5 cm) and clinical parameters (age / gender / family history); accurate evaluation of the diabetes risk is realized through a logistic regression model trained by mass clinical data; (2) AI voice interaction optimization: on the basis of an ASR model (the word error rate is less than or equal to 5%) of a Transform architecture, natural language parameter input is supported, the response time is less than or equal to 1 second, and the operation satisfaction of old users is improved to 92%; (3) full-automatic data acquisition: a weighing module, a body fat module and a waistline module are synchronously triggered, user intervention is not needed, and the data acquisition efficiency is improved; and (4) intelligent data management: local storage supports HIS system docking, and doctors can check three-month trend curves (such as BMI weekly change rate).
Owner:ZHEJIANG UNIV

A method and system for constructing a laos language grammar correction corpus based on error distribution to guide a large model, and an electronic device

The present application relates to a method and system for constructing a Lao grammar correction corpus for guiding a large model based on error distribution, and an electronic device. The present application simulates common grammar errors that actually occur by using an existing speech recognition model; then, a large language model is used to automatically generate data covering multiple error distributions according to rules and constraints; then, the data is cleaned and preprocessed; then, a small language model is used as a correction model to correct the fusion corpus; finally, through model evaluation, a Lao grammar correction corpus with a lower word error rate is selected, thereby effectively solving the problem of a lack of Lao grammar correction corpus. The present application effectively utilizes the error distribution of Lao grammar to guide the large model to generate a Lao grammar correction corpus, and good experimental results have been achieved in the Lao grammar correction task.
Owner:KUNMING UNIV OF SCI & TECH

Multi-dialect speech recognition method based on Transform self-attention mechanism

The invention discloses a multi-dialect speech recognition method based on a Transform self-attention mechanism, and belongs to the technical field of speech recognition, and the method comprises the steps: obtaining a multi-dialect audio data set, and carrying out the preprocessing and feature extraction; using a dialect topological space mapping module to construct topological space representation of dialect features; inputting the topological mapping features into a Transform encoder, and performing encoding processing by using a cross-dialect attention fusion mechanism and a dialect feature dynamic weighting mechanism; according to the method, Riemannian manifold mapping, graph attention fusion and gating dynamic weighting are introduced, the space structure relation of the speech features among dialects is accurately modeled, the recognition accuracy under a multi-dialect scene is remarkably improved, the word error rate is averagely reduced by 18%-25%, and the recognition accuracy of a few dialects is improved by 30%-40%.
Owner:YANGTZE UNIVERSITY

Quality estimation for automatic speech recognition

Methods and systems are provided for implementing quality estimation for automatic speech recognition, and more specifically training an ASR model, and training a QE model to perform word error rate prediction upon the trained ASR model. The ASR model may be a transformer learning model having an architecture including an encoder including multi-head attention layers, and a memory encoder including a masking multi-head attention layer. The QE model may include a binary classification model and a regression model, where the binary classification model is based on a discrete statistical distribution, and the regression model is based on a continuous statistical distribution. Training the ASR model may produce output having variable word error rates, and the QE model may be trained based on empirical word error rates of the ASR model. The QE model may predict performance of the ASR model without labor-intensive labeling to generate ground truth.
Owner:ALIBABA GROUP HOLDING LTD