Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Word error rate" patented technology

Word error rate (WER) is a common metric of the performance of a speech recognition or machine translation system. The general difficulty of measuring performance lies in the fact that the recognized word sequence can have a different length from the reference word sequence (supposedly the correct one). The WER is derived from the Levenshtein distance, working at the word level instead of the phoneme level.

Multi-mode voice interaction system and method for intelligent cockpit

PendingCN121862092ASpeech recognitionSpeech synthesisModal voiceEngineering
The invention discloses a multi-mode voice interaction system and method for an intelligent cabin. Firstly, various vehicle working condition parameters in a CAN bus are read in real time, a wavelet domain Wiener filtering and judgment statistic enhancement strategy is dynamically adjusted, and self-adaptive voice purification based on vehicle state prior is achieved. Secondly, a bidirectional long-short-term memory network is adopted as a semantic understanding core, semantic feature vectors are extracted from enhanced signals through segmented maximum pooling operation, and the problem of semantic confusion caused by manifold collapse in a traditional method is effectively solved. And the feature vector and a safety state signal are synchronously input into an improved LLM-CoT engine and a dynamic arbiter, so that context accurate understanding and safety-up priority response are realized. According to the method, vehicle working condition perception and biLSTM depth time sequence modeling are deeply adapted to a vehicle-mounted complex acoustic scene, the error rate of semantic understanding is remarkably reduced in actual measurement under multi-person dialect dialogue, and the accuracy, robustness and safety of intelligent cabin interaction are systematically improved.
Owner:SOUTHEAST UNIV

Speech recognition optimization method and system based on adaptive dynamic programming

The invention provides a speech recognition optimization method and system based on adaptive dynamic programming, and relates to the technical field of speech recognition. The method comprises the following steps: firstly, acquiring voice data, preprocessing the acquired voice data, and constructing an acoustic feature sequence; secondly, constructing an acoustic state transition model, modeling a speech recognition process as an optimal control problem of a finite time domain, and determining an optimization target and strategy distribution; then, through two-stage training of a conform speech recognition model, strategy distribution is learned and optimized; in the first stage, a connection time sequence classification CTC framework is adopted for pre-training, and in the second stage, an adaptive dynamic programming ADP algorithm and a self-evaluation sequence training method are introduced for training; and finally, performing voice recognition by adopting the trained conform voice recognition model, and outputting a recognition result. The method is easy to implement and high in compatibility with an existing ASR architecture, and the word error rate is remarkably reduced while stability is guaranteed.
Owner:NORTHEASTERN UNIV CHINA

Self-adaptive voice semantic communication method based on hierarchical time sequence importance

The invention relates to a self-adaptive voice semantic communication method based on hierarchy time sequence importance, which belongs to the field of voice analysis, and comprises the following steps: a sending end performs priority division on a discrete feature matrix according to hierarchy and time sequence attributes of voice features, and screens the features according to importance scores in a feature matrix packaging and selecting link; the method comprises the following steps: interacting with wireless communication through a channel adaptive scheduling module, introducing a channel feedback mechanism, and dynamically adjusting transmission power and a resource allocation strategy by sensing current channel state information; and after completing signal demodulation, a receiving end inputs the acquired sparse feature flow into a voice restoration module, and globally reconstructs the received features by using voice priori knowledge of deep pre-training. According to the method, a hierarchical speech feature extraction technology based on a discrete codebook and a generative semantic repair technology are combined, so that the speech word error rate in a severe channel environment is reduced to the maximum extent, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

Detecting unintended memorization in language-model-fused ASR systems

A method includes inserting a set of canary text samples into a corpus of training text samples and training an external language model on the corpus of training text samples and the set of canary text samples inserted into the corpus of training text samples. For each canary text sample, the method also includes generating a corresponding synthetic speech utterance and generating an initial transcription for the corresponding synthetic speech utterance. The method also includes rescoring the initial transcription generated for each corresponding synthetic speech utterance using the external language model. The method also includes determining a word error rate (WER) of the external language model based on the rescored initial transcriptions and the canary text samples and detecting memorization of the canary text samples by the external language model based on the WER of the external language model.
Owner:GOOGLE LLC

Medical science popularization live broadcast dialogue transcription method and device based on large language model

The invention relates to the technical field of medical information technology and natural language processing, in particular to a medical science popularization live broadcast dialogue transcription method and device based on a large language model. And the final accuracy and availability of the medical science popularization live broadcast dialogue transcription can be obviously improved. According to the method, the inherent defects of a traditional ASR system in terms recognition and speaker separation in the medical industry are effectively overcome, and the word error rate and the medical concept error rate are greatly reduced, so that the medical knowledge spreading risk and the high manual calibration cost caused by transcription errors are reduced. Meanwhile, according to the scheme, complex structure modification or additional data training does not need to be carried out on a large language model, the inherent text understanding and generating capacity of the large language model is fully utilized, and efficient production and propagation of high-quality medical science popularization content are greatly promoted.
Owner:BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD

Continuous sign language recognition method and system based on hierarchical attention feature fusion

This invention proposes a continuous sign language recognition method and system based on hierarchical attention feature fusion. The method involves processing sign language videos, outputting multi-level, multi-scale feature maps through a visual encoder. These maps are then processed by a cross-scale aggregation module containing a hierarchical attention network before being fed into a classifier for word classification. The CTC decoder outputs a word sequence. Simultaneously, a two-layer fully connected network performs binary classification prediction to determine whether each frame is a boundary frame for a sign language word. An online self-distillation mechanism based on exponential moving average (EMA) is established, using an exponential moving average version of the model parameters as the EMA teacher model. The student model continuously learns to obtain a more stable word sequence output. A multi-level regularization framework is constructed, coordinating optimization at four levels: label space, temporal space, feature space, and parameter space. This invention effectively improves the model's generalization ability under small sample conditions, thereby reducing the word error rate (WER) in continuous sign language recognition.
Owner:TIANJIN POLYTECHNIC UNIV

An adaptive speech semantic communication method based on hierarchical temporal importance

The application relates to an adaptive voice semantic communication method based on hierarchical time sequence importance, and belongs to the voice analysis field, which comprises the following steps: a sending end divides a discrete feature matrix according to the priority of voice features and time sequence attributes, and filters the features according to the importance score in the feature matrix encapsulation and selection link; a channel adaptive scheduling module is interacted with wireless communication, a channel feedback mechanism is introduced, the current channel state information is perceived, and the transmission power and resource allocation strategy are dynamically adjusted; after signal demodulation is completed, a receiving end inputs the obtained sparse feature stream to a voice repair module, and global reconstruction is performed on the received features by using deep pre-trained voice prior knowledge. The application combines a hierarchical voice feature extraction technology based on a discrete codebook and a generative semantic repair technology, so that the voice word error rate in a severe channel environment is maximally reduced, and the semantic intelligibility of a receiving end is improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

A multi-sound source separation system and method based on a generative adversarial network

The application discloses a multi-sound source separation system and method based on a generative adversarial network, which comprises a sound signal acquisition and processing all-in-one machine, a server platform, and a multi-sound source separation client. The sound signal acquisition and processing all-in-one machine is used for collecting original audio data emitted by multiple sound sources and processing the original audio data to determine the number of sound sources. The server platform determines whether to process the original audio data according to the determination result of the number of sound sources. If the number of sound sources is one, the original audio data is not processed. If the number of sound sources is multiple, the original audio data is subjected to cross-correlation to determine the position of each sound source, and then the multiple sound source signals are separated through adaptive beamforming. The speech signals after beamforming are optimized by using a generative adversarial network, and the output image data of the generative adversarial network is restored to audio data. The multi-sound source separation client acquires the separation result of the server platform on the multiple sound source signals. The application can effectively separate the speech signals of multiple sound sources, and the speech signals can obtain a higher speech recognition accuracy and a lower word error rate after being recognized.
Owner:NANJING UNIV

Method and system for evaluating voice recognition and directional sound transmission performance of open wireless earphone

PendingCN122266395AElectrical apparatusSpeech analysisNoiseSpeech recognition performance
The application discloses an open wireless earphone voice recognition and directional sound transmission performance evaluation method and system, relates to the technical field of earphone equipment, and comprises the following steps: voice recognition test preparation is carried out through a convolutional neural network noise suppression and an adaptive audio feedback mechanism, and a dynamic voice signal library is constructed to collect earphone voice signals; a word error rate is calculated based on the collected earphone voice signals, voice recognition evaluation is carried out in combination with short-time objective voice intelligibility and syllable rate estimation, and sound field directional sound transmission testing is carried out in combination with a directional gain formula. The application accurately evaluates the similarity between enhanced voice and real voice through multi-band filtering, syllable rate analysis and short-time objective voice intelligibility scoring, quantifies the voice recognition effect in combination with the word error rate, establishes the relationship between the signal-to-noise ratio and the voice recognition performance, and provides a comprehensive and dynamic performance evaluation system.
Owner:GUANGZHOU VIKEN COMM TECH CO LTD +1

Energy storage operation and maintenance multi-modal dialogue system and method based on large language model

The invention relates to the technical field of energy storage power station operation and maintenance, large language model reasoning and multi-mode man-machine interaction, in particular to an energy storage operation and maintenance multi-mode dialogue system and method based on a large language model, and the system comprises five layers of architectures: a data collection and multi-mode perception layer obtains BMS data, images and voice; a semantic fusion and dynamic Prompt layer constructs a Prompt containing a real-time working condition; an edge reasoning layer quantifies a 7B parameter LLM by using Q4 and is combined with LoRA fine tuning, and runs in a 10W Edge GPU; the security layer performs confidence evaluation, white list and intention classification filtering; the human-computer interaction layer supports local ASR and self-adaptive TTS, according to the energy storage operation and maintenance multi-mode dialogue system and method based on the large language model, the first response is smaller than or equal to 1.9 s, the Top-1 accuracy is larger than or equal to 92.4%, the voice bandwidth is smaller than or equal to 128 kb / s, and the word error rate is 1t; and off-line use can be realized, an unauthorized instruction is blocked, and the operation and maintenance efficiency of the energy storage power station is remarkably improved.
Owner:ALPHA ESS CO LTD

Evaluation platform and evaluation method for Chinese reading and spoken language annotation

The invention discloses a Chinese reading and spoken language annotation evaluation platform and evaluation method, and belongs to the technical field of intelligent education evaluation. The platform adopts a five-layer architecture, wherein a client layer comprises a teacher terminal, a student terminal, an access layer, a service layer, a data processing layer and a storage layer. The AI automatically generates a TextGrid format three-layer annotation file conforming to the Praat standard, namely a phoneme layer, a tone layer and an error layer, and forms a closed-loop optimization mechanism with manual review of a teacher. The platform aligns the voice features with the text through a forced alignment algorithm, and automatically marks pronunciation errors; and after the teacher corrects in the Praat, the difference comparison module calculates an error and triggers model incremental learning, and the recognition accuracy is continuously optimized. The platform solves the problems that in the prior art, non-native language speech recognition robustness is poor, automatic scoring and manual review are separated, and fine-grained feedback is lacked, marking efficiency is remarkably improved, the workload of manual review is reduced by about 70%, and the word error rate of non-standard accent can be reduced by 15%-25% through iterative optimization.
Owner:EAST CHINA NORMAL UNIV

Encoding and decoding method, encoder and decoder

Embodiments of the present application provide a method and device for encoding and decoding. The method comprises: obtaining a first code block (CB) and a second CB, wherein the first CB and the second CB are encoded using the same scheme, the first CB and the second CB correspond to different MIMO antennas respectively, the first CB comprises m information bits, and the m information bits are mapped to the second CB; and transmitting the first CB and the second CB. According to the present application, the technology of using related bits between multiple CBs in a transmission block is applied to a MIMO transmission system, wherein the multiple CBs are from different MIMO antennas respectively, which helps to improve the accuracy of CB transmission in the MIMO transmission system and reduce the word error rate (WER) of the CB.
Owner:HUAWEI TECH CO LTD

Electronic scale with functions of healthy weight management, diabetes risk assessment and AI voice interaction

The invention provides an electronic scale with functions of healthy weight management, diabetes risk assessment and AI voice interaction, and the electronic scale has the following technical advantages: (1) multi-modal health assessment: fusing weight, body fat, waistline (laser automatic measurement, error < = 0.5 cm) and clinical parameters (age / gender / family history); accurate evaluation of the diabetes risk is realized through a logistic regression model trained by mass clinical data; (2) AI voice interaction optimization: on the basis of an ASR model (the word error rate is less than or equal to 5%) of a Transform architecture, natural language parameter input is supported, the response time is less than or equal to 1 second, and the operation satisfaction of old users is improved to 92%; (3) full-automatic data acquisition: a weighing module, a body fat module and a waistline module are synchronously triggered, user intervention is not needed, and the data acquisition efficiency is improved; and (4) intelligent data management: local storage supports HIS system docking, and doctors can check three-month trend curves (such as BMI weekly change rate).
Owner:ZHEJIANG UNIV

Multi-dialect speech recognition method based on Transform self-attention mechanism

The invention discloses a multi-dialect speech recognition method based on a Transform self-attention mechanism, and belongs to the technical field of speech recognition, and the method comprises the steps: obtaining a multi-dialect audio data set, and carrying out the preprocessing and feature extraction; using a dialect topological space mapping module to construct topological space representation of dialect features; inputting the topological mapping features into a Transform encoder, and performing encoding processing by using a cross-dialect attention fusion mechanism and a dialect feature dynamic weighting mechanism; according to the method, Riemannian manifold mapping, graph attention fusion and gating dynamic weighting are introduced, the space structure relation of the speech features among dialects is accurately modeled, the recognition accuracy under a multi-dialect scene is remarkably improved, the word error rate is averagely reduced by 18%-25%, and the recognition accuracy of a few dialects is improved by 30%-40%.
Owner:YANGTZE UNIVERSITY