Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12 results about "Speech recognition performance" patented technology

Speech recognition apparatus, method, and program

ActiveCN115482822BSpeech recognitionData expansionSpeech recognition performance
Embodiments of the present application relate to a speech recognition apparatus, a method, and a program. A speech recognition apparatus, a method, and a program capable of improving speech recognition performance are provided. A speech recognition apparatus according to an embodiment includes a data expansion unit, a sound score calculation unit, an adjustment unit, a sound score merging unit, a lattice generation unit, and a search unit. The data expansion unit generates a plurality of expanded speech data based on input speech data. The sound score calculation unit generates a plurality of sound scores based on each of the plurality of expanded speech data and a sound model. The adjustment unit generates a plurality of adjusted sound scores by performing resampling on the plurality of sound scores, respectively. The sound score merging unit generates a merged sound score by merging the plurality of adjusted sound scores. The lattice generation unit generates a merged lattice based on the merged sound score, a pronunciation dictionary, and a language model. The search unit searches for a speech recognition result having the highest likelihood from the merged lattice.
Owner:KK TOSHIBA

Teacher classroom speech recognition method and system based on hot word guidance, and readable storage medium

PendingCN121838769ASemantic analysisBiological modelsSpeech recognition performanceSpeech sound
The invention relates to the technical field of semantic recognition, in particular to a teacher classroom speech recognition method and system based on hot word guidance and a readable storage medium. According to the method, the hot word bank strongly related to the teaching scene is constructed, the hot words in the hot word bank are subjected to data processing to obtain the fusion features with prominent features, and the fusion features are input into the large language model for speech recognition, so that the model can improve the recognition precision by using the hot word information. The objective of the invention is to improve the classroom speech recognition performance of teachers.
Owner:SOUTHWEST FORESTRY UNIVERSITY

Speech recognition large model text aided training method, speech recognition method and device

The invention discloses a voice recognition large model text auxiliary training method and device, and a voice recognition method and device, and relates to the technical field of voice recognition, and the method comprises the steps: obtaining a training set which is a text data set of a target field or an audio text pair data set of a source field, and training a trainable parameter entity through the training set, obtaining a trained parameter entity, obtaining a pseudo-audio feature corresponding to each piece of text data in the text data set based on the trained parameter entity, and performing parameter training on a text decoder and a text embedding layer in the initial speech recognition large model at least according to the text data set and the pseudo-audio features, and obtaining a target speech recognition large model. According to the method, the pseudo audio features are added in the training stage of the text decoder and the text embedding layer, and the real audio embedding features of the target field are simulated through the pseudo audio features, so that the optimization deviation of the text decoder is reduced, and the voice recognition performance of the target voice recognition large model on the target field audio is improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Speech recognition method, speech recognition model, speech recognition device and storage medium

PendingCN121281499ASpeech recognitionSpeech recognition performanceSpeech sound
The invention provides a speech recognition method, a speech recognition model, a speech recognition device and a storage medium. The speech recognition method comprises the following steps: acquiring a speech waveform to be recognized; inputting the feature vector of the voice waveform into a routing network, so that the routing network screens out at least one target voice recognition network corresponding to the feature vector from a plurality of voice recognition networks; inputting the feature vector of the voice waveform into at least one target voice recognition network to obtain at least one voice recognition result; and determining a final recognition result of the voice waveform according to the at least one voice recognition result. According to the embodiment of the invention, the voice recognition effect can be improved under the condition of considering the reasoning speed.
Owner:镁佳(北京)科技有限公司

Automatic speech recognition method in complex environment

The invention relates to an automatic speech recognition method in a complex environment, which comprises the following steps of: firstly, carrying out availability labeling classification on acquired speech data, and firstly excluding unavailable speech data; and then searching a comprehensive optimal VAD model to filter voice data, filtering out a non-voice part, and leaving effective voice data. Through the basic speech recognition model, preliminary automatic recognition and speaker labeling are carried out, dialects, environmental noise and background music are further recognized, and according to professional terms possibly appearing in different occasions, the actual application scene of the recognition model is generalized. Finally, secondary correction is carried out manually for model optimization, in the model optimization, through speaker voiceprint recognition, the voice of a speaker can be positioned more accurately, and environment noise and background music are regarded as individual speakers for identity recognition, so that the recognition precision and performance of the training model are remarkably improved. The method has more accurate speech recognition performance, and has better generalization at the same time.
Owner:CHENGDU LINGSHU YICHEN HEALTH TECHNOLOGY CO LTD

Identification model training method and device, identification method and device, equipment and medium

The invention provides a recognition model training method and device, a recognition method and device, equipment and a medium, and the method comprises the steps: training a pre-trained speech recognition model according to a sample audio and a text label, obtaining a first optimization parameter of a rank decomposition increment parameter corresponding to an optimization model parameter of an adapter module in a pre-trained speech recognition model and a model parameter of a large language model; training a pre-trained speech recognition model according to the sample text and the analog speech signal to obtain a second optimization parameter of the rank decomposition increment parameter; performing hierarchical fusion on the first optimization parameter and the second optimization parameter to obtain a fused optimization parameter; and updating model parameters of a pre-trained speech recognition model according to the optimization model parameters and the fusion optimization parameters to obtain a target speech recognition model. According to the invention, model fine tuning is carried out through cross-modal text adaptation and LoRA model parameter fusion, the dependence on paired data is broken through, and the speech recognition performance is effectively improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

A speech recognition method and system based on adaptive audiovisual modal fusion

PendingCN122313971ANetwork modelSpeech recognition performance
This invention discloses a speech recognition method and system based on adaptive audiovisual modal fusion, relating to the field of multimodal signal processing technology. The method includes: acquiring audiovisual speech data to be recognized; preprocessing the audiovisual speech data to obtain standardized input data; inputting the standardized input data into a pre-established modal expert network model, outputting feature representations of multiple experts, including audio feature representations, visual feature representations, and joint feature representations; inputting the feature representations of multiple experts into a pre-established soft routing module, outputting confidence weights of multiple experts; performing a weighted summation of the feature representations of multiple experts based on their confidence weights, obtaining a hybrid feature representation; inputting the hybrid feature representation into a pre-trained shared classifier, outputting the final speech recognition result, thereby improving the overall speech recognition performance.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method and system for evaluating voice recognition and directional sound transmission performance of open wireless earphone

PendingCN122266395AElectrical apparatusSpeech analysisNoiseSpeech recognition performance
The application discloses an open wireless earphone voice recognition and directional sound transmission performance evaluation method and system, relates to the technical field of earphone equipment, and comprises the following steps: voice recognition test preparation is carried out through a convolutional neural network noise suppression and an adaptive audio feedback mechanism, and a dynamic voice signal library is constructed to collect earphone voice signals; a word error rate is calculated based on the collected earphone voice signals, voice recognition evaluation is carried out in combination with short-time objective voice intelligibility and syllable rate estimation, and sound field directional sound transmission testing is carried out in combination with a directional gain formula. The application accurately evaluates the similarity between enhanced voice and real voice through multi-band filtering, syllable rate analysis and short-time objective voice intelligibility scoring, quantifies the voice recognition effect in combination with the word error rate, establishes the relationship between the signal-to-noise ratio and the voice recognition performance, and provides a comprehensive and dynamic performance evaluation system.
Owner:GUANGZHOU VIKEN COMM TECH CO LTD +1

Cooking apparatus including digital controller door and method of improving speech recognition performance

PendingCN121053978ADomestic stoves or rangesDoors for stoves/rangesEngineeringSpeech recognition performance
The present invention relates to a cooking apparatus including a digital controller door and a method of improving voice recognition performance, the cooking apparatus including the digital controller door according to an embodiment of the present invention includes a function unit providing a cooking function, the digital controller door including a microphone section receiving a voice of a user, and if the microphone section receives a wake-up word, the function unit provides a voice recognition function. If so, the controller executes a command word recognition pattern to receive a command word during a preset time period. Therefore, the technology of accurately recognizing voice in the process of executing voice recognition by the cooking equipment can be realized.
Owner:LG ELECTRONICS INC

Cooking appliance with digital controller door and method for enhancing voice recognition performance

A cooking appliance can include a main portion including a cavity for receiving food, and one or more functional components, the main portion being configured to provide a cooking function, a display door coupled to the main portion, the display door including a display configured to provide a user interface, and one or more door-mounted components including a microphone, and a controller. Also, the controller is configured to in response to receiving a wake-up word, execute a command word recognition mode for receiving a command word for a predetermined amount of time, in which the command word recognition mode includes adjusting an operation level of at least component among of the one or more functional components of the main portion or the one or more door-mounted components of the display door.
Owner:LG ELECTRONICS INC

Speech recognition and model distillation method, related equipment and program product

The invention discloses a voice recognition and model distillation method, related equipment and a program product, and the method comprises the steps: setting at least one stage of teaching assistant model between a teacher model and a student model (the size is between the teacher model and the student model), and carrying out the step-by-step downward training through a knowledge distillation method, and training the student model by using the last level of teaching assistant model to obtain a trained student model. By gradually compressing the model, key knowledge of the teacher model can be better migrated, and a better distillation effect is obtained. In the multi-stage knowledge distillation process, voice recognition performance evaluation is regularly carried out on the currently trained model, and when it is determined that a performance evaluation result does not meet requirements, the number and size of the multi-stage teaching assisting model are dynamically adjusted, so that the knowledge distillation effect is further promoted, and the voice recognition effect of the finally trained student model is improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Speech recognition model training method, using method and related device

The invention discloses a speech recognition model training method, a use method and a related device, and relates to the technical field of speech processing, and the method comprises the steps: processing a training data set through employing an initial speech recognition model and a first decoder, obtaining a token-level decoding sequence generated by the autoregression of the first decoder, according to a token-level decoding sequence generated by a first decoder in the initial speech recognition model and a frame-level decoding sequence generated by a second decoder in the initial speech recognition model, aligning the frame-level decoding sequence with the token-level decoding sequence by using the target attention information to obtain a first alignment sequence corresponding to the token-level decoding sequence and a second alignment sequence corresponding to the frame-level decoding sequence, and determining decoding consistency loss according to the first alignment sequence and the second alignment sequence so as to train the initial speech recognition model. According to the invention, the token-level decoding sequence is used as a strong supervision signal to train the initial speech recognition model, so that the encoder can encode the output rich in the context, and the speech recognition performance is improved on the premise of maintaining the efficient reasoning capability of the speech recognition model.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD