Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Speech classification" patented technology

Speech Disorders Classification System-Typology (SDCS-T) TheleftarmoftheSDCSshowninFigure1includesclassificationcategoriesforfourtypesof speech sound disorders based on a speaker’s age and current and/or prior speech errors. Normal(ized) Speech Acquisition (NSA) is assigned to speakers of any age with typical or normalized speech.

Method for training speech enhancement network, method for enhancing speech, and electronic device

PendingUS20250391419A1Speech analysisNoiseSpeech classification
A method for training a speech enhancement network, performed by an electronic device, includes: acquiring a first clean speech sample and a noise sample, and mixing them to generate a noisy speech sample; performing noise reduction on the noisy speech sample based on the speech enhancement network to obtain an enhanced speech sample; framing the enhanced speech sample into a plurality of enhanced speech frames, classifying speech effectiveness of the enhanced speech frames, and generating a first effectiveness distribution based on classification results of the enhanced speech frames; and determining a noise reduction accuracy based on the enhanced speech sample and the first clean speech sample, determining a speech classification accuracy based on the first effectiveness distribution, determining a speech enhancement accuracy based on the noise reduction accuracy and the speech classification accuracy, and training the speech enhancement network based on the speech enhancement accuracy.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Toxic speech detection method and system based on forgetting learning

The invention discloses a toxic speech detection method and system based on forgetting learning, and the method comprises the steps: firstly constructing a toxic speech classification model, dynamically tracking the classification state change of each sample in a training process, and carrying out the quantitative calculation of the total number of forgetting events; then sorting and analyzing the samples based on the forgetting frequency, and identifying and removing redundant samples which are difficult to forget so as to construct a high-quality simplified training set; and the model is retrained by using the simplified data set, the model parameters are optimized, and the detection efficiency is improved. The system comprises a data preprocessing module, a forgetting event calculation module, a sample screening module, a model training module and a detection generation module, and a complete toxicity speech detection and optimization process is formed. According to the method, a forgetting learning mechanism is introduced, so that the defects of training data redundancy, annotation noise interference and the like in a traditional method are effectively overcome, the model training speed and generalization performance are remarkably improved while the detection accuracy is ensured, and reliable technical support is provided for online content safety management and toxic speech real-time detection.
Owner:ZHEJIANG UNIV OF TECH

Adversarial training of keyword spotting to minimize TTS data overfitting

PendingUS20260051318A1Biological modelsSpeech recognitionSpeech trainingHidden layer
A method includes receiving training utterances that include non-synthetic speech training utterances and synthetic speech utterances. For each training utterance, the method includes processing, using a memorized neural network, a corresponding sequence of input audio frames to generate a hotword detection output indicating a likelihood the training utterance includes a hotword, determining a first loss based on the hotword detection output, obtaining a hidden layer feature vector for each corresponding input audio frame; processing, using a speech classification model, the hidden layer feature vectors to predict a classification output for the training utterance; and determining an adversarial loss based on the classification output predicted for the training utterance. The method also includes training the memorized neural network on the first losses and the adversarial losses to teach the memorized neural network to learn how to detect the hotword in audio and prevent overfitting of the synthetic speech training utterances.
Owner:GDM HOLDING LLC

A deep learning-based teaching quality evaluation method and system

The present application belongs to the technical field of intelligent teaching, and particularly relates to a teaching quality evaluation method and system based on deep learning. The method comprises extracting feature data to be evaluated from normal speech, evaluating the feature data to be evaluated by using a speech evaluation model to generate an evaluation result, and determining the quality grade of the normal speech according to the evaluation result, wherein the feature extraction from noise speech and noise speech comprises extracting amplitude information and frequency information from a sound production section. The present application performs screening on teaching speech to identify abnormal sound sections; through voiceprint comparison, the teaching speech in which the abnormal sound sections that can match the pre-stored voiceprint are classified as noise speech, and the teaching speech that cannot be matched is classified as noise speech, so as to distinguish the noise speech originating from the background environment from the noise speech originating from the teaching subject, overcome the evaluation error problem caused by regarding the two as noise without distinction, and lay a data foundation for subsequent evaluation.
Owner:CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP

Efficient speech detection method and device based on feature enhancement pre-trained model

ActiveCN119132337BInternal combustion piston enginesSpeech recognitionNoiseSpeech classification
This application relates to an effective speech detection method and apparatus based on a feature-enhanced pre-trained model. The method includes: acquiring speech to be detected containing different types of noise; inputting the speech to be detected into a first pre-trained model, and extracting effective speech features from the speech to be detected through the first pre-trained model; the first training data used by the first pre-trained model is obtained by enhancing the data features of unlabeled sample speech; inputting the effective speech features into a second pre-trained model, and performing effective speech classification through the second pre-trained model to obtain a classification result sequence; and outputting effective speech segments of the speech to be detected based on the classification result sequence; the effective speech segments are speech segments from which noise has been removed. This method can adapt to more application scenarios and noise types, effectively improving the effective speech detection effect and performance, thereby enhancing the performance of the speech recognition system.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Voice quality detection method and device, medium and equipment

The embodiment of the invention discloses a voice quality detection method, and the method comprises the steps: determining a voice classification result of each frame through a preset voice activity detection algorithm, dividing audio data into a voice segment and a non-voice segment, taking each frame of audio of which the classification result is the voice data in the non-voice segment as an interference frame, and carrying out the elimination. And calculating a signal-to-noise ratio of the voice segment to determine a quality detection result. And the interference frame which is misjudged as voice is eliminated from the non-voice segment, and purer noise estimation is obtained. Based on the calculated signal-to-noise ratio, the voice signal quality can be reflected more truly, and the problems of noise power overestimation and signal-to-noise ratio underestimation caused by VAD false detection are effectively solved, so that the accuracy and reliability of voice quality detection are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Method and apparatus for voice-controlled air conditioner, air conditioner, storage medium

ActiveCN114822529BMechanical apparatusSpeech recognitionHome applianceSpeech classification
This application relates to the field of smart home appliance technology, and discloses a method for voice-controlled air conditioners. The method includes, when a voice module receives an externally input voice command, sending the voice command to a voice classification module to determine whether the voice command is an offline voice command; if the voice command is an offline voice command, feeding the voice command back to the voice module to control the air conditioner's operation; if the voice command is not an offline voice command, transmitting the voice command to the cloud for parsing and recognition before feeding it back to the voice module to control the air conditioner's operation. By first determining whether the user's voice command is an offline voice command, and then performing the corresponding offline or online control, this avoids using online cloud AI to parse and recognize the voice command at the initial stage, reducing the frequency of using online AI for voice-controlled air conditioners and saving system resources. This application also discloses a device, an air conditioner, and a storage medium for voice-controlled air conditioners.
Owner:QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD

Method and system for robust processing of speech classifiers

PendingJP2026505579ASpeech analysisSpeech classificationSpeech sound
The present disclosure relates to a method and system for performing speech classification on an audio signal. The method includes obtaining the audio signal including a sequence of audio frames and determining, for each audio frame, a first speech confidence metric using a first speech classifier. For each given audio frame in at least a subset of the sequence of audio frames, the method includes classifying each respective audio frame in a first context window associated with the given audio frame as a speech frame or a non-speech frame by comparing the first speech confidence metric of the given audio frame to a first predetermined threshold, determining an adaptive threshold based on a number of speech frames in the first context window, and determining a first binary speech classification indicator for the given audio frame based on the first speech confidence metric and the adaptive threshold.
Owner:DOLBY LABORATORIES LICENSING CORP

Method for determining classification of a vehicle domain or an external domain based on user speech and a speech recognition system for a vehicle

ActiveUS12494193B2Speech recognitionElectric/fluid circuitSpeech classificationA domain
A method for determining a vehicle domain includes: converting a user's speech into text; and classifying the user's speech into a vehicle domain or an external domain based on the text, wherein the classifying of the user's speech into the vehicle domain or the external domain includes classifying a domain of the user's speech based on previously stored keyword-related information and then classifying the domain of the user's speech based on previously stored keyword-related information and then classifying the domain of the user's speech based on a trained domain classification model.
Owner:HYUNDAI MOTOR CO LTD +1

Digital semantic communication method and system for voice reconstruction and classification

The invention relates to a digital semantic communication method and system for voice reconstruction and classification, and the method comprises the steps: carrying out the preprocessing of voice data, and inputting the preprocessed data into a semantic encoder formed based on a double-attention residual mechanism network; after semantic features output by the semantic encoder are processed by the channel encoder, the semantic features are mapped into digital constellation symbols by a modulator adopting a soft quantization method; after transmission through the physical channel, the modulation symbols recover features by the soft decision demodulator, and signal distortion caused by the physical channel is eliminated through the channel decoder; and the demodulated semantic features are respectively input into a semantic decoder with a symmetric architecture to realize voice reconstruction, and input into a voice classifier consisting of a lightweight feature compression network and a classification head to realize voice classification. According to the method, the feature extraction integrity is improved through a double-attention architecture, the gradient fracture problem is solved through soft quantization and soft decision, and the channel robustness is enhanced in combination with progressive signal-to-noise ratio training.
Owner:GUANGZHOU UNIVERSITY

Automatically determining timing windows for speech captions in an audio stream

The invention relates to automatically determining timing windows for speech captions in an audio stream. A content system inputs segments of an audio stream to a speech classifier for classification, the speech classifier generating raw scores for the segments of the audio stream that represent a likelihood that the respective segment of the audio stream includes the presence of speech sounds. The content system generates binary scores for the audio stream based on the set of raw scores, each binary score generated based on an aggregation of the raw scores from a contiguous series of segments of the audio stream. The content system generates one or more timing windows for speech sounds in the audio stream based on the binary scores, each timing window indicating an estimate of a start and end timestamp for one or more speech sounds in the audio stream.
Owner:GOOGLE LLC

Multi-modal classification method and system for audio semantic understanding and rejection

PendingCN121747545ASpeech recognitionClassification methodsSpeech classification
The embodiment of the invention provides a multi-modal classification method and system for audio semantic understanding and rejection. The method comprises the following steps: receiving a collected original audio signal; acquiring a context historical memory of the original audio signal in voice interaction, and capturing multi-dimensional features related to semantic understanding and rejection from the original audio signal and the context historical memory; inputting the multi-dimensional features into a pre-trained multi-modal classification large model, wherein the multi-modal classification large model directly determines speech recognition content, semantic classification and rejection results of the original audio signals by using the multi-dimensional features; and determining a reply result for feeding back the original audio signal based on the speech recognition content, the semantic classification and the rejection result. According to the embodiment of the invention, end-to-end processing from audio to semantic understanding and rejection is realized. The speech recognition capability and the speech classification and rejection capability are unified in the multi-modal large model, so that the calculation resources are reduced, the overall system delay is greatly reduced, and the speech recognition accuracy is also improved.
Owner:AISPEECH CO LTD

Voice classification model optimization training method and device, computer equipment and medium

ActiveCN115579019BTime domainEngineering
The application relates to the technical field of artificial intelligence, in particular to a speech classification model optimization training method and device, computer equipment and a medium. The method obtains a speech sample and a sampling frequency, determines a frequency range according to the sampling frequency, constructs a filter function according to first and second cutoff frequencies determined in the frequency range, obtains a time domain function through inverse Fourier transform, inputs a classification model after convolution calculation of the speech sample and the time domain function, obtains an output result, trains the classification model and the time domain function according to the output result and the category of the speech sample, constructs the filter function with the first and second cutoff frequencies, and then performs convolution with the time domain function corresponding to the filter function, so that the trained time domain convolution function can more easily capture information of a target frequency band, and the trained classification model can be provided with more accurate speech features, thereby improving the accuracy of speech classification model classification.
Owner:PING AN TECH (SHENZHEN) CO LTD

Enhanced logits for natural language processing

ActiveUS12511492B2Natural language data processingTransmissionLogitSpeech classification
Techniques for using enhanced logit values for classifying utterances and messages input to chatbot systems in natural language processing. A method can include a chatbot system receiving an utterance generated by a user interacting with the chatbot system and inputting the utterance into a machine-learning model including a series of network layers. A final network layer of the series of network layers can include a logit function. The machine-learning model can map a first probability for a resolvable class to a first logit value using the logit function. The machine-learning model can map a second probability for a unresolvable class to an enhanced logit value. The method can also include the chatbot system classifying the utterance as the resolvable class or the unresolvable class based on the first logit value and the enhanced logit value.
Owner:ORACLE INT CORP

VEHICLE-BASED VOICE UNIT WITH HYBRID VOICE RECOGNITION

A hybrid speech recognition (HLI) system comprises one or more microphones configured to detect an acoustic utterance within the interior of a host system, a loudspeaker configured to transmit a request or response within the host system's interior, a processor, and memory. The processor executes a procedure that uses hybrid speech recognition logic stored in memory to classify the acoustic utterance as a hybrid utterance containing two or more languages, determines the relative language contribution of each language, and instructs the loudspeaker to transmit the request or response within the interior. The request or response contains the relative language contribution. The HLI system can be used as part of a vehicle, which has a body structure that defines a vehicle interior, with the hybrid speech recognition taking place within the vehicle interior.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Voice classification processing method and system for doctor-patient dialogue analysis

The invention discloses a voice classification processing method and system for doctor-patient dialogue analysis. The method comprises the following steps: acquiring voice data of a doctor user; based on the sound data, extracting voiceprint features of the doctor user; dialogue voice data including a dialogue between the doctor user and the patient user is acquired; inputting the voiceprint features and the dialogue voice data into a trained voice recognition segmentation model to obtain voice segmentation result data; the voice segmentation result data comprises a plurality of voice text data which are respectively marked as the doctor user and the patient user. Visibly, real-time doctor-patient dialogue structured processing based on voiceprint and deep learning can be realized, the accuracy and availability of medical dialogue records are improved, and the risk of medical record errors caused by confusion of speakers in traditional transcription is reduced.
Owner:GUANGZHOU UNIVERSITY OF CHINESE MEDICINE

Speech recognition method, apparatus, device, and storage medium

ActiveCN115512695BSpeech recognitionSpeech codeSpeech classification
The application discloses a speech recognition method, device and equipment and a storage medium. A speech recognition model configured by the application predicts an initial predicted text based on speech coding features output by a speech encoder through a first speech classification layer. A text encoder encodes the initial predicted text, fuses the text coding features and the speech coding features, inputs the fused coding features into a shared encoder for secondary coding, and obtains a final predicted text based on the secondary coding features by a second speech classification layer. Since the speech recognition model can extract more rich fused coding features as a whole, the recognition accuracy can be further improved. In addition, since the speech recognition model includes a text encoder and a shared encoder, the text encoder and the shared encoder can be additionally trained using pure text data in the training process. Pure text data is easier to obtain in large quantities than annotated text of speech, greatly reducing the cost of manual annotation.
Owner:IFLYTEK CO LTD

Intelligent outbound method and device, computer equipment and storage medium

The invention relates to the field of voice processing, and discloses an intelligent outbound method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a customer telephone voice stream when an intelligent outbound task is executed, and carrying out the slicing of the customer telephone voice stream, and obtaining a plurality of customer voice segments; calculating the voice trueness of the customer voice segment and further calculating the round trueness of the interaction round; processing the voice segments in the interaction round through an AI voice classification model to obtain an AI voice probability; and when the turn truth is smaller than a preset truth threshold value and the AI voice probability is greater than a preset AI voice threshold value, executing a preset on-hook process. The AI answering scene recognition accuracy and processing efficiency of the intelligent outbound call system are improved, call resources and operation cost are saved, and the method can be applied to financial science and technology and medical health old-age service scenes.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method for training a speech classification model, speech classification method and apparatus

ActiveCN117672196BSpeech classificationAcoustics
The application provides a speech classification model training method, a speech classification method and device, comprising: a plurality of sample speech segments obtained by segmenting and determining a speech stream with a plurality of speakers; inputting each sample speech segment into a pre-constructed speech classification model to generate a target feature vector, a first prediction result and a second prediction result; clustering each sample speech segment according to the target feature vector to obtain a plurality of clustering clusters, and sample speech segments with the same pseudo label belong to speech segments corresponding to the same speaker; calculating a first error of the pseudo label and the first prediction result; calculating a second error of the second prediction result and label information; and training the speech classification model according to the first error and the second error. According to the embodiment of the application, the accuracy of classification can be improved, so that the speakers corresponding to the plurality of speech segments in the speech stream can be clustered without waiting for the entire conversation to end.
Owner:CHINA SOUTHERN POWER GRID BIG DATA SERVICE CO LTD

Method and apparatus for classifying speech information

The application provides a voice information classification method and device, relates to the technical field of artificial intelligence, and can be applied to the technical field of finance or other technical fields.The voice information classification method comprises the following steps: obtaining voice information to be recognized; extracting voice feature information from the voice information; inputting the voice feature information into a voice classification model created based on business voice data and general voice data to obtain a voice classification result.The application can increase data credibility and realize accurate recognition in a specific business scenario.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Seamless roaming method and system

The application relates to the communication technical field, and provides a seamless roaming method and system. The method comprises the following steps: receiving a first frame sent by a station; determining a target access point based on the first frame; sending context information of the station to the target access point, and receiving response information sent by the target access point; setting a value of an access category (AC) to which a second frame belongs as a value of a voice category (AC_VO), or setting a value of an AC to which the second frame belongs as a new value; and sending the second frame to the station. In the process of switching the access point, the value of the AC to which the second frame belongs is set as the value of the voice category (AC_VO), or the value of the AC to which the second frame belongs is set as the new value, so that, compared with other to-be-executed processes of a source access point, the sending process of the second frame can be preferentially executed by the source access point, thereby the interaction time length of the overall switching process can be reduced, the packet loss rate can be reduced, the service lag can be avoided, and the user experience can be improved.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Speech recognition methods, speech recognition systems, computer equipment and storage media

ActiveCN116959424Bavoid collectingaccurate identificationSpeech recognitionSpeech codeSpeech classification
This application provides a speech recognition method, a speech recognition system, a computer device, and a storage medium, belonging to the field of financial technology. The method includes: inputting target speech with a preset emotion category into a pre-trained multi-task speech recognition model; encoding the target speech using a first speech coding sub-model to obtain initial speech features; performing speech attention processing on the initial speech features using a first attention sub-model to obtain first target attention features; encoding the initial speech features using a second speech coding sub-model to obtain hidden speech features; performing hidden attention processing on the first target attention features and the hidden speech features using a second attention sub-model to obtain second target attention features; and performing speech classification on the second target attention features using a multi-task classification sub-model to obtain a target speech label. This application embodiment can improve the recognition accuracy of multi-task speech recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD

A speech recognition method and system for use in driving test (Part 3)

This invention discloses a speech recognition method and system applied to the driving test (Part 3). Based on video and audio data, this invention filters out suspected cheating audio segments. Then, through a constructed speech recognition model, the speech data is processed as a simple image classification. While maintaining accuracy, the Squeeze-and-Excitatio module in MobileNetV3, which has a poor effect on speech classification, is removed from the speech recognition model's network structure to improve speed, balancing accuracy and speed. The ReLU activation function is used in the front part of the network, and the Hardwish function is used in the back part. Furthermore, the Focal Loss function is used to mitigate the imbalance problem in the speech data. This invention also addresses the problem in existing technologies that cannot monitor voice cheating behavior emitted by the safety officer in the driving test vehicle.
Owner:DUOLUN TECH CO LTD

Lightweight voice detection model training method, apparatus and device, and storage medium

PendingCN121708904ABiological modelsSpeech recognitionEngineeringSpeech classification
The invention relates to the technical field of lightweight voice model deployment, in particular to a lightweight voice detection model training method and device, equipment and a storage medium. Comprising the steps of initializing a student model, and obtaining voice input data; generating a corresponding first phoneme probability sequence through a preset teacher model, and performing voice annotation on the voice input data based on the first phoneme probability sequence to obtain voice annotation data; generating a second phoneme probability sequence and a voice classification probability sequence through the student model; calculating voice classification loss according to the voice annotation data and the voice classification probability sequence, constructing phoneme classification loss according to the first phoneme probability sequence and the second phoneme probability sequence, and fusing the phoneme classification loss into joint training loss; and parameter updating is performed on the student model based on the joint training loss, and lightweight processing is performed on the student model until a first iteration condition is satisfied, so that a lightweight voice detection model can be obtained. According to the invention, the balance between high precision and low resource consumption of the lightweight voice detection model can be realized.
Owner:SHENZHEN RAISOUND TECH