Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Speech classification" patented technology

Speech Disorders Classification System-Typology (SDCS-T) TheleftarmoftheSDCSshowninFigure1includesclassificationcategoriesforfourtypesof speech sound disorders based on a speaker’s age and current and/or prior speech errors. Normal(ized) Speech Acquisition (NSA) is assigned to speakers of any age with typical or normalized speech.

Method for training speech enhancement network, method for enhancing speech, and electronic device

PendingUS20250391419A1Speech analysisNoiseSpeech classification
A method for training a speech enhancement network, performed by an electronic device, includes: acquiring a first clean speech sample and a noise sample, and mixing them to generate a noisy speech sample; performing noise reduction on the noisy speech sample based on the speech enhancement network to obtain an enhanced speech sample; framing the enhanced speech sample into a plurality of enhanced speech frames, classifying speech effectiveness of the enhanced speech frames, and generating a first effectiveness distribution based on classification results of the enhanced speech frames; and determining a noise reduction accuracy based on the enhanced speech sample and the first clean speech sample, determining a speech classification accuracy based on the first effectiveness distribution, determining a speech enhancement accuracy based on the noise reduction accuracy and the speech classification accuracy, and training the speech enhancement network based on the speech enhancement accuracy.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Toxic speech detection method and system based on forgetting learning

The invention discloses a toxic speech detection method and system based on forgetting learning, and the method comprises the steps: firstly constructing a toxic speech classification model, dynamically tracking the classification state change of each sample in a training process, and carrying out the quantitative calculation of the total number of forgetting events; then sorting and analyzing the samples based on the forgetting frequency, and identifying and removing redundant samples which are difficult to forget so as to construct a high-quality simplified training set; and the model is retrained by using the simplified data set, the model parameters are optimized, and the detection efficiency is improved. The system comprises a data preprocessing module, a forgetting event calculation module, a sample screening module, a model training module and a detection generation module, and a complete toxicity speech detection and optimization process is formed. According to the method, a forgetting learning mechanism is introduced, so that the defects of training data redundancy, annotation noise interference and the like in a traditional method are effectively overcome, the model training speed and generalization performance are remarkably improved while the detection accuracy is ensured, and reliable technical support is provided for online content safety management and toxic speech real-time detection.
Owner:ZHEJIANG UNIV OF TECH

Adversarial training of keyword spotting to minimize TTS data overfitting

A method includes receiving training utterances that include non-synthetic speech training utterances and synthetic speech utterances. For each training utterance, the method includes processing, using a memorized neural network, a corresponding sequence of input audio frames to generate a hotword detection output indicating a likelihood the training utterance includes a hotword, determining a first loss based on the hotword detection output, obtaining a hidden layer feature vector for each corresponding input audio frame; processing, using a speech classification model, the hidden layer feature vectors to predict a classification output for the training utterance; and determining an adversarial loss based on the classification output predicted for the training utterance. The method also includes training the memorized neural network on the first losses and the adversarial losses to teach the memorized neural network to learn how to detect the hotword in audio and prevent overfitting of the synthetic speech training utterances.
Owner:GDM HOLDING LLC

A deep learning-based teaching quality evaluation method and system

The present application belongs to the technical field of intelligent teaching, and particularly relates to a teaching quality evaluation method and system based on deep learning. The method comprises extracting feature data to be evaluated from normal speech, evaluating the feature data to be evaluated by using a speech evaluation model to generate an evaluation result, and determining the quality grade of the normal speech according to the evaluation result, wherein the feature extraction from noise speech and noise speech comprises extracting amplitude information and frequency information from a sound production section. The present application performs screening on teaching speech to identify abnormal sound sections; through voiceprint comparison, the teaching speech in which the abnormal sound sections that can match the pre-stored voiceprint are classified as noise speech, and the teaching speech that cannot be matched is classified as noise speech, so as to distinguish the noise speech originating from the background environment from the noise speech originating from the teaching subject, overcome the evaluation error problem caused by regarding the two as noise without distinction, and lay a data foundation for subsequent evaluation.
Owner:CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP

Voice classification method and device, electronic equipment and storage medium

PendingCN120687862ASemantic analysisOther databases indexingFeature extractionSpeech classification
The embodiment of the invention provides a voice classification method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring target voice data and a target text of the target voice data; performing feature extraction on the target voice data to obtain emotion feature information; performing feature extraction on the semantics of the target text to obtain intermediate feature information, and converting the intermediate feature information into graph structure feature information; and classifying the target voice data according to the emotion feature information and the graph structure feature information to obtain a classification result of the target voice data. Therefore, the accuracy of voice classification can be improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Efficient speech detection method and device based on feature enhancement pre-trained model

ActiveCN119132337BInternal combustion piston enginesSpeech recognitionNoiseSpeech classification
This application relates to an effective speech detection method and apparatus based on a feature-enhanced pre-trained model. The method includes: acquiring speech to be detected containing different types of noise; inputting the speech to be detected into a first pre-trained model, and extracting effective speech features from the speech to be detected through the first pre-trained model; the first training data used by the first pre-trained model is obtained by enhancing the data features of unlabeled sample speech; inputting the effective speech features into a second pre-trained model, and performing effective speech classification through the second pre-trained model to obtain a classification result sequence; and outputting effective speech segments of the speech to be detected based on the classification result sequence; the effective speech segments are speech segments from which noise has been removed. This method can adapt to more application scenarios and noise types, effectively improving the effective speech detection effect and performance, thereby enhancing the performance of the speech recognition system.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Voice quality detection method and device, medium and equipment

The embodiment of the invention discloses a voice quality detection method, and the method comprises the steps: determining a voice classification result of each frame through a preset voice activity detection algorithm, dividing audio data into a voice segment and a non-voice segment, taking each frame of audio of which the classification result is the voice data in the non-voice segment as an interference frame, and carrying out the elimination. And calculating a signal-to-noise ratio of the voice segment to determine a quality detection result. And the interference frame which is misjudged as voice is eliminated from the non-voice segment, and purer noise estimation is obtained. Based on the calculated signal-to-noise ratio, the voice signal quality can be reflected more truly, and the problems of noise power overestimation and signal-to-noise ratio underestimation caused by VAD false detection are effectively solved, so that the accuracy and reliability of voice quality detection are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Method and apparatus for voice-controlled air conditioner, air conditioner, storage medium

ActiveCN114822529BMechanical apparatusSpeech recognitionHome applianceSpeech classification
This application relates to the field of smart home appliance technology, and discloses a method for voice-controlled air conditioners. The method includes, when a voice module receives an externally input voice command, sending the voice command to a voice classification module to determine whether the voice command is an offline voice command; if the voice command is an offline voice command, feeding the voice command back to the voice module to control the air conditioner's operation; if the voice command is not an offline voice command, transmitting the voice command to the cloud for parsing and recognition before feeding it back to the voice module to control the air conditioner's operation. By first determining whether the user's voice command is an offline voice command, and then performing the corresponding offline or online control, this avoids using online cloud AI to parse and recognize the voice command at the initial stage, reducing the frequency of using online AI for voice-controlled air conditioners and saving system resources. This application also discloses a device, an air conditioner, and a storage medium for voice-controlled air conditioners.
Owner:QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD

Voice assistant error detection system

ActiveUS12431129B2Speech recognitionEngineeringSpeech classification
A vehicle system for classifying spoken utterance within a vehicle cabin as one of system-directed and non-system directed, the system may include at least one microphone configured to detect at least one acoustic utterance from at least one occupant of a vehicle, at least one sensor to detect user behavior data indicative of user behavior, and a processor programmed to: receive the acoustic utterance, classify the acoustic utterance as one of a system-directed utterance and a non-system directed utterance, determine whether the acoustic utterance was properly classified based on user behavior observed via data received from the sensor after the classification, and apply a mitigating adjustment to classifications of subsequent acoustic utterances based on an improper classification.
Owner:CERENCE OPERATING CO

Method and system for robust processing of speech classifiers

PendingJP2026505579ASpeech analysisSpeech classificationSpeech sound
The present disclosure relates to a method and system for performing speech classification on an audio signal. The method includes obtaining the audio signal including a sequence of audio frames and determining, for each audio frame, a first speech confidence metric using a first speech classifier. For each given audio frame in at least a subset of the sequence of audio frames, the method includes classifying each respective audio frame in a first context window associated with the given audio frame as a speech frame or a non-speech frame by comparing the first speech confidence metric of the given audio frame to a first predetermined threshold, determining an adaptive threshold based on a number of speech frames in the first context window, and determining a first binary speech classification indicator for the given audio frame based on the first speech confidence metric and the adaptive threshold.
Owner:DOLBY LABORATORIES LICENSING CORP

Method for determining classification of a vehicle domain or an external domain based on user speech and a speech recognition system for a vehicle

ActiveUS12494193B2Speech recognitionElectric/fluid circuitSpeech classificationA domain
A method for determining a vehicle domain includes: converting a user's speech into text; and classifying the user's speech into a vehicle domain or an external domain based on the text, wherein the classifying of the user's speech into the vehicle domain or the external domain includes classifying a domain of the user's speech based on previously stored keyword-related information and then classifying the domain of the user's speech based on previously stored keyword-related information and then classifying the domain of the user's speech based on a trained domain classification model.
Owner:HYUNDAI MOTOR CO LTD +1

False voice detection method, false voice detection model acquisition method and related equipment

ActiveCN116403603BSustainable transportationSpeech analysisSpeech classificationSpeech sound
The present invention provides a false voice detection method, a false voice detection model acquisition method, and related equipment. The false voice detection method comprises: acquiring a target speech; and detecting whether the target speech is false voice based on a pre-acquired target false voice detection model. The target false voice detection model is trained using training speech labeled with speech categories to construct a false voice detection model. The constructed false voice detection model includes a speech encoder, a speaker characterization module that acquires a speaker characterization based on the output of the speech encoder, a false voice characterization module that acquires a false voice characterization based on the output of the speech encoder, and a speech classification module that performs speech classification based on the outputs of the speaker characterization module and the false voice characterization module. The speaker characterization module is acquired by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a pre-trained speech model obtained through pre-training. The false voice detection method provided by the present invention can accurately detect whether a speech is false voice.
Owner:IFLYTEK CO LTD

Digital semantic communication method and system for voice reconstruction and classification

The invention relates to a digital semantic communication method and system for voice reconstruction and classification, and the method comprises the steps: carrying out the preprocessing of voice data, and inputting the preprocessed data into a semantic encoder formed based on a double-attention residual mechanism network; after semantic features output by the semantic encoder are processed by the channel encoder, the semantic features are mapped into digital constellation symbols by a modulator adopting a soft quantization method; after transmission through the physical channel, the modulation symbols recover features by the soft decision demodulator, and signal distortion caused by the physical channel is eliminated through the channel decoder; and the demodulated semantic features are respectively input into a semantic decoder with a symmetric architecture to realize voice reconstruction, and input into a voice classifier consisting of a lightweight feature compression network and a classification head to realize voice classification. According to the method, the feature extraction integrity is improved through a double-attention architecture, the gradient fracture problem is solved through soft quantization and soft decision, and the channel robustness is enhanced in combination with progressive signal-to-noise ratio training.
Owner:GUANGZHOU UNIVERSITY

High-precision dynamic memristor model and application method thereof

PendingCN120808843ADigital storagePhysical realisationTerm memorySpeech classification
A high-precision dynamic memristor model and an application method thereof are characterized in that the memristor model is established based on a Ta2O5 memristor, and the Ta2O5 memristor comprises a bottom electrode, a functional layer and a top electrode which are sequentially attached from bottom to top; the bottom electrode is platinum, the functional layer is tantalum pentoxide, and the top electrode comprises lower tantalum and upper platinum. The memristor model has the effects that the memristor model can fit experimental data of a Ta2O5-based multi-layer device, and 2.9% of relative root-mean-square error is realized in resistance state conversion. In a voice classification task, the training accuracy rate of the method reaches 99.4%, the test accuracy rate reaches 91.6%, and the potential of neural morphology calculation of the method is proved. According to the work, the accuracy of memristor modeling is improved, and application in intelligent computing and memory systems is supported.
Owner:SOUTHWEST UNIV

Automatically determining timing windows for speech captions in an audio stream

The invention relates to automatically determining timing windows for speech captions in an audio stream. A content system inputs segments of an audio stream to a speech classifier for classification, the speech classifier generating raw scores for the segments of the audio stream that represent a likelihood that the respective segment of the audio stream includes the presence of speech sounds. The content system generates binary scores for the audio stream based on the set of raw scores, each binary score generated based on an aggregation of the raw scores from a contiguous series of segments of the audio stream. The content system generates one or more timing windows for speech sounds in the audio stream based on the binary scores, each timing window indicating an estimate of a start and end timestamp for one or more speech sounds in the audio stream.
Owner:GOOGLE LLC

Multi-modal classification method and system for audio semantic understanding and rejection

PendingCN121747545ASpeech recognitionClassification methodsSpeech classification
The embodiment of the invention provides a multi-modal classification method and system for audio semantic understanding and rejection. The method comprises the following steps: receiving a collected original audio signal; acquiring a context historical memory of the original audio signal in voice interaction, and capturing multi-dimensional features related to semantic understanding and rejection from the original audio signal and the context historical memory; inputting the multi-dimensional features into a pre-trained multi-modal classification large model, wherein the multi-modal classification large model directly determines speech recognition content, semantic classification and rejection results of the original audio signals by using the multi-dimensional features; and determining a reply result for feeding back the original audio signal based on the speech recognition content, the semantic classification and the rejection result. According to the embodiment of the invention, end-to-end processing from audio to semantic understanding and rejection is realized. The speech recognition capability and the speech classification and rejection capability are unified in the multi-modal large model, so that the calculation resources are reduced, the overall system delay is greatly reduced, and the speech recognition accuracy is also improved.
Owner:AISPEECH CO LTD

Deep fake audio detection and protection method based on large-scale pre-trained model Whisper

ActiveCN120126481BSpeech recognitionSpeech classificationSpeech sound
The present invention discloses a deep fake audio detection and protection method based on a large-scale pre-trained model Whisper, comprising the steps of step S1: inputting the audio to be detected into the pre-trained model Whisper, and using the transcribed text of the entire audio as a prompt during the fine-tuning process of each audio segment, so that the model Whisper can utilize global text information; step S2: pre-processing the audio data to adapt to the model Whisper during training and decoding; step S3: evaluating the performance of the model Whisper during training by designing a detection cross-entropy loss function. The deep fake audio detection and protection method based on the large-scale pre-trained model Whisper disclosed by the present invention realizes audio authenticity identification through transfer learning. Different from existing speech classification or recognition tasks, the complete transcribed text of the audio is embedded in the decoder as prompt information through an innovative fine-tuning strategy, which overcomes the problem of single reliance on traditional acoustic features.
Owner:YANGTZE DELTA REGION INST OF TSINGHUA UNIV ZHEJIANG

Korean language learning system and method for applying word spacing to japanese language

PCT designated stage expiredWO2025084470A3Data processing applicationsNatural language data processingPart of speechSpeech classification
The present invention relates to a Korean language learning system and method for applying word spacing to Japanese language, the system comprising: a database which collects and processes text data required for Korean and Japanese language learning; a learning sentence generation unit which generates grammar training sentences in Korean and Japanese, in which at least one element among non-spacing of words, marking of word spacing, word segmentation, part-of-speech classification, and color marking has been edited using the text data collected and processed by the database; a learned content management unit which collects and displays content learned by a learner on the basis of the grammar training sentences; a right / wrong determination unit which determines right / wrong answers on the content learned by the learner; and a learning step determination unit which determines the next learning step on the basis of the content learned by the learner.
Owner:WEKLEM INC

Voice classification model optimization training method and device, computer equipment and medium

ActiveCN115579019BTime domainEngineering
The application relates to the technical field of artificial intelligence, in particular to a speech classification model optimization training method and device, computer equipment and a medium. The method obtains a speech sample and a sampling frequency, determines a frequency range according to the sampling frequency, constructs a filter function according to first and second cutoff frequencies determined in the frequency range, obtains a time domain function through inverse Fourier transform, inputs a classification model after convolution calculation of the speech sample and the time domain function, obtains an output result, trains the classification model and the time domain function according to the output result and the category of the speech sample, constructs the filter function with the first and second cutoff frequencies, and then performs convolution with the time domain function corresponding to the filter function, so that the trained time domain convolution function can more easily capture information of a target frequency band, and the trained classification model can be provided with more accurate speech features, thereby improving the accuracy of speech classification model classification.
Owner:PING AN TECH (SHENZHEN) CO LTD

Enhanced logits for natural language processing

ActiveUS12511492B2Natural language data processingTransmissionLogitSpeech classification
Techniques for using enhanced logit values for classifying utterances and messages input to chatbot systems in natural language processing. A method can include a chatbot system receiving an utterance generated by a user interacting with the chatbot system and inputting the utterance into a machine-learning model including a series of network layers. A final network layer of the series of network layers can include a logit function. The machine-learning model can map a first probability for a resolvable class to a first logit value using the logit function. The machine-learning model can map a second probability for a unresolvable class to an enhanced logit value. The method can also include the chatbot system classifying the utterance as the resolvable class or the unresolvable class based on the first logit value and the enhanced logit value.
Owner:ORACLE INT CORP

Gnn-lstm method for chinese lip speech classification based on node multi-association graph information fusion

The application belongs to the technical field of lip language mouth shape analysis, and discloses a GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion. Lip key points are represented by three structures of adjacency graph, symmetric graph and upper and lower lip relationship graph. High-dimensional space-time features are extracted under the synergistic effect of graph convolutional neural network and long short-term memory network, the space-time global correlation between lip key points is effectively captured, and the mouth shape class is divided based on initial and final vowels. The multi-level collaborative relationship between initial and final vowels is considered, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, a mouth shape library is established by using the extracted high-dimensional space-time features, the lip shape and corresponding pinyin are more discriminatively mapped and induced, the influence of mouth shape similarity and visual ambiguity on model performance is reduced, the method can adapt to complex changes and many-to-one mapping phenomena in actual pronunciation processes, and effectively enhances the accuracy and robustness of subsequent mouth shape classification.
Owner:XIANGJIANG LAB

VEHICLE-BASED VOICE UNIT WITH HYBRID VOICE RECOGNITION

A hybrid speech recognition (HLI) system comprises one or more microphones configured to detect an acoustic utterance within the interior of a host system, a loudspeaker configured to transmit a request or response within the host system's interior, a processor, and memory. The processor executes a procedure that uses hybrid speech recognition logic stored in memory to classify the acoustic utterance as a hybrid utterance containing two or more languages, determines the relative language contribution of each language, and instructs the loudspeaker to transmit the request or response within the interior. The request or response contains the relative language contribution. The HLI system can be used as part of a vehicle, which has a body structure that defines a vehicle interior, with the hybrid speech recognition taking place within the vehicle interior.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Customer service voice quality management method and device

The invention discloses a customer service voice quality management method and device. The method comprises the following steps: acquiring a session voice of a target customer service staff in a target round session, and performing voice recognition on the session voice to obtain a target utterance text; analyzing the target utterance text by using an utterance classification model to obtain a target utterance classification result; determining a target personnel classification result of the target round of session according to the target utterance classification result, a historical utterance classification result corresponding to the target customer service personnel in the historical round of session and a historical personnel classification result; when the target utterance classification result and the target person classification result both meet service specifications, session voice is output; and when any classification result does not meet the service specification, performing intention recognition on the target utterance text, and outputting synthetic speech which is generated based on an intention recognition result and meets the service specification. The technical problem that improper utterances possibly exist in the process that customer service personnel serve users, and the enterprise image is affected is solved.
Owner:CHINA TELECOM CORP LTD

Voice task scheduling method, device, system, and storage medium

ActiveCN120091368BProgram initiation/switchingResource allocationEngineeringSpeech classification
The present application relates to the field of data processing technology, and discloses a method and device, system, and storage medium for scheduling voice tasks, which are applied to terminal air conditioners. The method includes: performing audio classification on received voice tasks to obtain voice classification tasks; wherein the language classification tasks include non-wake-up tasks; performing network classification on non-wake-up tasks to obtain offline interaction tasks and online interaction tasks; obtaining the network connection status and terminal load status of the terminal air conditioner; and according to the network connection status and terminal load status, unloading online interaction tasks or offline interaction tasks to the collaborative device end for task processing, or the terminal air conditioner performs task processing locally. The method performs flexible task scheduling on voice tasks, effectively meeting the requirements of low-latency voice wake-up and high-reliability voice execution.
Owner:QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD

Voice classification processing method and system for doctor-patient dialogue analysis

The invention discloses a voice classification processing method and system for doctor-patient dialogue analysis. The method comprises the following steps: acquiring voice data of a doctor user; based on the sound data, extracting voiceprint features of the doctor user; dialogue voice data including a dialogue between the doctor user and the patient user is acquired; inputting the voiceprint features and the dialogue voice data into a trained voice recognition segmentation model to obtain voice segmentation result data; the voice segmentation result data comprises a plurality of voice text data which are respectively marked as the doctor user and the patient user. Visibly, real-time doctor-patient dialogue structured processing based on voiceprint and deep learning can be realized, the accuracy and availability of medical dialogue records are improved, and the risk of medical record errors caused by confusion of speakers in traditional transcription is reduced.
Owner:GUANGZHOU UNIVERSITY OF CHINESE MEDICINE

Speech recognition method, apparatus, device, and storage medium

ActiveCN115512695BSpeech recognitionSpeech codeSpeech classification
The application discloses a speech recognition method, device and equipment and a storage medium. A speech recognition model configured by the application predicts an initial predicted text based on speech coding features output by a speech encoder through a first speech classification layer. A text encoder encodes the initial predicted text, fuses the text coding features and the speech coding features, inputs the fused coding features into a shared encoder for secondary coding, and obtains a final predicted text based on the secondary coding features by a second speech classification layer. Since the speech recognition model can extract more rich fused coding features as a whole, the recognition accuracy can be further improved. In addition, since the speech recognition model includes a text encoder and a shared encoder, the text encoder and the shared encoder can be additionally trained using pure text data in the training process. Pure text data is easier to obtain in large quantities than annotated text of speech, greatly reducing the cost of manual annotation.
Owner:IFLYTEK CO LTD

Intelligent outbound method and device, computer equipment and storage medium

The invention relates to the field of voice processing, and discloses an intelligent outbound method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a customer telephone voice stream when an intelligent outbound task is executed, and carrying out the slicing of the customer telephone voice stream, and obtaining a plurality of customer voice segments; calculating the voice trueness of the customer voice segment and further calculating the round trueness of the interaction round; processing the voice segments in the interaction round through an AI voice classification model to obtain an AI voice probability; and when the turn truth is smaller than a preset truth threshold value and the AI voice probability is greater than a preset AI voice threshold value, executing a preset on-hook process. The AI answering scene recognition accuracy and processing efficiency of the intelligent outbound call system are improved, call resources and operation cost are saved, and the method can be applied to financial science and technology and medical health old-age service scenes.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method for training a speech classification model, speech classification method and apparatus

ActiveCN117672196BSpeech classificationAcoustics
The application provides a speech classification model training method, a speech classification method and device, comprising: a plurality of sample speech segments obtained by segmenting and determining a speech stream with a plurality of speakers; inputting each sample speech segment into a pre-constructed speech classification model to generate a target feature vector, a first prediction result and a second prediction result; clustering each sample speech segment according to the target feature vector to obtain a plurality of clustering clusters, and sample speech segments with the same pseudo label belong to speech segments corresponding to the same speaker; calculating a first error of the pseudo label and the first prediction result; calculating a second error of the second prediction result and label information; and training the speech classification model according to the first error and the second error. According to the embodiment of the application, the accuracy of classification can be improved, so that the speakers corresponding to the plurality of speech segments in the speech stream can be clustered without waiting for the entire conversation to end.
Owner:CHINA SOUTHERN POWER GRID BIG DATA SERVICE CO LTD

Method and apparatus for classifying speech information

The application provides a voice information classification method and device, relates to the technical field of artificial intelligence, and can be applied to the technical field of finance or other technical fields.The voice information classification method comprises the following steps: obtaining voice information to be recognized; extracting voice feature information from the voice information; inputting the voice feature information into a voice classification model created based on business voice data and general voice data to obtain a voice classification result.The application can increase data credibility and realize accurate recognition in a specific business scenario.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Method and system for robust processing of speech classifiers

PendingCN120677526ASpeech recognitionSpeech classificationSpeech sound
The present disclosure relates to a method and system for performing speech classification on an audio signal. The method includes obtaining an audio signal comprising a sequence of audio frames, and determining, for each audio frame, a first speech confidence metric using a first speech classifier. For each given audio frame of at least a subset of the sequence of audio frames, the method comprises classifying each respective audio frame of a first context window associated with the given audio frame as a speech frame or a non-speech frame by comparing a first speech confidence metric for the respective audio frame with a first predetermined threshold, an adaptive threshold is determined based on a number of speech frames of the first context window, and a first binary speech classification indicator is determined for the given audio frame based on the first speech confidence metric and the adaptive threshold.
Owner:DOLBY LABORATORIES LICENSING CORP