Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Speech segmentation" patented technology

Speech segmentation is the process of identifying the boundaries between words, syllables, or phonemes in spoken natural languages. The term applies both to the mental processes used by humans, and to artificial processes of natural language processing.

Speech processing method and apparatus, device, and medium

PendingUS20250329334A1Speech analysisSpeech segmentationAcoustics
A speech processing method includes: obtaining overlapping speech data; obtaining reference speech data of a specified object; extracting a voiceprint representation vector of the specified object from the reference speech data, the voiceprint representation vector representing a voiceprint characteristic of the specified object, and inputting the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, and segmenting, by the speech segmentation model based on an attention mechanism, the overlapping speech data to obtain a target speech signal matching the voiceprint characteristic; and generating a speech file of the specified object based on the speech signal.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Speech translation using a wearable device

A speech translation system may provide real-time or near real-time translation of speech uttered by a person or emitted from a media device. The speech translation system may include a device that may receive audio representing speech in a source language and output audio representing speech in a target language. The speech translation system may translate the speech in portions representing semantically cohesive speech segments such that the target speech reflects the semantic meaning of words, phrases, and / or clauses as used in the context of the source speech. The speech translation system may condense the speech segments prior to or during translation to reduce verbosity. The speech translation system may selectively translate some speakers and not others, and may determine voice characteristics of source speech and apply identifying characteristics to the target speech that allow a user to differentiate respective target speech from different speakers based on the identifying characteristics.
Owner:AMAZON TECH INC

Doctor-patient speech communication model training method and system based on multi-modal corpus analysis

PendingCN121725770ASpeech recognitionPattern recognitionSpeech segmentation
The invention discloses a doctor-patient speech communication model training method and system based on multi-modal corpus analysis, and belongs to the crossing field of artificial intelligence and medical treatment, and the method comprises the steps: carrying out the extraction according to an OpenPose algorithm to obtain motion features, carrying out the muscle activity intensity detection to obtain expression features, carrying out the speech recognition to obtain text features, and carrying out the recognition of the text features; speech segmentation is carried out based on the time domain energy parameters, and a segmentation result is subjected to modal analysis to obtain acoustic features; performing feature fusion based on an attention mechanism to obtain fusion features, and inputting an obstacle recognition model to obtain an obstacle type; the method comprises the steps of obtaining an obstacle type, obtaining an intervention strategy and an intervention identity according to a solution mapping relation, taking the obstacle type, the intervention strategy and the intervention identity as multi-dimensional labels, carrying out time domain alignment and structured packaging to obtain target data, and training according to the target data to obtain a communication model used for providing dialogue prompt information. Obstacle recognition accuracy can be improved, and the intelligent agent application effect can be improved.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

System and method for enhanced customer service through automated real-time FAQ generation from call center interactions

An automated system for generating Frequently Asked Questions (FAQs) from call center interactions includes a call center device for recording audio conversations between agents and customers, and a backend system. The backend system segments the conversation between agent and customer speech, converts the segmented speech into text using an Automatic Speech Recognition engine, and generates FAQs using a Large Language Model. Each FAQ includes a query statement corresponding to the customer's speech and at least one answer statement corresponding to the agent's speech. The system also includes mechanisms for selecting relevant, non-duplicate FAQs, determining the importance of each FAQ based on frequency, sentiment, and coherence scores, and dynamically updating the FAQ database in real-time. A user interface displays the generated FAQs with dropdown arrows to view the answers, enhancing customer service efficiency and accuracy by providing immediate, relevant responses to common inquiries.
Owner:ELM INC

A method for generating a speaker diary based on audio-visual fusion clustering

ActiveCN119964596BSpeech recognitionSpeech segmentationSpeaker verification
The application discloses a speaker diary generation method based on audio-visual fusion clustering, and aims to solve the problem of "who speaks at what time" in a multi-speaker scene. The method is realized through the following steps: first, an overlapping-aware speech segmentation model is used to segment the audio segment, solving the problem of overlapping speech; second, an advanced speaker verification model is used to extract the speaker voiceprint features of each audio segment and a speaker score matrix generated through face tracking and speaker detection; then, through an audio-video joint clustering method, the number of clusters is optimized according to the audio features and visual information, and K-means clustering is used to complete speaker clustering; the experimental results show that the system adopting the method achieves the lowest diary error rate (DER) on the Ego4D validation set.
Owner:HUNAN UNIV

A real-time voice conversion method and device, electronic equipment and medium

ActiveCN115910083BSpeech recognitionSpeech synthesisSpeech segmentationEngineering
The application provides a real-time voice conversion method and device, electronic equipment and medium. The method comprises the following steps: intercepting first voice data meeting voice segmentation conditions from voice data of a source speaking object recorded in real time; processing the first voice data to extract first semantic information; inputting the first semantic information into a pre-trained voice conversion model, and converting and processing effective information of historical voice data before the first semantic information and the first voice data through the voice conversion model to obtain target voice feature information corresponding to the first semantic information and a voice factor of a target speaking object; reconstructing the target voice feature information to obtain second voice data converted from the first voice data, thereby realizing low-delay streaming inference and low-delay and high-performance real-time voice conversion.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

AI-based telephone answering system

The application discloses an AI-based telephone answering system, relates to the technical field of telephone answering, and comprises a real-time voice acquisition and preprocessing module, a speech speed detection and prediction module, a self-adaptive ASR dynamic regulation module, a refined voice slicing and decoding module, a keyword detection and voice recognition module, a semantic understanding and intention analysis module and an emergency response and dispatching module; the real-time voice acquisition and preprocessing module acquires user voice data of an emergency help telephone in real time, establishes a real-time audio stream transmission channel, and rapidly preprocesses the acquired audio signal. The application rapidly and accurately captures the voice features of a user under high speech speed, dynamically adjusts the voice segmentation length of an ASR engine, makes each voice segment more clear and accurate, avoids the splitting and recognition errors of cross-segment words caused by too fast speech speed, and effectively avoids the omission of user emergency information and the risk of misjudgment in an emergency help scene.
Owner:SHANDONG ZHIQUN INFORMATION TECH CO LTD

An audio separation and script violation reminding method, device and computer equipment

ActiveCN116631432BSpeech analysisSpeech segmentationFeature extraction
The specification relates to the technical field of artificial intelligence, in particular to an audio separation and script violation reminding method and device and a computer device. The audio separation method comprises the following steps: performing speech segmentation on a to-be-separated audio to obtain a plurality of audio segments; performing feature extraction on each audio segment to obtain an audio feature vector corresponding to the audio segment; processing the audio feature vector by using a trained speaker prediction model to determine a predicted speaker identifier corresponding to each audio segment; processing the audio feature vector corresponding to a target predicted speaker identifier by using a trained audio clustering model to obtain a target speaker identifier corresponding to each audio feature vector; and merging the audio segments corresponding to the same target speaker identifier to obtain a target audio corresponding to each target speaker identifier. By using the embodiment of the specification, the correction of the predicted speaker identifier is realized, and the accuracy of the audio separation is improved.
Owner:CHINA CITIC BANK CO LTD

System and method for enhanced customer service through automated real-time FAQ generation from call center interactions

An automated system for generating Frequently Asked Questions (FAQs) from call center interactions includes a call center device for recording audio conversations between agents and customers, and a backend system. The backend system segments the conversation between agent and customer speech, converts the segmented speech into text using an Automatic Speech Recognition engine, and generates FAQs using a Large Language Model. Each FAQ includes a query statement corresponding to the customer's speech and at least one answer statement corresponding to the agent's speech. The system also includes mechanisms for selecting relevant, non-duplicate FAQs, determining the importance of each FAQ based on frequency, sentiment, and coherence scores, and dynamically updating the FAQ database in real-time. A user interface displays the generated FAQs with dropdown arrows to view the answers, enhancing customer service efficiency and accuracy by providing immediate, relevant responses to common inquiries.
Owner:ELM INC

A method for regulatory speech segmentation based on speech recognition and end-point detection

ActiveCN117238279BAutomatic segmentationSpeech segmentation
The application provides a regulation voice segmentation method based on speech recognition and endpoint detection, which is applied to air traffic control voice audio stream segmentation, and comprises the following steps: step 1, constructing a punctuation model based on speech recognition and a speech endpoint detection model; step 2, using the punctuation model based on speech recognition to recognize the audio data stream of the regulation voice, and outputting the corresponding text and sentence end identifier of the audio data stream; step 3, using the speech endpoint detection model to judge the speech starting point and ending point contained in the audio data stream of the regulation voice; step 4, segmenting the audio data stream of the regulation voice into audio segments; and step 5, applying the audio segments as data materials to the speech recognition process of the air traffic control system. Through the combination of speech recognition and endpoint detection, the application realizes the automatic segmentation of the air traffic control voice audio stream, and improves the accuracy and efficiency of the segmentation.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

A small sample semantic segmentation method and system fusing class label semantics

ActiveCN118587440BCharacter and pattern recognitionSpeech segmentationImaging processing
The application relates to the field of image processing, and proposes a small sample semantic segmentation method and system fusing class label semantics, wherein a prior information generation module is designed, multi-modal data fusion of image data information and text data information serving as class labels is realized, a semantic segmentation model can more accurately understand image content, a multi-scale fusion module is further designed, original detail information of an image is further protected, the calculation performance and channel fusion capability of the semantic segmentation model are greatly improved, the parameter quantity of a decoder is greatly reduced while ensuring the speech segmentation precision, and the accuracy of target recognition and target positioning of semantic segmentation is greatly improved.
Owner:JIANGXI NORMAL UNIV

A conference management method and system based on natural language processing and retrieval enhancement

ActiveCN120634500BReservationsBiological modelsSpeech segmentationText stream
The application provides a conference management method and system based on natural language processing and retrieval enhancement, and belongs to the technical field of conference management. The method comprises the following steps: outputting a conference reservation result based on key parameters; if the reservation is successful, acquiring audio stream data in the conference process, performing speech segmentation, initial voiceprint extraction and initial clustering on the audio stream data to output original speech segments and speaker reference speech; performing target speaker extraction, accurate voiceprint extraction and final clustering based on the original speech segments and the speaker reference speech to obtain pure speech text stream with a voiceprint label; performing dynamic abstract generation, enhanced retrieval and verification on the pure speech text stream; performing responsibility recognition and voiceprint binding on the conference information to output final conference management data. The application can cover intelligent management of the whole conference process, and realize conference whole-process management from voice-driven reservation, real-time content analysis to automatic task distribution.
Owner:GUANGDONG KAMFU TECH CO LTD

Emotion recognition method in voice dialogue

PendingCN121191542ASpeech recognitionSpeech segmentationNoise
The invention relates to an emotion recognition method in voice dialogue, which comprises the following steps: S1, acquiring a voice stream, and converting the voice stream into a WAV format; s2, performing noise suppression on the voice stream in the WAV format; s3, performing voice segmentation processing on the voice stream after noise suppression; s4, carrying out standardization processing on the single-segment voice waveform after segmentation processing; s5, performing feature coding on the standardized voice waveform signal to obtain a feature vector of the input waveform; and S6, inputting the feature vector into a classifier, and outputting the most probable emotion in a given emotion category. According to the invention, automatic recognition of human emotional states is realized by analyzing acoustic features in voice signals, and the problem of poor robustness of an accompanying robot in a complex environment is solved; the dialogue of the accompanying robot is more natural and intelligent, the robot has the emotional sharing ability, and the robot is endowed with understanding and response to the emotion of human beings.
Owner:SHANGHAI HUAHUA INTELLIGENT TECH CO LTD

Binary natural speech segmentation method and system based on time domain analysis

PendingCN120998213ASpeech recognitionTime domainSpeech segmentation
The invention provides a binary natural speech segmentation method and system based on time domain analysis. The method comprises the following steps: acquiring a target binary natural voice file and generating corresponding binary natural voice data; performing code element-by-code analysis on the binary natural voice data, and identifying whether the current code element is an error code, a normal code or a cut-off position of a code word or a code block according to a relation between a time interval between the current code element and an adjacent code element and a preset time interval coefficient; and performing code word segmentation and code block segmentation on the binary natural voice data according to an identification result.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

A language data preprocessing method for multilingual and complex scenes

ActiveCN120913580BSpeech recognitionSpeech segmentationTransient noise
The application relates to the technical field of language data processing, and discloses a language data preprocessing method for multiple languages and complex scenes, a speech data preprocessing system for multiple languages and complex scenes based on an AutoPrep framework, and five modules of speech enhancement, speech segmentation, speaker clustering, target speech extraction and quality filtering are integrated to realize automatic and structured processing of speech data. The scheme realizes differentiated suppression of stable-state and transient-state noises in multiple language speech signals, especially in small language (such as Kazakh and Tagalog) scenes, effectively improves the speech signal-to-noise ratio and the language independence of speech characteristics, overcomes the problem that the existing technology lacks a special phonetic system processing module for small languages, resulting in a high phoneme mapping error rate, and enhances the availability and processing effect of low-resource language data.
Owner:SHENYI FUTURE TECHNOLOGY (GUANGDONG HENGQIN) CO LTD

Speaker voice segmentation method based on non-local space U-Net and mixed features

PendingCN121983030AImprove discrimination abilityFully explore the spatial and temporal correlationsSpeech recognitionSpeech segmentationSemantic feature
The invention belongs to the technical field of voice signal processing, and particularly provides a speaker voice segmentation method based on non-local space U-Net and mixed features. According to the invention, the non-local space U-Net network is constructed, and a non-local space attention module is introduced, so that the long-range dependency relationship in voice signals is effectively captured, and the expression ability of spatial features is improved; and meanwhile, a mixed feature fusion mechanism is adopted, and time-frequency domain features and deep semantic features are combined, so that the discrimination of voice features is enhanced. In addition, a dynamic weight loss function is designed, the fusion efficiency of the features under different scales is optimized, and the contribution degrees of various voice segments are balanced. According to the method, the space-time relevance in the voice signals can be fully mined, the dependence of a traditional method on local features and single features is overcome, and the accuracy and robustness of speaker voice segmentation are remarkably improved.
Owner:CHINA CRIMINAL POLICE UNIV

A Smart Classroom Interaction Analysis Method Based on Two-Layer Architecture Speech Segmentation

ActiveCN120783757BSpeech recognitionSpeech segmentationNerve network
This invention provides a smart classroom interaction analysis method based on a two-layer architecture speech segmentation, belonging to the field of speech segmentation technology. Specifically, it includes the following steps: extracting speech features from the speech signal using Mel-frequency cepstral coefficients (MFCC); designing a text-enhanced, multi-scale time-aware delay-based neural network to coarsely screen the speech features, dividing audio segments into single-speaker segments and multi-speaker segments; inputting the coarsely screened multi-speaker segments into a sliding window segmentation model (SW-NIF) that integrates neighbor window information to locate speaker transition points within the multi-speaker segments; training and validating the constructed model on a dataset. The technical solution of this invention overcomes the problem of neglecting classroom audio segmentation in existing technologies, which only perform simple segmentation of classroom audio for subsequent tasks, resulting in mixed speakers in the audio segments and affecting the analysis effect.
Owner:SHANDONG UNIV OF SCI & TECH

A multi-scene intelligent assistant system and interaction method based on dialect voice wake-up

The application relates to the technical field of intelligent voice interaction, in particular to a multi-scene intelligent assistant system based on dialect voice wake-up and an interaction method, which comprises a voice segmentation module, a scene screening module, a terminal access module, an intra-domain wake-up module and an entry determination module. The application establishes a continuous processing relationship by surrounding the voice segment boundary, the scene label residence state, the terminal release state and the dialect wake-up word structure, narrows the scene range of the current voice segment participating in identification according to the scene label residence state, checks the tone fluctuation, the syllable sequence and the tailing convergence in the limited range, judges the interaction entry attribution in combination with the terminal access state and the activation confirmation information, keeps the scene constraint and the dialect wake-up discrimination synchronous, keeps the terminal access and the entry confirmation connected, controls the cross-scene repeated matching, the near-sound trigger confusion and the terminal access conflict, and makes the dialect voice wake-up and the subsequent interaction more coherent.
Owner:XIAMEN LUJIANG TECHNOLOGY CO LTD

A method, apparatus, device, and readable storage medium for speech segmentation.

PendingCN122313977ASpeech segmentationConfusion
This application discloses a speech segmentation method, apparatus, device, and readable storage medium, relating to the field of speech interaction technology. It includes: performing real-time semantic analysis on the user's speech stream locally on the terminal device to determine the semantic integrity status, and dynamically selecting a target threshold from a preset set of silence thresholds to extract audio segments and send them to the cloud; the cloud performs endpoint detection on the audio segments, obtains the start and end timestamps of the speech, and calculates the time interval between the current segment and the previous segment; when the interval is less than or equal to a preset listening window, it performs a semantic correlation judgment on the two segments based on rules and a semantic model, and decides whether to splice them. This application solves the problems of response delay and semantic confusion caused by fixed thresholds through semantically driven dynamic threshold adjustment and cloud-based semantic-level splicing decision-making, thereby improving the response speed and recognition accuracy of speech interaction.
Owner:AISPEECH CO LTD

Speaker voice segmentation method based on non-local spatial U-Net and mixed features

ActiveCN121862087BSpeech segmentationSemantic feature
This invention belongs to the field of speech signal processing technology, specifically proposing a speaker speech segmentation method based on nonlocal spatial U-Net and hybrid features. This invention constructs a nonlocal spatial U-Net network, and by introducing a nonlocal spatial attention module, effectively captures long-range dependencies in the speech signal, enhancing the expressive power of spatial features. Simultaneously, a hybrid feature fusion mechanism is employed, combining time-frequency domain features and deep semantic features to enhance the discriminative power of speech features. Furthermore, a dynamic weight loss function is designed to optimize the fusion efficiency of features at different scales and balance the contribution of various speech segments. This invention can fully exploit the spatiotemporal correlations in speech signals, overcome the dependence of traditional methods on local and single features, and significantly improve the accuracy and robustness of speaker speech segmentation.
Owner:CHINA CRIMINAL POLICE UNIV