Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Speech segmentation" patented technology

Speech segmentation is the process of identifying the boundaries between words, syllables, or phonemes in spoken natural languages. The term applies both to the mental processes used by humans, and to artificial processes of natural language processing.

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Speech translation using a wearable device

A speech translation system may provide real-time or near real-time translation of speech uttered by a person or emitted from a media device. The speech translation system may include a device that may receive audio representing speech in a source language and output audio representing speech in a target language. The speech translation system may translate the speech in portions representing semantically cohesive speech segments such that the target speech reflects the semantic meaning of words, phrases, and / or clauses as used in the context of the source speech. The speech translation system may condense the speech segments prior to or during translation to reduce verbosity. The speech translation system may selectively translate some speakers and not others, and may determine voice characteristics of source speech and apply identifying characteristics to the target speech that allow a user to differentiate respective target speech from different speakers based on the identifying characteristics.
Owner:AMAZON TECH INC

Doctor-patient speech communication model training method and system based on multi-modal corpus analysis

PendingCN121725770ASpeech recognitionPattern recognitionSpeech segmentation
The invention discloses a doctor-patient speech communication model training method and system based on multi-modal corpus analysis, and belongs to the crossing field of artificial intelligence and medical treatment, and the method comprises the steps: carrying out the extraction according to an OpenPose algorithm to obtain motion features, carrying out the muscle activity intensity detection to obtain expression features, carrying out the speech recognition to obtain text features, and carrying out the recognition of the text features; speech segmentation is carried out based on the time domain energy parameters, and a segmentation result is subjected to modal analysis to obtain acoustic features; performing feature fusion based on an attention mechanism to obtain fusion features, and inputting an obstacle recognition model to obtain an obstacle type; the method comprises the steps of obtaining an obstacle type, obtaining an intervention strategy and an intervention identity according to a solution mapping relation, taking the obstacle type, the intervention strategy and the intervention identity as multi-dimensional labels, carrying out time domain alignment and structured packaging to obtain target data, and training according to the target data to obtain a communication model used for providing dialogue prompt information. Obstacle recognition accuracy can be improved, and the intelligent agent application effect can be improved.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

A real-time voice conversion method and device, electronic equipment and medium

ActiveCN115910083BSpeech recognitionSpeech synthesisSpeech segmentationEngineering
The application provides a real-time voice conversion method and device, electronic equipment and medium. The method comprises the following steps: intercepting first voice data meeting voice segmentation conditions from voice data of a source speaking object recorded in real time; processing the first voice data to extract first semantic information; inputting the first semantic information into a pre-trained voice conversion model, and converting and processing effective information of historical voice data before the first semantic information and the first voice data through the voice conversion model to obtain target voice feature information corresponding to the first semantic information and a voice factor of a target speaking object; reconstructing the target voice feature information to obtain second voice data converted from the first voice data, thereby realizing low-delay streaming inference and low-delay and high-performance real-time voice conversion.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

An audio separation and script violation reminding method, device and computer equipment

ActiveCN116631432BSpeech analysisSpeech segmentationFeature extraction
The specification relates to the technical field of artificial intelligence, in particular to an audio separation and script violation reminding method and device and a computer device. The audio separation method comprises the following steps: performing speech segmentation on a to-be-separated audio to obtain a plurality of audio segments; performing feature extraction on each audio segment to obtain an audio feature vector corresponding to the audio segment; processing the audio feature vector by using a trained speaker prediction model to determine a predicted speaker identifier corresponding to each audio segment; processing the audio feature vector corresponding to a target predicted speaker identifier by using a trained audio clustering model to obtain a target speaker identifier corresponding to each audio feature vector; and merging the audio segments corresponding to the same target speaker identifier to obtain a target audio corresponding to each target speaker identifier. By using the embodiment of the specification, the correction of the predicted speaker identifier is realized, and the accuracy of the audio separation is improved.
Owner:CHINA CITIC BANK CO LTD

A method for regulatory speech segmentation based on speech recognition and end-point detection

ActiveCN117238279BAutomatic segmentationSpeech segmentation
The application provides a regulation voice segmentation method based on speech recognition and endpoint detection, which is applied to air traffic control voice audio stream segmentation, and comprises the following steps: step 1, constructing a punctuation model based on speech recognition and a speech endpoint detection model; step 2, using the punctuation model based on speech recognition to recognize the audio data stream of the regulation voice, and outputting the corresponding text and sentence end identifier of the audio data stream; step 3, using the speech endpoint detection model to judge the speech starting point and ending point contained in the audio data stream of the regulation voice; step 4, segmenting the audio data stream of the regulation voice into audio segments; and step 5, applying the audio segments as data materials to the speech recognition process of the air traffic control system. Through the combination of speech recognition and endpoint detection, the application realizes the automatic segmentation of the air traffic control voice audio stream, and improves the accuracy and efficiency of the segmentation.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

A small sample semantic segmentation method and system fusing class label semantics

ActiveCN118587440BCharacter and pattern recognitionSpeech segmentationImaging processing
The application relates to the field of image processing, and proposes a small sample semantic segmentation method and system fusing class label semantics, wherein a prior information generation module is designed, multi-modal data fusion of image data information and text data information serving as class labels is realized, a semantic segmentation model can more accurately understand image content, a multi-scale fusion module is further designed, original detail information of an image is further protected, the calculation performance and channel fusion capability of the semantic segmentation model are greatly improved, the parameter quantity of a decoder is greatly reduced while ensuring the speech segmentation precision, and the accuracy of target recognition and target positioning of semantic segmentation is greatly improved.
Owner:JIANGXI NORMAL UNIV

A language data preprocessing method for multilingual and complex scenes

ActiveCN120913580BSpeech recognitionSpeech segmentationTransient noise
The application relates to the technical field of language data processing, and discloses a language data preprocessing method for multiple languages and complex scenes, a speech data preprocessing system for multiple languages and complex scenes based on an AutoPrep framework, and five modules of speech enhancement, speech segmentation, speaker clustering, target speech extraction and quality filtering are integrated to realize automatic and structured processing of speech data. The scheme realizes differentiated suppression of stable-state and transient-state noises in multiple language speech signals, especially in small language (such as Kazakh and Tagalog) scenes, effectively improves the speech signal-to-noise ratio and the language independence of speech characteristics, overcomes the problem that the existing technology lacks a special phonetic system processing module for small languages, resulting in a high phoneme mapping error rate, and enhances the availability and processing effect of low-resource language data.
Owner:SHENYI FUTURE TECHNOLOGY (GUANGDONG HENGQIN) CO LTD

Speaker voice segmentation method based on non-local space U-Net and mixed features

PendingCN121983030AImprove discrimination abilityFully explore the spatial and temporal correlationsSpeech recognitionSpeech segmentationSemantic feature
The invention belongs to the technical field of voice signal processing, and particularly provides a speaker voice segmentation method based on non-local space U-Net and mixed features. According to the invention, the non-local space U-Net network is constructed, and a non-local space attention module is introduced, so that the long-range dependency relationship in voice signals is effectively captured, and the expression ability of spatial features is improved; and meanwhile, a mixed feature fusion mechanism is adopted, and time-frequency domain features and deep semantic features are combined, so that the discrimination of voice features is enhanced. In addition, a dynamic weight loss function is designed, the fusion efficiency of the features under different scales is optimized, and the contribution degrees of various voice segments are balanced. According to the method, the space-time relevance in the voice signals can be fully mined, the dependence of a traditional method on local features and single features is overcome, and the accuracy and robustness of speaker voice segmentation are remarkably improved.
Owner:CHINA CRIMINAL POLICE UNIV

A multi-scene intelligent assistant system and interaction method based on dialect voice wake-up

The application relates to the technical field of intelligent voice interaction, in particular to a multi-scene intelligent assistant system based on dialect voice wake-up and an interaction method, which comprises a voice segmentation module, a scene screening module, a terminal access module, an intra-domain wake-up module and an entry determination module. The application establishes a continuous processing relationship by surrounding the voice segment boundary, the scene label residence state, the terminal release state and the dialect wake-up word structure, narrows the scene range of the current voice segment participating in identification according to the scene label residence state, checks the tone fluctuation, the syllable sequence and the tailing convergence in the limited range, judges the interaction entry attribution in combination with the terminal access state and the activation confirmation information, keeps the scene constraint and the dialect wake-up discrimination synchronous, keeps the terminal access and the entry confirmation connected, controls the cross-scene repeated matching, the near-sound trigger confusion and the terminal access conflict, and makes the dialect voice wake-up and the subsequent interaction more coherent.
Owner:XIAMEN LUJIANG TECHNOLOGY CO LTD

A method, apparatus, device, and readable storage medium for speech segmentation.

PendingCN122313977ASpeech segmentationConfusion
This application discloses a speech segmentation method, apparatus, device, and readable storage medium, relating to the field of speech interaction technology. It includes: performing real-time semantic analysis on the user's speech stream locally on the terminal device to determine the semantic integrity status, and dynamically selecting a target threshold from a preset set of silence thresholds to extract audio segments and send them to the cloud; the cloud performs endpoint detection on the audio segments, obtains the start and end timestamps of the speech, and calculates the time interval between the current segment and the previous segment; when the interval is less than or equal to a preset listening window, it performs a semantic correlation judgment on the two segments based on rules and a semantic model, and decides whether to splice them. This application solves the problems of response delay and semantic confusion caused by fixed thresholds through semantically driven dynamic threshold adjustment and cloud-based semantic-level splicing decision-making, thereby improving the response speed and recognition accuracy of speech interaction.
Owner:AISPEECH CO LTD

Speaker voice segmentation method based on non-local spatial U-Net and mixed features

ActiveCN121862087BSpeech segmentationSemantic feature
This invention belongs to the field of speech signal processing technology, specifically proposing a speaker speech segmentation method based on nonlocal spatial U-Net and hybrid features. This invention constructs a nonlocal spatial U-Net network, and by introducing a nonlocal spatial attention module, effectively captures long-range dependencies in the speech signal, enhancing the expressive power of spatial features. Simultaneously, a hybrid feature fusion mechanism is employed, combining time-frequency domain features and deep semantic features to enhance the discriminative power of speech features. Furthermore, a dynamic weight loss function is designed to optimize the fusion efficiency of features at different scales and balance the contribution of various speech segments. This invention can fully exploit the spatiotemporal correlations in speech signals, overcome the dependence of traditional methods on local and single features, and significantly improve the accuracy and robustness of speaker speech segmentation.
Owner:CHINA CRIMINAL POLICE UNIV