Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

42 results about "Speech segmentation" patented technology

Speech segmentation is the process of identifying the boundaries between words, syllables, or phonemes in spoken natural languages. The term applies both to the mental processes used by humans, and to artificial processes of natural language processing.

Classroom behavior analysis method and system, electronic equipment and storage medium

The invention provides a classroom behavior analysis method and system, electronic equipment and a storage medium, and relates to the technical field of educational informationization, and the method comprises the steps: collecting videos and audios, carrying out image enhancement, noise filtering and frame segmentation on the videos, carrying out noise reduction, sound source positioning and voice segmentation on the audios, and generating standardized images and voice sequences; an improved MTCNN cascade network is combined with a posture estimation technology, facial key points of students are extracted from videos, class arrival states, head postures and facial micro-expressions are recognized, audios are analyzed through a bidirectional LSTM network, and speaking duration and frequency are obtained; constructing individual behavior indexes according to attendance, head actions and expressions of the students; constructing an interaction index according to the teacher and student speaking data; the individual and interaction indexes are summarized to generate the multi-dimensional classroom behavior portrait, and the relation between the portrait and the teaching effect is analyzed, so that the classroom behavior analysis precision can be improved, the data processing capability can be enhanced, and the multi-dimensional classroom behavior comprehensive evaluation can be realized.
Owner:NANJING LANZHONG INTELLIGENT TECH CO LTD

Short voice-based voiceprint clustering method guided by speaker recognition pre-training model

PendingCN120375834ASpeech analysisSpeech segmentationFeature Dimension
The invention discloses a voiceprint clustering method guided by a speaker recognition pre-training model based on short voices, and the method comprises the following steps: obtaining an original voice signal, and randomly combining a plurality of enhancement strategies to achieve data enhancement; performing voice segmentation on the voice signal after data enhancement based on a uniform segmentation mode; extracting voiceprint features of the segmented voice based on an attention mechanism of global time-frequency domain context modeling; and obtaining a clustering result based on K-means clustering and spectral clustering, matching the clustering result with a real speaker tag, performing reverse transmission based on angle-dependent AAM-Softmax loss, and outputting a voiceprint clustering result. According to the method, the influence of environmental interference on feature extraction can be overcome, feature dimensions which are more effective for identity identification can be screened, a clustering output effect which is superior to that of a mainstream algorithm can be obtained with a relatively low parameter quantity, and the robustness under a noise interference condition is improved.
Owner:SOUTH CHINA UNIV OF TECH

Electronic medical record automatic generation method based on voice recognition

InactiveCN120690205ASpeech recognitionPatient-specific dataMedical recordSpeech segmentation
The invention relates to the technical field of electronic medical record generation, and discloses an electronic medical record automatic generation method based on voice recognition, which comprises the following steps: S1, initializing voice input, distributing a unique voice acquisition identifier for a medical session, and completing identifier generation, input, storage, association and identity verification; s2, voice information intelligent recognition: converting voice into a text by using a voice recognition engine, and ensuring semantic consistency through a context verification unit and a semantic analysis unit; and S3, performing multi-speaker processing based on intelligent interference detection and resolution, positioning an interference time period and an interference source through an interference detection unit, and realizing time period distribution and priority ranking of multi-speaker voices by using a voice segmentation protocol and a linear weighting model. And finally, extracting related information from the text generated by voice conversion, and filling the related information into a medical record template of a hospital. The method improves the efficiency and accuracy of electronic medical record generation, solves the problems of multi-speaker interference and semantic logic, and is suitable for medical informatization scenes.
Owner:THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)

Speech processing method and apparatus, device, and medium

PendingUS20250329334A1Speech analysisSpeech segmentationAcoustics
A speech processing method includes: obtaining overlapping speech data; obtaining reference speech data of a specified object; extracting a voiceprint representation vector of the specified object from the reference speech data, the voiceprint representation vector representing a voiceprint characteristic of the specified object, and inputting the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, and segmenting, by the speech segmentation model based on an attention mechanism, the overlapping speech data to obtain a target speech signal matching the voiceprint characteristic; and generating a speech file of the specified object based on the speech signal.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Speech translation using a wearable device

A speech translation system may provide real-time or near real-time translation of speech uttered by a person or emitted from a media device. The speech translation system may include a device that may receive audio representing speech in a source language and output audio representing speech in a target language. The speech translation system may translate the speech in portions representing semantically cohesive speech segments such that the target speech reflects the semantic meaning of words, phrases, and / or clauses as used in the context of the source speech. The speech translation system may condense the speech segments prior to or during translation to reduce verbosity. The speech translation system may selectively translate some speakers and not others, and may determine voice characteristics of source speech and apply identifying characteristics to the target speech that allow a user to differentiate respective target speech from different speakers based on the identifying characteristics.
Owner:AMAZON TECH INC

Speech synthesis method, speech synthesis device, electronic equipment and storage medium

PendingCN120636366ASpeech recognitionSpeech synthesisSpeech segmentationSynthesis methods
The embodiment of the invention provides a speech synthesis method, a speech synthesis device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for the field of financial science and technology and the field of digital medical treatment. The method comprises the steps of obtaining original voice, performing voice segmentation on the original voice to obtain candidate voice segments and speaker identifiers of the candidate voice segments, merging the candidate voice segments according to the speaker identifiers to obtain reference voice segments, screening the reference voice segments to obtain target voice segments, and performing voice recognition on the target voice segments to obtain voice recognition results. The method comprises the steps of obtaining a target text of a target voice segment, constructing a voice text pair according to the target voice segment and the target text, performing model updating on an original voice synthesis model according to the voice text pair to obtain a target voice synthesis model, and performing voice synthesis on a preset reference text through the target voice synthesis model, so that the voice synthesis quality can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Doctor-patient speech communication model training method and system based on multi-modal corpus analysis

PendingCN121725770ASpeech recognitionPattern recognitionSpeech segmentation
The invention discloses a doctor-patient speech communication model training method and system based on multi-modal corpus analysis, and belongs to the crossing field of artificial intelligence and medical treatment, and the method comprises the steps: carrying out the extraction according to an OpenPose algorithm to obtain motion features, carrying out the muscle activity intensity detection to obtain expression features, carrying out the speech recognition to obtain text features, and carrying out the recognition of the text features; speech segmentation is carried out based on the time domain energy parameters, and a segmentation result is subjected to modal analysis to obtain acoustic features; performing feature fusion based on an attention mechanism to obtain fusion features, and inputting an obstacle recognition model to obtain an obstacle type; the method comprises the steps of obtaining an obstacle type, obtaining an intervention strategy and an intervention identity according to a solution mapping relation, taking the obstacle type, the intervention strategy and the intervention identity as multi-dimensional labels, carrying out time domain alignment and structured packaging to obtain target data, and training according to the target data to obtain a communication model used for providing dialogue prompt information. Obstacle recognition accuracy can be improved, and the intelligent agent application effect can be improved.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

System and method for enhanced customer service through automated real-time FAQ generation from call center interactions

An automated system for generating Frequently Asked Questions (FAQs) from call center interactions includes a call center device for recording audio conversations between agents and customers, and a backend system. The backend system segments the conversation between agent and customer speech, converts the segmented speech into text using an Automatic Speech Recognition engine, and generates FAQs using a Large Language Model. Each FAQ includes a query statement corresponding to the customer's speech and at least one answer statement corresponding to the agent's speech. The system also includes mechanisms for selecting relevant, non-duplicate FAQs, determining the importance of each FAQ based on frequency, sentiment, and coherence scores, and dynamically updating the FAQ database in real-time. A user interface displays the generated FAQs with dropdown arrows to view the answers, enhancing customer service efficiency and accuracy by providing immediate, relevant responses to common inquiries.
Owner:ELM INC

A method for generating a speaker diary based on audio-visual fusion clustering

ActiveCN119964596BSpeech recognitionSpeech segmentationSpeaker verification
The application discloses a speaker diary generation method based on audio-visual fusion clustering, and aims to solve the problem of "who speaks at what time" in a multi-speaker scene. The method is realized through the following steps: first, an overlapping-aware speech segmentation model is used to segment the audio segment, solving the problem of overlapping speech; second, an advanced speaker verification model is used to extract the speaker voiceprint features of each audio segment and a speaker score matrix generated through face tracking and speaker detection; then, through an audio-video joint clustering method, the number of clusters is optimized according to the audio features and visual information, and K-means clustering is used to complete speaker clustering; the experimental results show that the system adopting the method achieves the lowest diary error rate (DER) on the Ego4D validation set.
Owner:HUNAN UNIV

A real-time voice conversion method and device, electronic equipment and medium

ActiveCN115910083BSpeech recognitionSpeech synthesisSpeech segmentationEngineering
The application provides a real-time voice conversion method and device, electronic equipment and medium. The method comprises the following steps: intercepting first voice data meeting voice segmentation conditions from voice data of a source speaking object recorded in real time; processing the first voice data to extract first semantic information; inputting the first semantic information into a pre-trained voice conversion model, and converting and processing effective information of historical voice data before the first semantic information and the first voice data through the voice conversion model to obtain target voice feature information corresponding to the first semantic information and a voice factor of a target speaking object; reconstructing the target voice feature information to obtain second voice data converted from the first voice data, thereby realizing low-delay streaming inference and low-delay and high-performance real-time voice conversion.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

AI-based telephone answering system

The application discloses an AI-based telephone answering system, relates to the technical field of telephone answering, and comprises a real-time voice acquisition and preprocessing module, a speech speed detection and prediction module, a self-adaptive ASR dynamic regulation module, a refined voice slicing and decoding module, a keyword detection and voice recognition module, a semantic understanding and intention analysis module and an emergency response and dispatching module; the real-time voice acquisition and preprocessing module acquires user voice data of an emergency help telephone in real time, establishes a real-time audio stream transmission channel, and rapidly preprocesses the acquired audio signal. The application rapidly and accurately captures the voice features of a user under high speech speed, dynamically adjusts the voice segmentation length of an ASR engine, makes each voice segment more clear and accurate, avoids the splitting and recognition errors of cross-segment words caused by too fast speech speed, and effectively avoids the omission of user emergency information and the risk of misjudgment in an emergency help scene.
Owner:SHANDONG ZHIQUN INFORMATION TECH CO LTD

Conference management method and system based on natural language processing and retrieval enhancement

ActiveCN120634500AReservationsBiological modelsSpeech segmentationText stream
The invention provides a conference management method and system based on natural language processing and retrieval enhancement, and belongs to the technical field of conference management, and the method comprises the steps: outputting a conference reservation result based on key parameters; if the reservation is successful, acquiring audio stream data in the conference process, and performing voice segmentation, initial voiceprint extraction and initial clustering on the audio stream data to output an original voice segment and speaker reference voice; target speaker extraction, accurate voiceprint extraction and final clustering are carried out based on the original voice segment and the speaker reference voice to obtain a pure voice text stream with a voiceprint label; performing dynamic abstract generation, enhanced retrieval and verification on the pure voice text stream; according to the conference management system and method, the intelligent management of the whole process of the conference can be covered, and meanwhile, the whole process management of the conference from voice-driven reservation, real-time content analysis to automatic task distribution is realized.
Owner:GUANGDONG KAMFU TECH CO LTD

An audio separation and script violation reminding method, device and computer equipment

ActiveCN116631432BSpeech analysisSpeech segmentationFeature extraction
The specification relates to the technical field of artificial intelligence, in particular to an audio separation and script violation reminding method and device and a computer device. The audio separation method comprises the following steps: performing speech segmentation on a to-be-separated audio to obtain a plurality of audio segments; performing feature extraction on each audio segment to obtain an audio feature vector corresponding to the audio segment; processing the audio feature vector by using a trained speaker prediction model to determine a predicted speaker identifier corresponding to each audio segment; processing the audio feature vector corresponding to a target predicted speaker identifier by using a trained audio clustering model to obtain a target speaker identifier corresponding to each audio feature vector; and merging the audio segments corresponding to the same target speaker identifier to obtain a target audio corresponding to each target speaker identifier. By using the embodiment of the specification, the correction of the predicted speaker identifier is realized, and the accuracy of the audio separation is improved.
Owner:CHINA CITIC BANK CO LTD

Voice segmentation method, device, server and storage medium

ActiveCN115101056BSpeech recognitionSpeech segmentationAcoustics
The present application is applicable to the field of artificial intelligence technology and provides a speech segmentation method, device, server, and storage medium. The method includes: obtaining speech data and determining whether the speech data includes multiple speech frequencies; if the speech data includes multiple speech frequencies, performing speech separation on the speech data based on a preset frequency set and the speech frequencies of each part in the speech data to obtain multiple target speech data; determining target speech segmentation points of the target speech data based on speech pause information in the target speech data and target speech text corresponding to the target speech data, and performing speech segmentation on the target speech data based on the target speech segmentation points. The present application can automatically perform speech segmentation on the speech data of the target user, which helps to improve the efficiency and accuracy of speech segmentation.
Owner:PING AN BANK CO LTD

System and method for enhanced customer service through automated real-time FAQ generation from call center interactions

An automated system for generating Frequently Asked Questions (FAQs) from call center interactions includes a call center device for recording audio conversations between agents and customers, and a backend system. The backend system segments the conversation between agent and customer speech, converts the segmented speech into text using an Automatic Speech Recognition engine, and generates FAQs using a Large Language Model. Each FAQ includes a query statement corresponding to the customer's speech and at least one answer statement corresponding to the agent's speech. The system also includes mechanisms for selecting relevant, non-duplicate FAQs, determining the importance of each FAQ based on frequency, sentiment, and coherence scores, and dynamically updating the FAQ database in real-time. A user interface displays the generated FAQs with dropdown arrows to view the answers, enhancing customer service efficiency and accuracy by providing immediate, relevant responses to common inquiries.
Owner:ELM INC

A method for regulatory speech segmentation based on speech recognition and end-point detection

ActiveCN117238279BAutomatic segmentationSpeech segmentation
The application provides a regulation voice segmentation method based on speech recognition and endpoint detection, which is applied to air traffic control voice audio stream segmentation, and comprises the following steps: step 1, constructing a punctuation model based on speech recognition and a speech endpoint detection model; step 2, using the punctuation model based on speech recognition to recognize the audio data stream of the regulation voice, and outputting the corresponding text and sentence end identifier of the audio data stream; step 3, using the speech endpoint detection model to judge the speech starting point and ending point contained in the audio data stream of the regulation voice; step 4, segmenting the audio data stream of the regulation voice into audio segments; and step 5, applying the audio segments as data materials to the speech recognition process of the air traffic control system. Through the combination of speech recognition and endpoint detection, the application realizes the automatic segmentation of the air traffic control voice audio stream, and improves the accuracy and efficiency of the segmentation.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

An intelligent real-time language synchronous translation system and its terminal

This application relates to an intelligent real-time language synchronous translation system and its terminal. The system includes: a speech segmentation module for performing semantic analysis on the target speech to be translated, segmenting the target speech with the obtained semantic information to obtain speech units; a scene recognition module for analyzing the speech units to obtain scene information; a syntactic rhythm analysis module for converting the speech units into text and then performing syntactic analysis, and analyzing the speech units in combination with this information and the semantic information to obtain speech rhythm information and sentence structure information; a semantic optimization module for correcting the semantic information based on the sentence structure information, scene information, and speech rhythm information to obtain optimized semantic information; a speech translation module for translating the optimized semantic information based on the target language to obtain a target language translation. Through the cooperation of multiple modules, the system can accurately grasp the speech semantics, fit different scenarios, and effectively improve the accuracy, fluency, and practicality of translation.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

A small sample semantic segmentation method and system fusing class label semantics

ActiveCN118587440BCharacter and pattern recognitionSpeech segmentationImaging processing
The application relates to the field of image processing, and proposes a small sample semantic segmentation method and system fusing class label semantics, wherein a prior information generation module is designed, multi-modal data fusion of image data information and text data information serving as class labels is realized, a semantic segmentation model can more accurately understand image content, a multi-scale fusion module is further designed, original detail information of an image is further protected, the calculation performance and channel fusion capability of the semantic segmentation model are greatly improved, the parameter quantity of a decoder is greatly reduced while ensuring the speech segmentation precision, and the accuracy of target recognition and target positioning of semantic segmentation is greatly improved.
Owner:JIANGXI NORMAL UNIV

Voice data processing method and device, electronic equipment and storage medium

PendingCN120183395ASpeech recognitionSpeech segmentationAcoustics
The embodiment of the invention provides a voice data processing method and device, electronic equipment and a storage medium, relates to the technical field of natural language processing, and is suitable for the fields of financial science and technology and medical health. The method comprises the following steps: acquiring original mixed voice; converting the original mixed voice into a transcriptional text; extracting a target host keyword from the transcriptional text according to the candidate host keyword; according to the time point of the target host keyword in the original mixed voice, performing voice segmentation on the original mixed voice to obtain at least two mixed voice segments; performing voice separation on the mixed voice segment to obtain single speaking object voice sub-segments; performing speech object classification on the speech sub-fragments of the single speech object according to the speech features of the target speech object to obtain a classification result; and carrying out voice fragment splicing on the single speaking object voice sub-fragments of each classification result to obtain target single speaking object voice. According to the embodiment of the invention, the quality of voice data can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method and device for removing voice broadcast in voice signal, equipment and medium

PendingCN120340502ASpeech recognitionSpeech segmentationEngineering
The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a method, a device, equipment and a medium for removing voice broadcast in a voice signal, which comprises the following steps: establishing a pre-stored voice broadcast sample and a voice feature library thereof, segmenting a target voice into a plurality of voice segments, and generating a window voice segment based on the segmentation starting point and a pre-stored voice duration, identifying a voice broadcast segment through correlation analysis, recording the starting and ending positions of the voice broadcast segment, removing the voice broadcast according to the recorded positions, and generating a voice signal after the voice broadcast is removed. According to the method and the device, the voice is divided into the plurality of voice segments, and the window voice segments are generated based on the duration of the pre-stored voice broadcast sample, so that the voice broadcast segments in the target voice are quickly positioned, the voice broadcast is accurately recognized through correlation analysis, and the voice broadcast removal accuracy and the processing efficiency are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

A conference management method and system based on natural language processing and retrieval enhancement

ActiveCN120634500BReservationsBiological modelsSpeech segmentationText stream
The application provides a conference management method and system based on natural language processing and retrieval enhancement, and belongs to the technical field of conference management. The method comprises the following steps: outputting a conference reservation result based on key parameters; if the reservation is successful, acquiring audio stream data in the conference process, performing speech segmentation, initial voiceprint extraction and initial clustering on the audio stream data to output original speech segments and speaker reference speech; performing target speaker extraction, accurate voiceprint extraction and final clustering based on the original speech segments and the speaker reference speech to obtain pure speech text stream with a voiceprint label; performing dynamic abstract generation, enhanced retrieval and verification on the pure speech text stream; performing responsibility recognition and voiceprint binding on the conference information to output final conference management data. The application can cover intelligent management of the whole conference process, and realize conference whole-process management from voice-driven reservation, real-time content analysis to automatic task distribution.
Owner:GUANGDONG KAMFU TECH CO LTD

A Speech Segmentation Method and System for an Air Traffic Control Voice Recorder

ActiveCN118968970BSpeech recognitionSpeech segmentationTimestamp
The present invention discloses a method and system for voice segmentation of an air traffic control voice recorder, including: obtaining mixed-channel voice data and control-channel voice data output by the voice recorder frame by frame, each frame of voice data being accompanied by a timestamp and a channel number, where only control voice data exists in the control channel, and the mixed channel includes crew voice data and control voice data in the control channel; extracting voice data features frame by frame, and identifying whether the voice data in the mixed channel with the same timestamp is control voice data according to the control-channel voice data features frame by frame, and removing the voice data corresponding to the frame in the mixed channel to obtain the crew voice data in the mixed channel. The present invention can remove the control voice in the mixed channel in the air traffic control system according to the voice data features, and finally obtain separate control voice data and separate crew data, and output them through different channels, ensuring the accuracy of other applications such as subsequent voice recognition.
Owner:GUANGZHOU ZHONGNANMIN AVIATION GUAN COMM NETWORK TECH +1

Emotion recognition method in voice dialogue

PendingCN121191542ASpeech recognitionSpeech segmentationNoise
The invention relates to an emotion recognition method in voice dialogue, which comprises the following steps: S1, acquiring a voice stream, and converting the voice stream into a WAV format; s2, performing noise suppression on the voice stream in the WAV format; s3, performing voice segmentation processing on the voice stream after noise suppression; s4, carrying out standardization processing on the single-segment voice waveform after segmentation processing; s5, performing feature coding on the standardized voice waveform signal to obtain a feature vector of the input waveform; and S6, inputting the feature vector into a classifier, and outputting the most probable emotion in a given emotion category. According to the invention, automatic recognition of human emotional states is realized by analyzing acoustic features in voice signals, and the problem of poor robustness of an accompanying robot in a complex environment is solved; the dialogue of the accompanying robot is more natural and intelligent, the robot has the emotional sharing ability, and the robot is endowed with understanding and response to the emotion of human beings.
Owner:SHANGHAI HUAHUA INTELLIGENT TECH CO LTD

Interactive speech segmentation and clustering method, device and equipment

ActiveCN114708850BSpeech recognitionSpeech segmentationSpeech sound
The present invention discloses an interactive speech segmentation and clustering method, apparatus, device, and storage medium, which includes: preprocessing audio data to be processed to obtain N types of speech; auditing the N types of speech and merging speech belonging to the same person to obtain M types of speech, wherein the M types of speech correspond to the number of people in the audio conversation; calculating the center vector of each type of speech and the similarity of each speech segment contained in each type of speech based on the M types of speech, and marking speech segments whose similarity is lower than a preset value; auditing the marked speech segments and reallocating the marked speech segments to obtain audio classification results. This method can improve the accuracy of speech segmentation and clustering results.
Owner:XIAMEN KUAISHANGTONG TECH CORP LTD

Partially-forged speech detection method and system based on speech spectrum characteristics and deep learning

PendingCN120452450ASpeech analysisSpeech segmentationSpeech sound
The invention discloses a partially-forged voice detection method and system based on speech spectrum features and deep learning, relates to the technical field of voice authenticity detection, and solves the problem that the existing forged voice detection method cannot effectively detect forged voice carrying short and small forged fragments, so that the detection accuracy is low. The method comprises the following steps of: extracting a Mel language spectrogram of voice, and dividing the Mel language spectrogram into a plurality of language spectrum sub-graphs according to a time direction; inputting the plurality of speech score sub-graphs into a pre-trained deep learning forged speech detection model composed of a linear projection layer, a Transformer encoder network and an xLSTM discriminator network, and obtaining real scores representing the corresponding speech score sub-graphs; and real scores of each speech spectrum sub-graph are fused to obtain a speech-level real score, so that a detection result is obtained, deep representation of the speech spectrum can be learned, the difference between true and false speech segments can be effectively captured, and high-accuracy detection of part of forged speech is realized.
Owner:TONGXIANG GENERAL ARTIFICIAL INTELLIGENCE RESEARCH INSTITUTE +1

Miao language speech adaptive segmentation method and system based on time domain features

The invention provides a Miao language speech adaptive segmentation method and system based on time domain features, and relates to the technical field of speech segmentation. The method comprises the following steps: pre-recording voice audio of a single channel, and preprocessing the voice audio to obtain short-time energy and a short-time zero-crossing rate; the syllables are extracted through the short-time energy and the short-time zero-crossing rate, and voice syllables are obtained; calculating the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculating the short-time zero-crossing rate to obtain a minimum zero-crossing rate; pre-segmenting the voice frame through the maximum short-time energy, the minimum short-time energy and the minimum zero-crossing rate to obtain a plurality of pre-segmented voice segments; and constructing a fitness evaluation function model, and performing adaptive segmentation on the plurality of pre-segmented voice segments through the fitness evaluation function model to obtain a plurality of optimal voice segments. A syllable boundary fuzzy problem is converted into an actual relation problem between a syllable real boundary and a predicted boundary, so that the speech syllable adaptive boundary search capability is improved.
Owner:GUIZHOU MINZU UNIV

Binary natural speech segmentation method and system based on time domain analysis

PendingCN120998213ASpeech recognitionTime domainSpeech segmentation
The invention provides a binary natural speech segmentation method and system based on time domain analysis. The method comprises the following steps: acquiring a target binary natural voice file and generating corresponding binary natural voice data; performing code element-by-code analysis on the binary natural voice data, and identifying whether the current code element is an error code, a normal code or a cut-off position of a code word or a code block according to a relation between a time interval between the current code element and an adjacent code element and a preset time interval coefficient; and performing code word segmentation and code block segmentation on the binary natural voice data according to an identification result.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University