Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Audio segmentation" patented technology

System and method for adaptive audio segmentation for contextual speech signal processing

ActiveUS20250391420A1Speech recognitionAudio segmentationNoise
A system for audio segmentation for context change detection in speech is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system detects a potential context change between a first audio frame and a second audio frame of the speech signal. The system splits the speech signal into a first set of split audio frames based on the potential context change. The system generates a noisy speech signal by modulating a flicker noise signal into the speech signal. The system splits the noisy speech signal into a second set of split audio frames. The system detects a difference between the first and second sets of split audio frames. The system reconfigures a modulation of the speech signal with the flicker noise signal to reduce the difference between the first and second sets of audio frames.
Owner:BANK OF AMERICA CORP

System and method for dynamic audio slicing window selection based on context and speech patterns

ActiveUS20250391404A1Speech recognitionAudio segmentationSpeech patterns
A system for an audio slicing window selection for contextually splitting a speech signal is disclosed. The system identifies a first audio processing software algorithm that is assigned to a user. The system identifies a set of audio processing software algorithms and configures each of them with a respective audio slicing window. The system selects a second audio processing software algorithm, from among the set of audio processing software algorithms. The system selects one of the first and second audio processing software algorithms that is assigned an audio slicing window associated with the context of the speech signal. The system splits the speech signal using the selected audio processing software algorithm. The system determines whether the speech signal is split contextually. In response to determining that the speech signal is not split contextually, the selected audio processing software algorithm and / or the audio slicing window may be updated.
Owner:BANK OF AMERICA CORP

Audio processing method and device, terminal equipment, storage medium and program product

PendingCN121054023ASpeech analysisAudio segmentationTerminal equipment
The invention relates to an audio processing method and apparatus, a terminal device, a storage medium and a program product. The audio processing method comprises the steps of obtaining a first audio; segmenting the first audio into a plurality of second audios; determining playing parameters of each second audio; wherein the playing parameters are different, and sound effects obtained by playing the same second audio are different; and according to the playing parameters of the second audios, playing the corresponding second audios so as to realize playing of the first audios after audio processing. According to the embodiment of the invention, the sound effect after each second audio is played is improved, the situation that part of audio in the first audio is not matched with the playing parameter due to the fact that the first audio is played according to the same playing parameter is reduced, the situation that the sound effect is poor due to the fact that the playing parameter is not matched with the audio is also reduced, and the use experience of a user is improved. Manual operation of a user is not needed, so that operation steps of the user are reduced, the method is more intelligent and convenient, and the use experience of the user is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Automated segmentation and transcription of unlabeled audio speech corpus

ActiveUS12512100B2Speech recognitionTimestampAudio segmentation
A method includes obtaining initial transcription for input natural speech; performing segmentation of initial transcription into text portions, based on punctuation marks in initial transcription; determining segment-level timestamps for text portions based on the input natural speech; performing audio segmentation on input natural speech, by cutting input natural speech based on segment-level timestamps, to obtain audio chunks; generating transcription portions for each of the audio chunks; merging transcription portions to form re-transcription; determining word-level timestamps for re-transcription, by aligning input natural speech against re-transcription; calculating silence time periods, each corresponding to silence between each two adjacent words of input natural speech, based on word-level timestamps; performing a final segmentation on input natural speech and re-transcription, based on silence time periods, to generate final audio segments and corresponding final transcription portions. The final audio segments and corresponding final transcription portions may be included in training dataset for training a model.
Owner:ORACLE INT CORP

Psychological state monitoring method and system based on deep learning voice emotion vector

The invention discloses a psychological state monitoring method and system based on a deep learning voice emotion vector, which can accurately measure the psychological state of a pilot in a simulated flight training process, and overcome the problems of strong subjectivity, complex equipment, easy interference, non-deep voice emotion analysis and the like in the existing measurement method. And a powerful guarantee is provided for flight safety. The method comprises the following steps: acquiring voice data of a tested person, and performing audio processing on the acquired voice data to obtain an effective voice audio file set after audio segmentation and de-muting; adopting a pre-trained deep voice emotion calculation model to extract a high-dimensional voice emotion vector of each audio clip; and sequentially inputting the extracted voice emotion vectors into the trained psychological state evaluation model to generate psychological characteristic indexes.
Owner:CHINA EASTERN TECH APPL RES & DEV CENT CO LTD

High privacy DSP-based audio anonymization with audio segmentation and randomization

PendingUS20250384891A1Speech analysisAudio segmentationAudio frequency
A method and an electronic device for generating an anonymized audio output are provided. The method, executable by the electronic device, comprises acquiring an audio recording of a speaker; stochastically determining a base pitch value based on at least a first probabilistic function; segmenting the original audio input into a plurality of audio segments, each of the plurality of audio segments being associated with a respective pitch. For each audio segment, the method further comprises generating a pitch adjustment value using a combination of the base pitch value of the segment and a value determined using a second probabilistic function; generating an adjusted audio segment by adjusting the pitch of the audio segment using the pitch adjustment value, the adjusted audio segment having an adjusted pitch that is different from the original pitch; generating the anonymized audio output by combining the adjusted audio segments.
Owner:HUAWEI TECH CO LTD

Audio segmentation method and device, electronic equipment and storage medium

The application provides an audio segmentation method and device, electronic equipment and storage medium, wherein the method comprises: obtaining audio to be segmented; extracting acoustic features of each frame in the audio to be segmented, and based on the acoustic features of each frame, performing semantic boundary sequence labeling on the audio to be segmented to obtain semantic boundary labeling results of each frame; and based on the semantic boundary labeling results of each frame, performing segmentation on the audio to be segmented. The method, device, electronic equipment and storage medium provided by the application can assist semantic segmentation based on tone and pause information in the acoustic features of each frame, retain complete semantic information of the audio, and avoid punctuation recognition errors, thereby improving the accuracy and reliability of audio segmentation. Furthermore, the method can be applied to a cascaded speech translation system and an end-to-end speech translation system, thereby expanding the application range of audio segmentation.
Owner:HKUST IFLYTEK (SHANGHAI) TECH CO LTD

Audio transcription method and device based on intelligent analysis, equipment and medium

PendingCN121600931ASpeech recognitionSpeech synthesisAudio segmentationEngineering
The invention relates to the technical field of audio recognition, and relates to an audio transcription method and system based on intelligent analysis, and the method comprises the steps: carrying out the file verification and file slicing uploading of a to-be-transcribed file, and obtaining an uploaded audio file; performing audio segmentation on the uploaded audio file, and extracting a voice text after audio segmentation to obtain a primary transcription text; calculating an analysis transcription progress and a real-time transcription progress of the uploaded audio file, performing maximum screening on the analysis transcription progress and the real-time transcription progress to obtain a transcription progress, and updating the primary transcription text into a standard transcription text according to the transcription progress; performing feature activation on the text context feature of the standard transcription text to obtain a text abstract feature; performing feature style decoding on the text abstract features according to a style keyword input by a user to obtain a decoded text abstract; and combining and splicing the standard transcriptional text and the decoded text abstract to obtain an audio transcriptional file. According to the invention, the audio transcription efficiency can be improved.
Owner:SHENZHEN LEXIN SOFTWARE TECH CO LTD

Computer-implemented method for adaptive decryption of audio stream, and device and storage medium

Disclosed in the present invention are a computer-implemented method for adaptive decryption of an audio stream, and a device and a storage medium. The method comprises: acquiring an MPEG-DASH manifest file and parsing same; initializing a standard DRM interface; performing feature extraction on audio segment information in an identified audio segmentation mode, and on the basis of the extracted features, dynamically adjusting a segment request and processing logic, so as to perform adaptive segment downloading and pre-processing; extracting encryption parameters to identify an encryption flag of a target audio stream platform, and on the basis of the encryption flag, selecting a corresponding decryption algorithm to decrypt a current audio segment; and performing audio frame reconstruction on decrypted data to form an updated audio stream, buffering the updated audio stream into a preset buffer on the basis of an adaptive buffer policy, and outputting a final audio stream, so as to ensure the continuity of the final audio stream and realize efficient parallel processing of decryption and playback, thereby improving the decryption efficiency, and enhancing the user experience especially in high-concurrent scenarios.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Audio segmentation method and device, electronic equipment and storage medium

ActiveCN115719596BSpeech analysisAudio segmentationEngineering
The application provides an audio segmentation method and device, electronic equipment and storage medium, wherein the method comprises: determining a double-channel audio to be segmented; respectively performing mute section labeling on a first channel audio and a second channel audio in the double-channel audio to obtain a mute section in the first channel audio and a mute section in the second channel audio; determining a common mute separation point in the double-channel audio based on the mute section in the first channel audio and the mute section in the second channel audio, and performing segmentation on the first channel audio based on the common mute separation point to obtain a plurality of first segmented audio sections; performing mute section removal on each first segmented audio section to obtain each second segmented audio section; and performing customer audio combination based on a voiceprint feature of each second segmented audio section to obtain customer audio in units of customers, thereby overcoming the defect that a fixed-time segmentation cannot distinguish customers, realizing audio segmentation in units of customers, and providing assistance for different service quality inspection and service evaluation.
Owner:IFLYTEK CO LTD

System and method for adaptive audio segmentation for contextual speech signal processing

ActiveUS12646523B2Speech recognitionAudio segmentationNoise
A system for audio segmentation for context change detection in speech is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system detects a potential context change between a first audio frame and a second audio frame of the speech signal. The system splits the speech signal into a first set of split audio frames based on the potential context change. The system generates a noisy speech signal by modulating a flicker noise signal into the speech signal. The system splits the noisy speech signal into a second set of split audio frames. The system detects a difference between the first and second sets of split audio frames. The system reconfigures a modulation of the speech signal with the flicker noise signal to reduce the difference between the first and second sets of audio frames.
Owner:BANK OF AMERICA CORP

A system and method for enhancing a call characterization of voice calls

A system and method for enhancing a call characterization of voice calls is disclosed. The method comprises: receiving, at the platform (1), the voice call (30) segmented in audio packets; generating a Call Detail Record "CDR" (12) comprising call parameters; processing, by the Digital Signal Processing "DSP" engine (300), the audio packets, wherein the processing step in turn comprises: computing audio spectral representations; and, accumulating a predetermined number of audio packets (30.1). The method further comprises: evaluating, by the Audio Segmentation "AS" engine (301), an audio type present in the accumulated audio packets; generating temporal marks (31), by the Audio Segmentation "AS" engine (301), indicating which audio segment correspond to each audio type; computing metrics (32), by the Metrics Computation "MC" engine (302), using the accumulated audio packets and the temporal marks; generating an enhanced Call Detail Record "eCDR" (33) by adding the metrics (32) to the "CDR".
Owner:BTS TECHNOLOGY SERVICES SA

Multi-level audio segmentation using deep embeddings

ActiveUS12586552B2Electrophonic musical instrumentsAudio segmentationAudio frequency
Embodiments are disclosed for generating an audio segmentation of an audio sequence using deep embeddings. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving an input including an audio sequence and extracting features for each frame of the audio sequence, where each frame is associated with a beat of the audio sequence. The method may further comprise clustering frames of the audio sequence into one or more clusters based on the extracted features and generating segments of the audio sequence based on the clustered frames, where each segment includes frames of the audio sequence from a same cluster. The method may further comprise constructing a multi-level audio segmentation of the audio sequence and performing a segment fusioning process that merges shorter segments with neighboring segments based on cluster assignments.
Owner:ADOBE INC

Multi-level data modeling-based pathological voice detection model training method

The embodiment of the invention provides a pathological voice detection model training method based on multi-level data modeling. The method comprises the following steps: inputting a dialogue audio into a pathological voice detection model; in the session layer, an encoder is used for encoding continuous audio clips obtained through dialogue audio segmentation, extracting depth features of the audio clips, capturing global context information, obtaining pathological pseudo-labels of the audio clips through depth feature classification, and determining session level loss based on the pathological pseudo-labels; in the fragment layer, modeling is carried out by taking the pathology pseudo labels as supervision signals so as to position audio fragments containing symptom features, depth features of the audio fragments are classified, a predicted pathology classification result is obtained, and fragment-level loss is determined based on the predicted pathology classification result; and performing end-to-end training on the pathological voice detection model based on the loss. According to the embodiment of the invention, the training remarkably improves the performance of the model under a very small amount of annotation data, and achieves the generalization of cross-language / disease species.
Owner:SHANGHAI JIAOTONG UNIV

Surgical safety verification method and system based on artificial intelligence voiceprint recognition

ActiveCN115938371BSpeech analysisAlgorithmAudio segmentation
This invention provides a surgical security verification method and system based on artificial intelligence voiceprint recognition, belonging to the field of artificial intelligence technology. In this invention, voice information is collected from the user to be verified to output corresponding audio of the user to be identified. Audio segmentation is performed on the audio of the user to be identified to form multiple audio segments corresponding to the audio of the user to be identified. Based on a pre-trained voiceprint recognition model, each audio segment of the multiple audio segments is recognized to output multiple audio recognition results corresponding to the multiple audio segments. Then, based on the multiple audio recognition results, an identity verification operation is performed on the user to be verified to output the user identity verification result. Based on the above method, the reliability of surgical security verification can be improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

A speech video generation method based on audio and video structure alignment

The application discloses a speech video generation method based on audio and video structure alignment and belongs to the virtual digital person field.The application comprises an audio segmentation module, an audio conversion module, an audio coding module, a video coding module and a video fusion decoding module.The audio conversion module is used for converting segmented phonemes into a mel-frequency spectrogram which is more in line with the frequency range of human ears according to Fourier transform.In the audio coding process, the frames of the same phonemes are taken as a continuous time module, and the time module is taken as time consistency to constrain the change of the lips, so that the fine-grained control of the lips is realized at the phoneme level through the time consistency constraint.In the video coding, the part region of the multi-pose change face in the input video is set as a mask region, the mask region is used as spatial consistency to accurately control the change amplitude of the lips, the position of the speaker's lips is aligned, the visual artifacts of the video are reduced, the facial details are optimized, and a high-quality speech video with audio-visual synchronization is generated.
Owner:BEIJING INST OF TECH

Audio segmentation method, device, equipment, medium and product

PendingCN121214919ABiological modelsSpeech recognitionAudio segmentationEngineering
The invention discloses an audio segmentation method and device, equipment, a medium and a product, and the method comprises the steps: constructing an initial meaning group segmentation prediction model, obtaining meaning group segmentation audio data, training the initial meaning group segmentation prediction model through employing the meaning group segmentation audio data, obtaining a target meaning group segmentation prediction model, and obtaining a target meaning group segmentation prediction model; and inputting the to-be-predicted audio data into the target meaning group segmentation prediction model to obtain a target segmentation position of the to-be-predicted audio data. According to the invention, the boundary of the audio data can be delimited through voice meaning group segmentation, the boundary delimitation is accurate, and the accuracy of audio segmentation can be improved.
Owner:IFLYTEK CO LTD

Hybrid multi-modal learning heterogeneous multi-codebook quantization video retrieval method and system

The present application relates to the technical field of multi-modal video retrieval, in particular to a heterogeneous multi-codebook quantization video retrieval method and system based on hybrid multi-modal learning, comprising the following steps: S1: importing a video to be retrieved; S2: performing key frame sampling and audio segmentation on the input video, and extracting video features and audio features; S3: introducing low-cost semantic knowledge and designing a prompt-driven multi-modal understanding module to encourage the model to focus on event-related information and enhance the correlation between the video and audio modalities; S4: mapping the heterogeneous modalities to a shared semantic space by designing a contrastive learning loss constructed by fusion video and audio prompts, thereby simultaneously enhancing cross-modal consistency and the transferability of the model; S5: designing a two-stage feature interaction and fusion module to fuse multi-modal features, inputting the fused features into a heterogeneous multi-codebook quantization module, and based on the Euclidean distance, retrieving the most similar database video to the to-be-tested video quantization code and outputting the retrieval result.
Owner:WUHAN INST OF TECH

Language interaction text information calibration method based on humiture large language model Agent

PendingCN121687015ASemantic analysisSpeech recognitionAudio segmentationEngineering
The invention discloses a language interaction text information calibration method based on a humiture large language model Agent, which relates to the field of speech recognition, solves the problem of poor calibration effect of the existing language interaction text information calibration method, and comprises the following steps: S1, acquiring real-time interaction audio and sampling to obtain a sampled audio clip, and storing the sampled audio clip in a database; the method comprises the following steps: S1, sampling an audio sample, carrying out audio analysis to obtain a speech loudness intrusive ratio and a fragment statement recognition coincidence ratio, and obtaining interactive audio collection data, S2, carrying out type division on real-time interactive audio by analyzing the interactive audio collection data, respectively carrying out audio statement segmentation to obtain real-time interactive audio segmentation data, and carrying out audio analysis on the real-time interactive audio segmentation data; and S3, performing vocabulary semantic verification on the text vocabularies according to the real-time interaction audio segmentation data, and performing vocabulary translation and outputting according to the verification result. The method can improve the pertinence and accuracy of the language interaction text information calibration method.
Owner:XINJIANG UYGUR AUTONOMOUS REGION INST OF MEASUREMENT & TESTING

Audio segmentation system performance evaluation method, device, equipment and medium

The invention relates to an audio segmentation system performance evaluation method and device, equipment and a medium. The method comprises the following steps: acquiring a reference segmentation result of a target audio signal and a prediction segmentation result obtained by predicting the target audio signal by an audio segmentation system; the reference segmentation result comprises a plurality of categories of reference audio clips, and the prediction segmentation result comprises a plurality of categories of prediction audio clips; performing alignment processing on the reference audio clip and the predicted audio clip according to categories to obtain audio clip pairs of multiple categories; and performing performance evaluation on the basis of the multiple types of audio clip pairs to obtain a performance evaluation result of the audio segmentation system. By adopting the method, the system performance evaluation accuracy can be improved.
Owner:CHINA TELECOM CLOUD TECH CO LTD

An audio segmentation and classification method based on multi-granularity slicing

This invention discloses an audio segmentation and classification method based on multi-granularity slicing, comprising: preprocessing the audio to obtain an audio file with a uniform sampling rate; slicing the audio file at different time granularities; extracting MFCC features from each slice at different time granularities and then performing image processing; establishing an image classification convolutional neural network model and training and validating it; inputting the processed audio into the image classification convolutional neural network model to obtain the classification result of each slice; and performing aggregation analysis based on the classification results to obtain the segmentation points and segment types of the audio file. This invention, by segmenting long audio files at different time granularities, utilizing an image classification convolutional neural network model for type judgment and classification, and finally performing aggregation analysis, can quickly and accurately find the segmentation points between different types of audio and determine the audio types of the audio segments before and after the segmentation points.
Owner:SICHUAN ZHONGYUN ZHIWANG TECH CO LTD

Cough sound recognition method based on large model parameter efficient fine tuning

The invention discloses a cough sound recognition method based on large model parameter efficient fine tuning, and belongs to the field of artificial intelligence auxiliary diagnosis. The method specifically comprises the following steps: constructing a data set containing cough sound and non-cough sound; segmenting the long audio into audio clips with fixed lengths, removing mute clips and unifying the sampling rates of all the audio clips; converting the audio clip into a two-dimensional time-frequency spectrogram through an audio feature extractor, and taking the two-dimensional time-frequency spectrogram as an input feature of a subsequent model; the method comprises the following steps: taking a pre-trained audio classification large model as a backbone network, and inserting a parameter efficient fine tuning module into an Encoder layer of a standard Transform in the backbone network; main body parameters of the backbone network are frozen, only parameters of a parameter efficient fine tuning module and a classification head of the model are trained, and an optimal model is selected for cough sound recognition; according to the method, a parameter efficient fine tuning technology is used, a parameter efficient fine tuning module is inserted into an audio classification large model, main body parameters of a backbone network are frozen, meanwhile, parameters of a classification head of the parameter efficient fine tuning module and the model are trained, and high-precision cough sound recognition is achieved while the training parameter quantity is reduced.
Owner:ZHENGZHOU UNIV

Audio emotion deep analysis system and method based on artificial intelligence

PendingCN121641075ASpeech analysisAlgorithmAudio segmentation
The invention provides an audio emotion deep analysis system and method based on artificial intelligence. The environment configuration channel stabilization module is used for constructing a voice emotion recognition model and constructing and forming a video voice emotion analysis tool based on deep learning; the video and audio extraction driving module is used for losslessly separating an audio track from a video file; an input video file needing to be processed is selected, and analysis task parameters are set; the audio segmentation and sentiment analysis module is used for segmenting and serializing the audio, adjusting interval values corresponding to sentiment dimension tendency information through linear conversion of original sentiment values, calling a voice sentiment recognition model through an interface, and performing sentiment dimension prediction on each audio segment according to a timestamp; and the result processing and presenting module gathers the timestamp audio clip sentiment analysis prediction information, automatically calculates and adds statistical information, draws the trend of a plurality of sentiment dimensions along with time into a curve graph, and obtains a time-varying trend curve graph of the sentiment dimensions.
Owner:SHANDONG XINYU TECH DEV CO LTD

Artificial intelligence-based automatic generation of lip-synced dubbing

A display device for generating real-time translated lip-synced dubbing comprises: a speaker diarization model for separating input audio into a background audio feed and an individual speaker audio feed; a face detection model for cropping a frame including a face; an active speaker detection model for pairing the individual speaker audio feed with a cropped frame corresponding to an active speaker; a translation model for obtaining translated speech audio; and a lip-sync model for generating lip-synced video frames including a facial movement of the active speaker synchronized with the translated speech audio by using a prediction model, and outputting the translated speech audio while displaying the lip-synced video frames in real time on the basis of the generated video frames, wherein the facial movement of the active speaker is synchronized with the output translated speech audio.
Owner:LG ELECTRONICS INC

Methods and Apparatus to Fingerprint an Audio Signal

Methods, apparatus, systems, and articles of manufacture to fingerprint an audio signal. An example apparatus disclosed herein includes an audio segmenter to divide an audio signal into a plurality of audio segments, a bin normalizer to normalize the second audio segment to thereby create a first normalized audio segment, a subfingerprint generator to generate a first subfingerprint from the first normalized audio segment, the first subfingerprint including a first portion corresponding to a location of an energy extremum in the normalized second audio segment, a portion strength evaluator to determine a likelihood of the first portion to change, and a portion replacer to, in response to determining the likelihood does not satisfy a threshold, replace the first portion with a second portion to thereby generate a second subfingerprint.
Owner:GRACENOTE INC

A Smart Classroom Interaction Analysis Method Based on Two-Layer Architecture Speech Segmentation

ActiveCN120783757BSpeech recognitionSpeech segmentationNerve network
This invention provides a smart classroom interaction analysis method based on a two-layer architecture speech segmentation, belonging to the field of speech segmentation technology. Specifically, it includes the following steps: extracting speech features from the speech signal using Mel-frequency cepstral coefficients (MFCC); designing a text-enhanced, multi-scale time-aware delay-based neural network to coarsely screen the speech features, dividing audio segments into single-speaker segments and multi-speaker segments; inputting the coarsely screened multi-speaker segments into a sliding window segmentation model (SW-NIF) that integrates neighbor window information to locate speaker transition points within the multi-speaker segments; training and validating the constructed model on a dataset. The technical solution of this invention overcomes the problem of neglecting classroom audio segmentation in existing technologies, which only perform simple segmentation of classroom audio for subsequent tasks, resulting in mixed speakers in the audio segments and affecting the analysis effect.
Owner:SHANDONG UNIV OF SCI & TECH

A stage lighting automatic programming control method and system

PendingCN122640907AStage lightingAudio segmentation
The present application relates to the technical field of stage performance equipment control, and particularly relates to a stage lighting automatic programming control method and system. The method comprises the following steps: obtaining GDTF format stage lamp data, stage building data and performance site data of a target stage scene, associating and encapsulating to generate a stage site data file; analyzing lamp capacity constraints, lamp position space constraints and site coverage constraints to generate a lighting material file and a lamp position layout diagram; writing material screening results and lighting parameter adjustment records to generate target lighting materials; combining target audio files to extract audio features, performing audio segmentation and lighting action segment matching, lamp control parameter mapping and time axis generation to obtain a lighting show execution file; checking DMX addresses, lamp channel fields and lamp parameter boundaries, converting into DMX execution control data and returning to generate a cloud-generated sample. The present application forms a continuous corresponding relationship among lighting materials, audio segments and stage lamp control data.
Owner:GUANGDONG LEYI ELECTRONICS CO LTD

Spatial voiceprint recognition method and system based on spatial acoustic feature extraction

The invention belongs to the technical field of audio processing, and provides a spatial voiceprint recognition method and system based on spatial acoustic feature extraction, and the method comprises the steps: collecting the environment sound of a target space; the collected sound is resampled and human voice is filtered out; segmenting the preprocessed audio into a plurality of segments; extracting spatial acoustic features from the audio clips; removing the abnormal audio clip; evaluating feature importance and selecting an optimal feature subset; generating a spatial voiceprint representation; and space identification and matching are realized. The voiceprint recognition field is expanded, the voiceprint technology is applied to environment space recognition for the first time, and the defect that traditional voiceprint recognition only pays attention to human voice or specific object voice is overcome; abundant acoustic information is provided for space recognition by comprehensively extracting spatial acoustic features; the unique voiceprint identification of different spaces is realized, and a new technical means is provided for the fields of space comparison, judicial case series-parallel connection, anti-fraud call early warning and the like.
Owner:SHANXI POLICE ACAD

Audio enhancement method and electronic device

PCT designated stageWO2026089258A1Speech analysisTime domainAudio segmentation
An audio enhancement method is provided. The method may comprise the steps of: acquiring a plurality of first chunks by dividing input audio; acquiring first features by converting the first chunks into a frequency domain; acquiring, on the basis of the first features, enhanced second features by using a diffusion model, the diffusion model sequentially processing the first features and, when each feature is processed, processing an extended feature including a region partially overlapping that of the preceding feature; acquiring second chunks by converting each of the second features into a time domain; adjusting the signal strength of an overlapping region between adjacent second chunks by applying a window function to the second chunks; and generating enhanced audio by overlap-and-add of the second chunks.
Owner:SAMSUNG ELECTRONICS CO LTD

Multi-speaker speech recognition method and system suitable for end-side equipment

PendingCN121838775ASpeech recognitionAudio segmentationSpeech sound
The invention discloses a multi-speaker speech recognition method and system suitable for an end side device, and the method comprises the steps: collecting multi-speaker audio data, decomposing the multi-speaker audio into two single-speaker audios through a sound channel separation method, carrying out the activity detection of each single-speaker audio, and carrying out the recognition of the single-speaker audios. Segmenting each single speaker audio into a plurality of effective sound segments; performing overlapping judgment on each effective voice segment on the time sequence, performing voice correction on the overlapped effective voice segments, and outputting a confidence coefficient optimization result of character transcription of each effective voice segment; aligning the text transcription result of the audio of each speaker on the time sequence, and optimally calculating the merging confidence value of the text transcription of each effective tone segment after the time sequence alignment; and according to the merging confidence value, carrying out text content splicing on the time sequence on the text transcription result of each effective tone segment, and calculating the optimal time offset in the splicing process of the transcribed text content of each single speaker by adopting a least square method.
Owner:HANGZHOU TIANKUAN TECH