Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

49 results about "Audio segmentation" patented technology

System and method for adaptive audio segmentation for contextual speech signal processing

ActiveUS20250391420A1Speech recognitionAudio segmentationNoise
A system for audio segmentation for context change detection in speech is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system detects a potential context change between a first audio frame and a second audio frame of the speech signal. The system splits the speech signal into a first set of split audio frames based on the potential context change. The system generates a noisy speech signal by modulating a flicker noise signal into the speech signal. The system splits the noisy speech signal into a second set of split audio frames. The system detects a difference between the first and second sets of split audio frames. The system reconfigures a modulation of the speech signal with the flicker noise signal to reduce the difference between the first and second sets of audio frames.
Owner:BANK OF AMERICA CORP

System and method for dynamic audio slicing window selection based on context and speech patterns

ActiveUS20250391404A1Speech recognitionAudio segmentationSpeech patterns
A system for an audio slicing window selection for contextually splitting a speech signal is disclosed. The system identifies a first audio processing software algorithm that is assigned to a user. The system identifies a set of audio processing software algorithms and configures each of them with a respective audio slicing window. The system selects a second audio processing software algorithm, from among the set of audio processing software algorithms. The system selects one of the first and second audio processing software algorithms that is assigned an audio slicing window associated with the context of the speech signal. The system splits the speech signal using the selected audio processing software algorithm. The system determines whether the speech signal is split contextually. In response to determining that the speech signal is not split contextually, the selected audio processing software algorithm and / or the audio slicing window may be updated.
Owner:BANK OF AMERICA CORP

Audio processing method and device, terminal equipment, storage medium and program product

PendingCN121054023ASpeech analysisAudio segmentationTerminal equipment
The invention relates to an audio processing method and apparatus, a terminal device, a storage medium and a program product. The audio processing method comprises the steps of obtaining a first audio; segmenting the first audio into a plurality of second audios; determining playing parameters of each second audio; wherein the playing parameters are different, and sound effects obtained by playing the same second audio are different; and according to the playing parameters of the second audios, playing the corresponding second audios so as to realize playing of the first audios after audio processing. According to the embodiment of the invention, the sound effect after each second audio is played is improved, the situation that part of audio in the first audio is not matched with the playing parameter due to the fact that the first audio is played according to the same playing parameter is reduced, the situation that the sound effect is poor due to the fact that the playing parameter is not matched with the audio is also reduced, and the use experience of a user is improved. Manual operation of a user is not needed, so that operation steps of the user are reduced, the method is more intelligent and convenient, and the use experience of the user is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

OSAHS screening method based on sleep breathing sounds and related equipment

PendingCN120690234ASpeech analysisBiological modelsNoiseAudio segmentation
The invention discloses an OSAHS screening method and related equipment based on sleep breath sounds, and the method comprises the steps: inputting an obtained audio signal into an OSAHS screening model, outputting an AHI value, and then obtaining an OSAHS symptom level; the training steps of the model include: performing audio segmentation and automatic labeling on sleep recording signals, and dividing tags into reliable tags and unreliable tags according to the reliability of the tags; training a noise label classifier by using a label correction algorithm, and iteratively identifying and correcting potential error labels in the unreliable labels in the training process of the noise label classifier; after the tag correction is completed, training an audio clip classifier by adopting the corrected tag; and fitting a linear regression model to predict an AHI value based on a classification result of the whole sleep recording signal. The label correction link is introduced in the model training process, potential wrong labels in unreliable labels are corrected, the accuracy of the model is effectively improved, and the method can be widely applied to the technical field of disease diagnosis equipment.
Owner:SOUTH CHINA UNIV OF TECH

Automated segmentation and transcription of unlabeled audio speech corpus

ActiveUS12512100B2Speech recognitionTimestampAudio segmentation
A method includes obtaining initial transcription for input natural speech; performing segmentation of initial transcription into text portions, based on punctuation marks in initial transcription; determining segment-level timestamps for text portions based on the input natural speech; performing audio segmentation on input natural speech, by cutting input natural speech based on segment-level timestamps, to obtain audio chunks; generating transcription portions for each of the audio chunks; merging transcription portions to form re-transcription; determining word-level timestamps for re-transcription, by aligning input natural speech against re-transcription; calculating silence time periods, each corresponding to silence between each two adjacent words of input natural speech, based on word-level timestamps; performing a final segmentation on input natural speech and re-transcription, based on silence time periods, to generate final audio segments and corresponding final transcription portions. The final audio segments and corresponding final transcription portions may be included in training dataset for training a model.
Owner:ORACLE INT CORP

Psychological state monitoring method and system based on deep learning voice emotion vector

The invention discloses a psychological state monitoring method and system based on a deep learning voice emotion vector, which can accurately measure the psychological state of a pilot in a simulated flight training process, and overcome the problems of strong subjectivity, complex equipment, easy interference, non-deep voice emotion analysis and the like in the existing measurement method. And a powerful guarantee is provided for flight safety. The method comprises the following steps: acquiring voice data of a tested person, and performing audio processing on the acquired voice data to obtain an effective voice audio file set after audio segmentation and de-muting; adopting a pre-trained deep voice emotion calculation model to extract a high-dimensional voice emotion vector of each audio clip; and sequentially inputting the extracted voice emotion vectors into the trained psychological state evaluation model to generate psychological characteristic indexes.
Owner:CHINA EASTERN TECH APPL RES & DEV CENT CO LTD

Audio synchronous playing method and device based on remote WiFi

The invention relates to the technical field of audio synchronous playing, in particular to an audio synchronous playing method and device based on remote WiFi. The method comprises the following steps: acquiring an audio uploading risk value of audio source equipment and an audio receiving risk value of audio receiving equipment based on an audio synchronous playing request, and judging whether to respond to the audio synchronous playing request based on the audio uploading risk value and the audio receiving risk value; according to the audio segment management sequence, performing content segment labeling on target audio data to be uploaded to obtain audio labeling segments, calculating audio feature parameters of each audio labeling segment, forming an information frame by the audio labeling segments, the audio feature parameters corresponding to the audio labeling segments and a preset playing timestamp, and sending the information frame to the server; the audio synchronization control system generates the information frame instruction and sends the information frame instruction to the audio source device through WiFi so as to instruct the audio source device to send the information frame to the audio receiving device, and the accuracy of audio synchronization playing can be improved.
Owner:深圳市迈远科技有限公司

High privacy DSP-based audio anonymization with audio segmentation and randomization

PendingUS20250384891A1Speech analysisAudio segmentationAudio frequency
A method and an electronic device for generating an anonymized audio output are provided. The method, executable by the electronic device, comprises acquiring an audio recording of a speaker; stochastically determining a base pitch value based on at least a first probabilistic function; segmenting the original audio input into a plurality of audio segments, each of the plurality of audio segments being associated with a respective pitch. For each audio segment, the method further comprises generating a pitch adjustment value using a combination of the base pitch value of the segment and a value determined using a second probabilistic function; generating an adjusted audio segment by adjusting the pitch of the audio segment using the pitch adjustment value, the adjusted audio segment having an adjusted pitch that is different from the original pitch; generating the anonymized audio output by combining the adjusted audio segments.
Owner:HUAWEI TECH CO LTD

A Risk Content Identification Method Based on Multimodal Large Model

This invention relates to the field of artificial intelligence technology and provides a risk content identification method based on a multimodal large model. The method includes: identifying forged parts in audio; extracting background noise features and high-frequency features from a target face image and inputting them into an image segmentation model to locate the forged region; segmenting the target face image into blocks, and inputting the resulting local region images and global images into a visual encoder to extract visual features; calculating the attention of text features and global image features to local image features, discarding local image features with low attention; and inputting the outputs of the audio segmentation model and image segmentation model, image features, and questions into a large language model to summarize the risk points. This invention improves the accuracy and robustness of risk identification by integrating multiple data sources and using a multimodal large model for risk identification. It can also effectively counter various fraud methods and solves the problems of existing technologies being unable to handle multimodal data and lacking interpretability.
Owner:HUAZHONG UNIV OF SCI & TECH

Audio segmentation method and device, electronic equipment and storage medium

The application provides an audio segmentation method and device, electronic equipment and storage medium, wherein the method comprises: obtaining audio to be segmented; extracting acoustic features of each frame in the audio to be segmented, and based on the acoustic features of each frame, performing semantic boundary sequence labeling on the audio to be segmented to obtain semantic boundary labeling results of each frame; and based on the semantic boundary labeling results of each frame, performing segmentation on the audio to be segmented. The method, device, electronic equipment and storage medium provided by the application can assist semantic segmentation based on tone and pause information in the acoustic features of each frame, retain complete semantic information of the audio, and avoid punctuation recognition errors, thereby improving the accuracy and reliability of audio segmentation. Furthermore, the method can be applied to a cascaded speech translation system and an end-to-end speech translation system, thereby expanding the application range of audio segmentation.
Owner:HKUST IFLYTEK (SHANGHAI) TECH CO LTD

Audio transcription method and device based on intelligent analysis, equipment and medium

PendingCN121600931ASpeech recognitionSpeech synthesisAudio segmentationEngineering
The invention relates to the technical field of audio recognition, and relates to an audio transcription method and system based on intelligent analysis, and the method comprises the steps: carrying out the file verification and file slicing uploading of a to-be-transcribed file, and obtaining an uploaded audio file; performing audio segmentation on the uploaded audio file, and extracting a voice text after audio segmentation to obtain a primary transcription text; calculating an analysis transcription progress and a real-time transcription progress of the uploaded audio file, performing maximum screening on the analysis transcription progress and the real-time transcription progress to obtain a transcription progress, and updating the primary transcription text into a standard transcription text according to the transcription progress; performing feature activation on the text context feature of the standard transcription text to obtain a text abstract feature; performing feature style decoding on the text abstract features according to a style keyword input by a user to obtain a decoded text abstract; and combining and splicing the standard transcriptional text and the decoded text abstract to obtain an audio transcriptional file. According to the invention, the audio transcription efficiency can be improved.
Owner:SHENZHEN LEXIN SOFTWARE TECH CO LTD

Marine mammal audio retrieval method and system based on biohashing ciphertext

This application discloses a method and system for marine mammal audio retrieval based on biohash ciphertext, relating to the field of audio processing. The method comprises: extracting sub-band CQT energy-entropy ratio features from the original audio, constructing a key-address index table, and constructing a biohash sequence based on the key-address index table; segmenting the original audio using a dual-threshold segmentation method based on short-time energy and spectral distribution variance to obtain short-time audio segments, mapping the short-time audio segments to the key-address index table, and reconstructing the biohash sequence; performing an improved AES encryption on the original audio to obtain ciphertext audio; constructing a hash index table based on the mapping relationship between the reconstructed biohash sequence, key, and ciphertext audio, and uploading the ciphertext audio and hash index table to the cloud for retrieval. This application uses an audio segmentation algorithm to reconstruct the biohash sequence, improving retrieval efficiency, while also using an improved AES encryption algorithm to ensure data security.
Owner:HAINAN UNIV

Computer-implemented method for adaptive decryption of audio stream, and device and storage medium

Disclosed in the present invention are a computer-implemented method for adaptive decryption of an audio stream, and a device and a storage medium. The method comprises: acquiring an MPEG-DASH manifest file and parsing same; initializing a standard DRM interface; performing feature extraction on audio segment information in an identified audio segmentation mode, and on the basis of the extracted features, dynamically adjusting a segment request and processing logic, so as to perform adaptive segment downloading and pre-processing; extracting encryption parameters to identify an encryption flag of a target audio stream platform, and on the basis of the encryption flag, selecting a corresponding decryption algorithm to decrypt a current audio segment; and performing audio frame reconstruction on decrypted data to form an updated audio stream, buffering the updated audio stream into a preset buffer on the basis of an adaptive buffer policy, and outputting a final audio stream, so as to ensure the continuity of the final audio stream and realize efficient parallel processing of decryption and playback, thereby improving the decryption efficiency, and enhancing the user experience especially in high-concurrent scenarios.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Audio segmentation method and device, electronic equipment and storage medium

ActiveCN115719596BSpeech analysisAudio segmentationEngineering
The application provides an audio segmentation method and device, electronic equipment and storage medium, wherein the method comprises: determining a double-channel audio to be segmented; respectively performing mute section labeling on a first channel audio and a second channel audio in the double-channel audio to obtain a mute section in the first channel audio and a mute section in the second channel audio; determining a common mute separation point in the double-channel audio based on the mute section in the first channel audio and the mute section in the second channel audio, and performing segmentation on the first channel audio based on the common mute separation point to obtain a plurality of first segmented audio sections; performing mute section removal on each first segmented audio section to obtain each second segmented audio section; and performing customer audio combination based on a voiceprint feature of each second segmented audio section to obtain customer audio in units of customers, thereby overcoming the defect that a fixed-time segmentation cannot distinguish customers, realizing audio segmentation in units of customers, and providing assistance for different service quality inspection and service evaluation.
Owner:IFLYTEK CO LTD

System and method for adaptive audio segmentation for contextual speech signal processing

ActiveUS12646523B2Speech recognitionAudio segmentationNoise
A system for audio segmentation for context change detection in speech is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system detects a potential context change between a first audio frame and a second audio frame of the speech signal. The system splits the speech signal into a first set of split audio frames based on the potential context change. The system generates a noisy speech signal by modulating a flicker noise signal into the speech signal. The system splits the noisy speech signal into a second set of split audio frames. The system detects a difference between the first and second sets of split audio frames. The system reconfigures a modulation of the speech signal with the flicker noise signal to reduce the difference between the first and second sets of audio frames.
Owner:BANK OF AMERICA CORP

Methods and apparatus to fingerprint an audio signal

Methods, apparatus, systems, and articles of manufacture to fingerprint an audio signal. An example apparatus disclosed herein includes an audio segmenter to divide an audio signal into a plurality of audio segments, a bin normalizer to normalize the second audio segment to thereby create a first normalized audio segment, a subfingerprint generator to generate a first subfingerprint from the first normalized audio segment, the first subfingerprint including a first portion corresponding to a location of an energy extremum in the normalized second audio segment, a portion strength evaluator to determine a likelihood of the first portion to change, and a portion replacer to, in response to determining the likelihood does not satisfy a threshold, replace the first portion with a second portion to thereby generate a second subfingerprint.
Owner:GRACENOTE INC

A system and method for enhancing a call characterization of voice calls

A system and method for enhancing a call characterization of voice calls is disclosed. The method comprises: receiving, at the platform (1), the voice call (30) segmented in audio packets; generating a Call Detail Record "CDR" (12) comprising call parameters; processing, by the Digital Signal Processing "DSP" engine (300), the audio packets, wherein the processing step in turn comprises: computing audio spectral representations; and, accumulating a predetermined number of audio packets (30.1). The method further comprises: evaluating, by the Audio Segmentation "AS" engine (301), an audio type present in the accumulated audio packets; generating temporal marks (31), by the Audio Segmentation "AS" engine (301), indicating which audio segment correspond to each audio type; computing metrics (32), by the Metrics Computation "MC" engine (302), using the accumulated audio packets and the temporal marks; generating an enhanced Call Detail Record "eCDR" (33) by adding the metrics (32) to the "CDR".
Owner:BTS TECHNOLOGY SERVICES SA

Multi-level audio segmentation using deep embeddings

ActiveUS12586552B2Electrophonic musical instrumentsAudio segmentationAudio frequency
Embodiments are disclosed for generating an audio segmentation of an audio sequence using deep embeddings. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving an input including an audio sequence and extracting features for each frame of the audio sequence, where each frame is associated with a beat of the audio sequence. The method may further comprise clustering frames of the audio sequence into one or more clusters based on the extracted features and generating segments of the audio sequence based on the clustered frames, where each segment includes frames of the audio sequence from a same cluster. The method may further comprise constructing a multi-level audio segmentation of the audio sequence and performing a segment fusioning process that merges shorter segments with neighboring segments based on cluster assignments.
Owner:ADOBE INC

Multi-level data modeling-based pathological voice detection model training method

The embodiment of the invention provides a pathological voice detection model training method based on multi-level data modeling. The method comprises the following steps: inputting a dialogue audio into a pathological voice detection model; in the session layer, an encoder is used for encoding continuous audio clips obtained through dialogue audio segmentation, extracting depth features of the audio clips, capturing global context information, obtaining pathological pseudo-labels of the audio clips through depth feature classification, and determining session level loss based on the pathological pseudo-labels; in the fragment layer, modeling is carried out by taking the pathology pseudo labels as supervision signals so as to position audio fragments containing symptom features, depth features of the audio fragments are classified, a predicted pathology classification result is obtained, and fragment-level loss is determined based on the predicted pathology classification result; and performing end-to-end training on the pathological voice detection model based on the loss. According to the embodiment of the invention, the training remarkably improves the performance of the model under a very small amount of annotation data, and achieves the generalization of cross-language / disease species.
Owner:SHANGHAI JIAOTONG UNIV

Surgical safety verification method and system based on artificial intelligence voiceprint recognition

ActiveCN115938371BSpeech analysisAlgorithmAudio segmentation
This invention provides a surgical security verification method and system based on artificial intelligence voiceprint recognition, belonging to the field of artificial intelligence technology. In this invention, voice information is collected from the user to be verified to output corresponding audio of the user to be identified. Audio segmentation is performed on the audio of the user to be identified to form multiple audio segments corresponding to the audio of the user to be identified. Based on a pre-trained voiceprint recognition model, each audio segment of the multiple audio segments is recognized to output multiple audio recognition results corresponding to the multiple audio segments. Then, based on the multiple audio recognition results, an identity verification operation is performed on the user to be verified to output the user identity verification result. Based on the above method, the reliability of surgical security verification can be improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

A speech video generation method based on audio and video structure alignment

The application discloses a speech video generation method based on audio and video structure alignment and belongs to the virtual digital person field.The application comprises an audio segmentation module, an audio conversion module, an audio coding module, a video coding module and a video fusion decoding module.The audio conversion module is used for converting segmented phonemes into a mel-frequency spectrogram which is more in line with the frequency range of human ears according to Fourier transform.In the audio coding process, the frames of the same phonemes are taken as a continuous time module, and the time module is taken as time consistency to constrain the change of the lips, so that the fine-grained control of the lips is realized at the phoneme level through the time consistency constraint.In the video coding, the part region of the multi-pose change face in the input video is set as a mask region, the mask region is used as spatial consistency to accurately control the change amplitude of the lips, the position of the speaker's lips is aligned, the visual artifacts of the video are reduced, the facial details are optimized, and a high-quality speech video with audio-visual synchronization is generated.
Owner:BEIJING INST OF TECH

Audio segmentation method, device, equipment, medium and product

PendingCN121214919ABiological modelsSpeech recognitionAudio segmentationEngineering
The invention discloses an audio segmentation method and device, equipment, a medium and a product, and the method comprises the steps: constructing an initial meaning group segmentation prediction model, obtaining meaning group segmentation audio data, training the initial meaning group segmentation prediction model through employing the meaning group segmentation audio data, obtaining a target meaning group segmentation prediction model, and obtaining a target meaning group segmentation prediction model; and inputting the to-be-predicted audio data into the target meaning group segmentation prediction model to obtain a target segmentation position of the to-be-predicted audio data. According to the invention, the boundary of the audio data can be delimited through voice meaning group segmentation, the boundary delimitation is accurate, and the accuracy of audio segmentation can be improved.
Owner:IFLYTEK CO LTD

Hybrid multi-modal learning heterogeneous multi-codebook quantization video retrieval method and system

The present application relates to the technical field of multi-modal video retrieval, in particular to a heterogeneous multi-codebook quantization video retrieval method and system based on hybrid multi-modal learning, comprising the following steps: S1: importing a video to be retrieved; S2: performing key frame sampling and audio segmentation on the input video, and extracting video features and audio features; S3: introducing low-cost semantic knowledge and designing a prompt-driven multi-modal understanding module to encourage the model to focus on event-related information and enhance the correlation between the video and audio modalities; S4: mapping the heterogeneous modalities to a shared semantic space by designing a contrastive learning loss constructed by fusion video and audio prompts, thereby simultaneously enhancing cross-modal consistency and the transferability of the model; S5: designing a two-stage feature interaction and fusion module to fuse multi-modal features, inputting the fused features into a heterogeneous multi-codebook quantization module, and based on the Euclidean distance, retrieving the most similar database video to the to-be-tested video quantization code and outputting the retrieval result.
Owner:WUHAN INST OF TECH

Language interaction text information calibration method based on humiture large language model Agent

PendingCN121687015ASemantic analysisSpeech recognitionAudio segmentationEngineering
The invention discloses a language interaction text information calibration method based on a humiture large language model Agent, which relates to the field of speech recognition, solves the problem of poor calibration effect of the existing language interaction text information calibration method, and comprises the following steps: S1, acquiring real-time interaction audio and sampling to obtain a sampled audio clip, and storing the sampled audio clip in a database; the method comprises the following steps: S1, sampling an audio sample, carrying out audio analysis to obtain a speech loudness intrusive ratio and a fragment statement recognition coincidence ratio, and obtaining interactive audio collection data, S2, carrying out type division on real-time interactive audio by analyzing the interactive audio collection data, respectively carrying out audio statement segmentation to obtain real-time interactive audio segmentation data, and carrying out audio analysis on the real-time interactive audio segmentation data; and S3, performing vocabulary semantic verification on the text vocabularies according to the real-time interaction audio segmentation data, and performing vocabulary translation and outputting according to the verification result. The method can improve the pertinence and accuracy of the language interaction text information calibration method.
Owner:XINJIANG UYGUR AUTONOMOUS REGION INST OF MEASUREMENT & TESTING

Audio segmentation system performance evaluation method, device, equipment and medium

The invention relates to an audio segmentation system performance evaluation method and device, equipment and a medium. The method comprises the following steps: acquiring a reference segmentation result of a target audio signal and a prediction segmentation result obtained by predicting the target audio signal by an audio segmentation system; the reference segmentation result comprises a plurality of categories of reference audio clips, and the prediction segmentation result comprises a plurality of categories of prediction audio clips; performing alignment processing on the reference audio clip and the predicted audio clip according to categories to obtain audio clip pairs of multiple categories; and performing performance evaluation on the basis of the multiple types of audio clip pairs to obtain a performance evaluation result of the audio segmentation system. By adopting the method, the system performance evaluation accuracy can be improved.
Owner:CHINA TELECOM CLOUD TECH CO LTD

An audio segmentation and classification method based on multi-granularity slicing

This invention discloses an audio segmentation and classification method based on multi-granularity slicing, comprising: preprocessing the audio to obtain an audio file with a uniform sampling rate; slicing the audio file at different time granularities; extracting MFCC features from each slice at different time granularities and then performing image processing; establishing an image classification convolutional neural network model and training and validating it; inputting the processed audio into the image classification convolutional neural network model to obtain the classification result of each slice; and performing aggregation analysis based on the classification results to obtain the segmentation points and segment types of the audio file. This invention, by segmenting long audio files at different time granularities, utilizing an image classification convolutional neural network model for type judgment and classification, and finally performing aggregation analysis, can quickly and accurately find the segmentation points between different types of audio and determine the audio types of the audio segments before and after the segmentation points.
Owner:SICHUAN ZHONGYUN ZHIWANG TECH CO LTD

Cough sound recognition method based on large model parameter efficient fine tuning

The invention discloses a cough sound recognition method based on large model parameter efficient fine tuning, and belongs to the field of artificial intelligence auxiliary diagnosis. The method specifically comprises the following steps: constructing a data set containing cough sound and non-cough sound; segmenting the long audio into audio clips with fixed lengths, removing mute clips and unifying the sampling rates of all the audio clips; converting the audio clip into a two-dimensional time-frequency spectrogram through an audio feature extractor, and taking the two-dimensional time-frequency spectrogram as an input feature of a subsequent model; the method comprises the following steps: taking a pre-trained audio classification large model as a backbone network, and inserting a parameter efficient fine tuning module into an Encoder layer of a standard Transform in the backbone network; main body parameters of the backbone network are frozen, only parameters of a parameter efficient fine tuning module and a classification head of the model are trained, and an optimal model is selected for cough sound recognition; according to the method, a parameter efficient fine tuning technology is used, a parameter efficient fine tuning module is inserted into an audio classification large model, main body parameters of a backbone network are frozen, meanwhile, parameters of a classification head of the parameter efficient fine tuning module and the model are trained, and high-precision cough sound recognition is achieved while the training parameter quantity is reduced.
Owner:ZHENGZHOU UNIV

Audio emotion deep analysis system and method based on artificial intelligence

PendingCN121641075ASpeech analysisAlgorithmAudio segmentation
The invention provides an audio emotion deep analysis system and method based on artificial intelligence. The environment configuration channel stabilization module is used for constructing a voice emotion recognition model and constructing and forming a video voice emotion analysis tool based on deep learning; the video and audio extraction driving module is used for losslessly separating an audio track from a video file; an input video file needing to be processed is selected, and analysis task parameters are set; the audio segmentation and sentiment analysis module is used for segmenting and serializing the audio, adjusting interval values corresponding to sentiment dimension tendency information through linear conversion of original sentiment values, calling a voice sentiment recognition model through an interface, and performing sentiment dimension prediction on each audio segment according to a timestamp; and the result processing and presenting module gathers the timestamp audio clip sentiment analysis prediction information, automatically calculates and adds statistical information, draws the trend of a plurality of sentiment dimensions along with time into a curve graph, and obtains a time-varying trend curve graph of the sentiment dimensions.
Owner:SHANDONG XINYU TECH DEV CO LTD

Artificial intelligence-based automatic generation of lip-synced dubbing

A display device for generating real-time translated lip-synced dubbing comprises: a speaker diarization model for separating input audio into a background audio feed and an individual speaker audio feed; a face detection model for cropping a frame including a face; an active speaker detection model for pairing the individual speaker audio feed with a cropped frame corresponding to an active speaker; a translation model for obtaining translated speech audio; and a lip-sync model for generating lip-synced video frames including a facial movement of the active speaker synchronized with the translated speech audio by using a prediction model, and outputting the translated speech audio while displaying the lip-synced video frames in real time on the basis of the generated video frames, wherein the facial movement of the active speaker is synchronized with the output translated speech audio.
Owner:LG ELECTRONICS INC

Methods and Apparatus to Fingerprint an Audio Signal

Methods, apparatus, systems, and articles of manufacture to fingerprint an audio signal. An example apparatus disclosed herein includes an audio segmenter to divide an audio signal into a plurality of audio segments, a bin normalizer to normalize the second audio segment to thereby create a first normalized audio segment, a subfingerprint generator to generate a first subfingerprint from the first normalized audio segment, the first subfingerprint including a first portion corresponding to a location of an energy extremum in the normalized second audio segment, a portion strength evaluator to determine a likelihood of the first portion to change, and a portion replacer to, in response to determining the likelihood does not satisfy a threshold, replace the first portion with a second portion to thereby generate a second subfingerprint.
Owner:GRACENOTE INC