Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

376 results about "Audio segment" patented technology

Audio processing method and device, electronic equipment, computer readable storage medium and computer program product

The invention provides an audio processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a structured list pre-generated for an audio to be edited; analyzing the structured list to obtain metadata of each audio clip; in response to an editing operation for the first audio clip, updating the metadata of the first audio clip to obtain updated metadata; obtaining the first audio clip based on the index of the first audio clip, and decoding the first audio clip to obtain corresponding audio data; creating an audio buffer source node, configuring playing parameters of the audio buffer source node according to the updated metadata, and calling the configured audio buffer source node to play audio data; and in response to receiving the file export request, constructing an offline audio context based on the updated metadata, and generating an audio file through the offline audio context. According to the invention, the audio editing efficiency can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio identification method and device based on marking and backtracking correction

The invention relates to an audio recognition method and device based on marking and backtracking correction, and the method comprises the steps: segmenting a target audio, and enabling adjacent segments after segmentation to have a partial overlapping region; if a slice point exists in a non-mute segment of the target audio and the acoustic feature similarity of a preset number of frames before and after the slice point is smaller than a set similarity threshold value, the slice point is marked as a cut-off risk point, and the cut-off risk point is used for indicating that continuous semantics before and after the slice point has a cut-off risk; after a text corresponding to the audio of each slice is recognized through the speech recognition model, the language confidence degree of an overlapping area where the truncation risk point is located is determined, and the language confidence degree is used for indicating context language logic of the overlapping area; and if it is determined that the language confidence is lower than a set confidence threshold, correcting the recognition texts of the slices before and after the truncation risk point. According to the invention, the accuracy of audio recognition is improved.
Owner:FIBOCOM WIRELESS

Multi-application audio splicing

Systems and methods for multi-application audio splicing may include (1) receiving, from a third-party application installed on a user's device, an audio segment of an audio file selected by the user via an audio-segment-selection interface presented by the third-party application in association with the audio file, (2) loading the audio segment into a post-creation interface of a social media application installed on the user's device and prompting the user to add visual content via the post-creation interface, (3) creating a social media post that includes (i) the audio segment and (ii) visual content added by the user via the post-creation interface, and (4) posting the social media post to a social media consumption channel. Various other methods, systems, and computer-readable media are also disclosed.
Owner:META PLATFORMS INC

Chinese acoustic text depression detection system based on channel attention mechanism fusion

The invention provides a Chinese acoustic text depression detection system based on channel attention mechanism fusion, and the system obtains an original audio signal of a target object through each unit in the system, carries out the preprocessing, obtains audio data, carries out the text transcription, obtains the corresponding text data, carries out the fragment division of the audio data, and carries out the segmentation of the audio data. Corresponding audio clips are obtained, and audio-text data are formed; performing feature extraction on the audio clips and the text data to obtain low-level descriptor features, high-level time-frequency features, audio waveform features and text features; fusing the low-level descriptor features, the high-level time-frequency features and the audio waveform features step by step to determine acoustic features, performing cross-modal fusion on the acoustic features and the text features to obtain fusion features, performing depression detection according to the fusion features, and determining a depression detection result of the target object. According to the scheme, more comprehensive acoustic expression and text features are used, voice text modes are effectively fused, depression detection is achieved, and clinical doctors are assisted in diagnosis.
Owner:SHENZHEN NANSHAN DISTRICT CHRONIC DISEASE CONTROL CENT (SHENZHEN NANSHAN DISTRICT MENTAL HEALTH CENT)

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Dynamic systems and methods for media-aware transport of fragment of content in low-latency, over-the-top, and adaptive bitrate streaming

Low latency, over-the-top (OTT), and / or adaptive bitrate (ABR) content streaming is provided. Content delivery is enhanced by determining if a fragment of a content segment at a content delivery network (CDN) edge node meets a threshold for preferential encapsulation and transport. If met, preferential encapsulation and transport to the client device is provided; otherwise, it defaults to non-preferential encapsulation. The size of the fragment is quantified at a parser of the CDN edge node or an ABR segment encryption system. The ABR system may be connected between a content source and a CDN origin and may include an encryptor that sends CMAF video and audio segment's fragment byte offsets metadata. Also, the CDN edge node may include the ABR system and an encryptor that sends an encrypted CMAF segment's fragment size to a threshold calculator of an HTTP server. Related apparatuses, devices, techniques, and articles are also described.
Owner:ADEIA GUIDES INC

Display control method of word mentioning content, terminal equipment and storage medium

The embodiment of the invention provides a prompt content display control method, terminal equipment and a storage medium. The method comprises the following steps: acquiring a target audio clip; obtaining K text blocks most similar to the semantics of the target audio clip from a plurality of pre-stored text blocks as K candidate text blocks; determining context text blocks of the candidate text blocks, and determining a context consistency coefficient according to the similarity between the target audio clip and each context text block and the similarity between the candidate text block and each context text block; according to the similarity between the target audio clip and the candidate text blocks and the context consistency coefficient, determining the target similarity between the target audio clip and the candidate text blocks, and determining the candidate text block corresponding to the highest target similarity in the K candidate text blocks as a target text block; and switching the displayed prompt content into the manuscript content related to the target text block in the electronic manuscript. And the accuracy of the displayed word mentioning content is improved.
Owner:ZHUHAI MOJIE TECH CO LTD

AI-based short video creative material cutting method and device, equipment and medium

PendingCN121284336ASelective content distributionArtistic renderingAlgorithm
The invention relates to an AI-based short video creative material cutting method and device, equipment and a medium. The method comprises the following steps: performing space-time segmentation on an original video material through a pre-training model to generate a video clip set carrying semantic tags and emotion vectors, and performing beat segmentation on an audio material to generate an audio clip set containing semantic tags, emotion vectors and rhythm features; calling knowledge graph analysis based on the target artistic effect type to generate an inconsistency rule set; calculating a conflict value of video and audio clip combinations through a creative tension scoring function, and screening high-score combinations to construct a candidate set; utilizing a sequence optimization algorithm to generate a fragment sequence for maximizing artistic tension and a cutting time point and a transition instruction of the fragment sequence; and synthesis is realized through an editing engine. Therefore, control over artistic audio-visual conflicts is achieved on the premise that basic continuity is guaranteed, the limitation that an existing AI editing system can only be mechanically matched with sounds and pictures is broken through, and the ability of building deep emotion expression is given to a machine for the first time.
Owner:CHENGDU YISHI INNOVATION INFORMATION TECHNOLOGY CO LTD

Text segmentation voice streaming processing method, system and device, medium and program product

The invention discloses a text segmentation voice streaming processing method, system and device, a medium and a program product, and the method comprises the steps: firstly, carrying out segmentation processing according to a semantic structure of a long text, and obtaining ordered text segments with complete semantics; speech synthesis is carried out on the first text segment to obtain a first audio segment, and SSE connection with the client is established while the first audio segment is synthesized; after the synthesis of the first audio clip is completed, the first audio clip is immediately transmitted to a segmentation buffer queue of the client through SSE connection, and meanwhile, speech synthesis is sequentially performed on subsequent text clips; and after the segmented buffering queue receives the first audio clip, the first audio clip is played immediately, and subsequent audio clips are continuously received and buffered while the first audio clip is played. According to the invention, the coordinated control of the AI speech synthesis progress and the streaming transmission progress is realized, the response speed of AI speech synthesis is obviously improved, and the switching fluency of segmented audios is effectively improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Audio data processing method and device, equipment and storage medium

The invention relates to the field of audio data analysis, in particular to an audio data processing method and device, equipment and a storage medium. The method comprises the following steps: carrying out multi-angle audio continuous acquisition and scene noise suppression on a conference room through a surrounding microphone array, and generating a filtering optimization audio signal; performing three-dimensional time difference positioning calculation according to the filtered and optimized audio signal to obtain accurate sound source positioning information; personalized voiceprint analysis is carried out on the filtered and optimized audio signals, audio distribution modeling is carried out based on accurate sound source positioning information, and a multi-person conference audio field is constructed; performing parallel audio stream separation according to the multi-person conference audio field to generate an intelligent spliced audio segment; and performing deep semantic analysis and semantic logic correction on the intelligent spliced audio segment to generate an audio analysis result. According to the invention, rapid and accurate speaker audio recognition of a multi-person parallel conference is improved.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Automated segmentation and transcription of unlabeled audio speech corpus

ActiveUS12512100B2Speech recognitionTimestampAudio segmentation
A method includes obtaining initial transcription for input natural speech; performing segmentation of initial transcription into text portions, based on punctuation marks in initial transcription; determining segment-level timestamps for text portions based on the input natural speech; performing audio segmentation on input natural speech, by cutting input natural speech based on segment-level timestamps, to obtain audio chunks; generating transcription portions for each of the audio chunks; merging transcription portions to form re-transcription; determining word-level timestamps for re-transcription, by aligning input natural speech against re-transcription; calculating silence time periods, each corresponding to silence between each two adjacent words of input natural speech, based on word-level timestamps; performing a final segmentation on input natural speech and re-transcription, based on silence time periods, to generate final audio segments and corresponding final transcription portions. The final audio segments and corresponding final transcription portions may be included in training dataset for training a model.
Owner:ORACLE INT CORP

Multi-language voice content recognition method and system

The invention relates to a multi-language voice content recognition method and system, and belongs to the technical field of voice signal processing, and the recognition method comprises the steps: collecting original audio stream data, executing the noise reduction filtering processing, and segmenting the original audio stream data into a plurality of audio segments; extracting acoustic feature vectors of the audio clips; inputting the audio clip into the speech recognition model group, and obtaining a text clip and a confidence score; fusing all the text fragments to generate a preliminary recognition text, segmenting the preliminary recognition text into text blocks, and translating the text blocks into a target language text block by block through a streaming translation model in combination with a semantic feedback tag received in real time; executing bidirectional translation verification on the text blocks of which the confidence scores are lower than a set threshold value; inputting the translated text stream and the acoustic feature vector into a semantic analysis model, and fusing an output result to obtain a semantic feedback tag; and constructing a cross-modal association graph through the graph neural network, outputting a key information abstract and triggering an alarm signal. The speech recognition accuracy can be improved while the recognition efficiency is guaranteed, and risk response is carried out.
Owner:BEIJING HIZHI TECH CO LTD

Personalized voice interaction method and related equipment

The embodiment of the invention provides a personalized voice interaction method and related equipment, and belongs to the technical field of intelligent voice interaction. The method comprises the following steps: receiving an object question audio stream by adopting a streaming processing technology in response to a voice interaction service request; performing segmentation operation and preprocessing operation on the object question audio stream to obtain object question audio clips; performing object feature analysis on the object question audio clip to obtain a service object tag; carrying out object group classification on the service object labels, and selecting a corresponding voice recognition model according to a classification result of an object group; sequentially performing voice recognition on the object question audio clips by adopting a voice recognition model to obtain an object question text; and performing text analysis on the object question text to obtain an answer response. According to the scheme, the questioning audio can be accurately recognized, and the user experience of intelligent voice interaction service is improved.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD

System and method for real-time detection of deepfakes

This document describes a system and method for detecting deepfakes in real-time. In particular, the described system and method is configured to detect, in real-time, if a captured image and / or audio segment comprises a deepfake
Owner:ENSIGN INFOSECURITY PTE LTD

Voice operation method of device, apparatus, and electronic device

A voice operation method of a device, comprising: acquiring a video collected by a camera; acquiring voice information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; intercepting a corresponding audio segment in the voice information according to the time interval; acquiring voiceprint information according to the face feature; and performing voice recognition on the audio segment according to the voiceprint information to acquire voice information. The voice operation method of the device does not need to pre-determine a target user and pre-record voiceprint information of the target user, can autonomously extract voiceprint information of multiple users when multiple users use the device at the same time, and can separate voices one by one. Meanwhile, the voiceprint can be autonomously updated and registered. The voice operation method of the device can significantly improve the voice recognition effect of the device in a noisy or multi-user speaking scene.
Owner:HUAWEI TECH CO LTD

A method, apparatus, device, medium and program product for generating an audio file

Embodiments of the present disclosure provide a method, device, medium and program product for generating an audio file. The method comprises: displaying an editing control on a playing page; in response to an interaction operation on the editing control, displaying an audio editing page, the audio editing page comprising audio information of a preset audio file and a generation control; in response to a selection operation on the audio information, determining a reference audio segment in the preset audio file; and in response to an interaction operation on the generation control, generating a target audio segment according to the reference audio segment, and determining a target audio file according to the target audio segment. The technical solution of the embodiments of the present disclosure generates a target audio segment of similar style according to a reference audio segment selected by a user, which can meet the user's expectation of music generation, reduce the difficulty of music creation, and improve the user experience.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Tellurion capable of accurately positioning

The utility model relates to the technical field of globes, and particularly discloses a globe capable of accurate positioning, which comprises a sphere, a support, a base, a laser lamp and a lamp strip, the lamp strip is arranged on the inner surface of the sphere, and the lamp strip surrounds different countries and regions on the sphere; the voice recognition module is installed in the base and used for conducting semantic recognition on voice of a user. The loudspeaker is mounted in the base and is used for outputting a corresponding audio band according to the voice of a user; the main board is installed in the base and used for extracting keywords of user voice and displaying corresponding positions in a highlight mode, controlling the ball body to rotate and display the corresponding positions and synchronously controlling the laser pen to irradiate the position, mentioned currently, of the audio segment according to the output audio segment. According to the utility model, the position corresponding to the voice of a user is displayed by rotating the globe, the related position of the introduction can be displayed through real-time irradiation of the laser pen while the introduction audio is output, the position can be accurately displayed, and the interactivity in the use process of the user is improved.
Owner:赵天奇 +5

Multi-modal data real-time analysis and feedback method and system

The invention provides a multi-modal data real-time analysis and feedback method and system, and the method comprises the steps: collecting a video frame and an audio frame, reading a count value of the same monotonic timer when the collection of the video frame and the audio frame is completed, and generating an audio and video sequence; calculating the behavior popularity of each region in each time slice based on the sequence, and generating a region set for multi-modal event analysis in combination with a preset threshold and a quantity upper limit; scheduling the corresponding video sub-blocks to a visual computing power unit for target positioning and action classification, and executing voice activity detection and keyword category judgment on the time-aligned audio clips at the same time; visual and audio results are fused, and a structured event sequence is generated according to time slice and region compression; and constructing a classroom teaching chain through the sequence, identifying a key event, and finally generating a teaching event description. According to the invention, under the condition of domestic chip combination, cost, time delay, multi-modal consistency and data security controllability are considered, and real-time perception and feedback of teaching behaviors are effectively supported.
Owner:GUANGZHOU KINDLINK INTELLIGENT TECHNOLOGY CO LTD

A method for generating a speaker diary based on audio-visual fusion clustering

ActiveCN119964596BSpeech recognitionSpeech segmentationSpeaker verification
The application discloses a speaker diary generation method based on audio-visual fusion clustering, and aims to solve the problem of "who speaks at what time" in a multi-speaker scene. The method is realized through the following steps: first, an overlapping-aware speech segmentation model is used to segment the audio segment, solving the problem of overlapping speech; second, an advanced speaker verification model is used to extract the speaker voiceprint features of each audio segment and a speaker score matrix generated through face tracking and speaker detection; then, through an audio-video joint clustering method, the number of clusters is optimized according to the audio features and visual information, and K-means clustering is used to complete speaker clustering; the experimental results show that the system adopting the method achieves the lowest diary error rate (DER) on the Ego4D validation set.
Owner:HUNAN UNIV

A to-do task generation method and device

The application discloses a to-do task generation method and device, and relates to the technical field of speech recognition. The method comprises the following steps: segmenting a recording audio signal and extracting features of the recording audio signal to obtain recording audio segment features; grouping the recording audio segment features based on similarities in a similarity matrix, and taking the recording audio segment features in each recording audio segment feature group as a cluster; calculating the similarity between the clusters based on the similarity matrix; clustering based on the similarity between the clusters to obtain a clustering tree; assigning a speaker label based on the clustering tree to obtain a target recording audio signal segment, and then performing speech-to-text processing on the target recording audio signal segment to obtain text content corresponding to the target recording audio signal segment; dynamically constructing a prompt word based on a keyword table, a historical task template and the text content; taking the text generation type as a task type corresponding to a preloaded model; and generating a to-do task by using a task generation model according to the prompt word. In this way, the efficiency of to-do task generation is improved.
Owner:SHAANXI ZHIYUAN INTERNET SOFTWARE CO LTD

A method, apparatus, and electronic device for detecting attacks targeting identity authentication.

This application provides a method, apparatus, and electronic device for detecting attacks on identity authentication, relating to the field of identity authentication technology. The method for detecting attacks on identity authentication includes: performing specified segmentation processing on target audio and target video to obtain multiple data segments, wherein the multiple data segments include various audio segments and various video segments, and the target audio and target video are content recorded during user authentication based on reading specified verification text; extracting speech features from each audio segment to obtain feature vectors for each audio segment, and extracting facial motion features from each video segment to obtain feature vectors for each video segment; determining the deviation result corresponding to each target data segment of a specified media type based on the obtained feature vectors; and determining the detection result based on the deviation result corresponding to each target data segment. Therefore, this solution can improve the accuracy of attack detection.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Video generation method and device, electronic equipment and storage medium

The embodiment of the invention provides a video generation method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, and the specific implementation scheme is as follows: obtaining a first commentary of a reference commentary video; performing text rewriting on the first commentary to obtain a second commentary different from the first commentary; determining picture segments matched with all the commentary sentences in the second commentary in the video to be commented, and generating audio segments of all the commentary sentences in the second commentary; based on the picture segments matched with the commentary sentences and the audio segments of the commentary sentences, derivative commentary video segments of the commentary sentences are generated; and based on the derivative explanation video clip of each explanation sentence, generating a derivative explanation video according to the sequence of each explanation sentence in the second explanation. By applying the scheme provided by the embodiment of the invention, the efficiency of generating the explanation video can be improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Subtitle processing methods and devices

This disclosure relates to a subtitle processing method and apparatus. The method includes: during the editing of a multimedia material segment, obtaining subtitle text corresponding to the audio and timestamp information of audio segments corresponding to each text element in the subtitle text through speech recognition; determining the material segment matching the text element in the multimedia material segment based on the timestamp information of the audio segment corresponding to each text element; and then synthesizing each text element with the matching material segment within the specified time to obtain a target multimedia material with a subtitle text appearing word by word in an animation effect. The solution of this disclosure can achieve a subtitle animation effect where the corresponding text subtitle appears when a certain word is spoken; furthermore, user input commands can automatically generate dynamic subtitles, simplifying user operation and improving user experience.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Dynamic killing method based on cough audio detection and analysis

PendingCN122075751AAvoid Ineffective Disinfectionsave energySpeech analysisLavatory sanitoryEngineeringEmergency medicine
The invention discloses a dynamic sterilization method based on cough audio detection and analysis, which comprises the following steps: acquiring hospital audio, dividing the hospital audio into audio segments, identifying audio frames containing cough audio, and merging the audio frames containing the cough audio to obtain cough audio segments; overlapping coughs are separated, and samples of all the single-sound coughs are obtained and correspond to the individuals to which the single-sound coughs belong; judging whether the individual coughs accidentally or not on the basis of the quantity and the time interval of the single-sound coughs, and further judging whether a hospital environment sterilization early warning program needs to be started or not; and for each environment index of the hospital, presetting a corresponding disinfection trigger threshold, and if it is judged that the hospital environment disinfection early warning program needs to be started and the hospital environment index reaches the corresponding disinfection trigger threshold, starting a corresponding disinfection program to complete disinfection of the hospital environment. According to the invention, 'on-demand disinfection 'can be realized accurately and dynamically in real time, and interference to normal medical activities is reduced to the greatest extent while the transmission path of aerosol in a hospital is blocked.
Owner:CHANGZHOU NO 2 PEOPLES HOSPITAL

Structured video documents

A method includes receiving a content feed that includes audio data corresponding to speech utterances and processing the content feed to generate a semantically-rich, structured document. The structured document includes a transcription of the speech utterances and includes a plurality of words each aligned with a corresponding audio segment of the audio data that indicates a time when the word was recognized in the audio data. During playback of the content feed, the method also includes receiving a query from a user requesting information contained in the content feed and processing, by a large language model, the query and the structured document to generate a response to the query. The response conveys the requested information contained in the content feed. The method also includes providing, for output from a user device associated with the user, the response to the query.
Owner:GOOGLE LLC

Abnormal sleep audio segment identification method, electronic device, and program product

The method for identifying abnormal sleep audio clips, the electronic device and the program product provided by the present disclosure relate to a deep learning technology, and include: obtaining a plurality of initial audio clips, and determining target audio clips meeting a preset sleep state in the plurality of initial audio clips; determining first snoring sound information before each initial audio clip determines the target audio clip, and second snoring sound information after the target audio clip; determining a confidence value of the target audio clip according to the first snoring sound information and the second snoring sound information; and determining whether the target audio clip is an abnormal sleep audio clip according to the confidence value of each target audio clip. In the scheme provided by the present disclosure, the target audio clip that may be an abnormal sleep audio clip can be initially identified in the plurality of initial audio clips, and then whether the target audio clip is indeed an abnormal sleep audio clip is determined by using the snoring sound information before and after the target audio clip, so that the abnormal sleep audio clip can be accurately determined.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Audio processing method and related device

The invention discloses an audio processing method and a related device. The method comprises the following steps: performing voice activity detection on an obtained first audio clip; if the first audio clip is non-voice data, determining a first continuous non-voice duration according to the duration of the first audio clip; if the first continuous non-voice duration is greater than or equal to the first dynamic duration threshold, acquiring a first text set; generating a first intermediate text according to the first text set, and adding the first intermediate text into the first sentence segmentation result set; generating a first to-be-detected text according to each intermediate text included in the first sentence segmentation result set; inputting the first to-be-detected text and the first interaction text into a trained semantic sentence segmentation model to obtain an output result; and if the output result represents that the semantics of the first to-be-detected text is complete, taking the first to-be-detected text as a first sentence segmentation result, and outputting the first sentence segmentation result. According to the invention, the flexibility and accuracy of voice sentence segmentation of the electronic equipment can be improved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

Intelligent ward intercom call system with background noise suppression function

The present application relates to the technical field of audio denoising, and particularly relates to an intercom calling system with background noise suppression function for intelligent ward. The system comprises: an audio data acquisition module, which is used for acquiring initial audio data and performing frame processing to obtain audio frames; an audio segment division module, which is used for segmenting the initial audio data based on amplitude energy difference characteristics between the audio frames to obtain audio information segments; a crosstalk index analysis module, which is used for analyzing superposition conditions such as prompt tones in other intercom devices that may be contained in the audio information segments, analyzing frequency domain information and phase information of the audio information segments, distinguishing crosstalk audio segments and normal audio segments, and calculating crosstalk indexes of each crosstalk audio segment; and an audio denoising module, which is used for performing targeted denoising on the crosstalk audio segments according to the crosstalk indexes, improving the accuracy of background noise suppression, and effectively improving the denoising effect.
Owner:HUNAN SHANGYIKANG MEDICAL TECH CO LTD

User emotion prediction processing method and device, equipment and storage medium

The invention provides a user emotion prediction processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring a face image and audio data of a user; preprocessing the face image and the audio data to obtain an aligned face image and each audio clip; performing visual feature extraction processing on the aligned face image to obtain visual representation features; performing audio feature extraction processing on each audio clip to obtain Mel spectrogram features; preprocessing the visual representation features and the Mel spectrogram features, and determining corresponding final condition features; inputting the final condition features, the current noise value and the time step number into a pre-trained back diffusion process function for processing to obtain a preliminary prediction titer-wake-up value; correcting and updating the preliminary predicted titer-wake-up value to obtain a final titer-wake-up value; and the emotion of the user is predicted according to the final titer-wake-up value, so that the prediction accuracy of the emotion of the user is improved.
Owner:BEIHANG UNIV

Methods and systems to provide a playlist for simultaneous presentation of a plurality of media assets

Systems and methods are described herein for generating a playlist for a simultaneous presentation of a plurality of media assets. The system retrieves a user preference associated with a user profile and receives a selection of a first media asset and a second media asset from the plurality of media assets for presentation on a user device. The system parses the respective audio streams of the first media asset and the second media asset to identify one or more preferred audio segments based on the user preference and generates the playlist of the identified one or more preferred audio segments. Based on a generated audio playlist, the system generates, for presentation on the user device, the video stream for each of the first media asset and the second media asset and the playlist of the identified one or more preferred audio segments.
Owner:ADEIA GUIDES INC