Audio data processing method and device based on sound large model

CN122551833APending Publication Date: 2026-08-11ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是现有方案中仅关注声学信号与文本符号之间的单向转换识别,缺乏对副语言信息(如语调、停顿、气息)及跨模态语义关联的理解能力;导致获取到的音频片段质量差

Benefits of technology

[0017]本申请实施例可以应用在音频数据筛选场景中,具体可以应用在新闻场景的同期声的音频筛选的过程中,本方案可以对音频数据进行声纹识别,提取出目标人物的音频片段,并利用声音大模型分析语气词确定表达状态和连续情绪向量,并利用语音转文本的文本识别结果进行数据对齐,形成结构化数据。之后,利用结构化数据和用户输入的目标文本,筛选出多段音频片段,并利用音频片段的结构化数据,分析语义相关性、情绪表达强度、表达流畅度和音频质量,以筛选出多段高质量的音频片段。本方案可以对音频片段的语义和音频质量进行分析,还可以结合语气词和上下文等信息,分析用户的连续情绪和表达情况,从而可以提取出高质量的音频片段。并且,本方案是应用在新闻场景,则音频数据包含有记者的声纹和受访人的声纹,本方案可以预先设置记者的声纹,并利用记者声纹确定时间段,以按照记者声纹结束时间来更准确的提取出受访人的音频片段。本方案还可以分析音频片段中的语气词,以学习语气词特征,并通过上下文分析语气词上下文的特征,以结合语气词特征和语气词上下文特征分析表达状态,结合音频片段的语义声学特征和语气词上下文特征分析连续情绪变量,以综合多个维度来对音频片段进行评价,确定更高质量的音频片段。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551833A_ABST
    Figure CN122551833A_ABST
Patent Text Reader

Abstract

This invention discloses an audio data processing method and apparatus based on a large voice model, belonging to the field of computer technology. The method includes: acquiring audio data and extracting audio segments of a target person through voiceprint recognition; processing the audio segments based on the large voice model to determine the segment recognition result, which includes: expression state, continuous emotion vector, and text recognition result; performing structured processing on the segment recognition result to determine structured data, which includes: aligned text content, time information, continuous emotion vector, and modal particle function markers, the modal particle function markers being determined based on the expression state; analyzing semantic relevance, emotional expression intensity, expression fluency, and audio quality based on the target text and structured data, and sorting and filtering multiple audio segments; this solution can filter out high-quality audio segments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to an audio data processing method and apparatus based on a large sound model. Background Technology

[0002] In the production process of television news and short videos, reporters need to extract high-quality synchronous interview footage and key statements from interviewees' answers from lengthy synchronous audio materials, following a pre-written script, using a non-linear editor. However, traditional non-linear editing systems rely primarily on manual review when selecting synchronous audio. Reporters must play the entire video and locate the start and end points of the required statements by listening. This timeline-based approach cannot efficiently select high-quality audio segments.

[0003] In recent years, with the rise of multimodal large-scale models, artificial intelligence has moved from single-modal to multimodal processing, including large-scale audio models. Early news editing mainly focused on text, video, or images, lacking the extraction of deep audio information and having limitations in application scenarios. This patented technology utilizes large-scale audio models to deeply mine and analyze audio data, improving the speed and effectiveness of synchronous sound screening.

[0004] The development of Audio Large Language Models (Audio LLMs) has undergone a paradigm shift from "perception" to "cognition." Early speech technologies focused on Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). Context-Aware Modeling Plus Plus (CAM++) is a deep learning-based speaker verification technology. It is primarily used to extract discriminative speaker embeddings from speech signals to achieve accurate identification and separation of different speakers. In applications such as news audio editing, the CAM++ model can automatically distinguish and separate speech segments from different speakers by modeling the voiceprints of reporters and interviewees, thus significantly improving the efficiency of material selection and editing. However, existing solutions only focus on the one-way conversion between acoustic signals and text symbols, lacking the ability to understand paralinguistic information (such as intonation, pauses, and breath) and cross-modal semantic relationships; resulting in poor quality of the acquired audio segments. Summary of the Invention

[0005] In view of the problems existing in the prior art, the purpose of this invention is to provide an audio data processing method and apparatus based on a large sound model, which can obtain audio segments of a target person through voiceprint recognition, identify the text content, emotion and other information of the audio segments, and analyze semantic relevance, emotional expression intensity, expression fluency and audio quality in order to extract high-quality audio data.

[0006] Firstly, this application provides an audio data processing method based on a large voice model. The method includes: acquiring audio data and extracting audio segments of a target person through voiceprint recognition; processing the audio segments based on the large voice model to determine segment recognition results, the segment recognition results including: expression state and continuous emotion vector determined based on interjections, and text recognition results of speech-to-text conversion, the text recognition results including text content and time information; performing structured processing on the segment recognition results to determine structured data, the structured data including: aligned text content, time information, continuous emotion vector, and interjection function markers, the interjection function markers being determined based on the expression state; acquiring the target text corresponding to the audio data, and analyzing semantic relevance, emotional expression intensity, expression fluency, and audio quality based on the target text and structured data to determine evaluation information of multiple audio segments related to the target text, so as to sort and filter the multiple audio segments.

[0007] Optionally, the step of extracting the audio segment of the target person through voiceprint recognition includes: obtaining a preset first voiceprint vector of the first person, and identifying the first segment information corresponding to the first person in the audio data, wherein the first segment information includes a first start time and a first end time; determining the second voiceprint vector of the second person based on the first segment information, and extracting the audio segment of the second person as the audio segment of the target person.

[0008] Optionally, the large audio model includes a shared encoder, a modal particle detection head, and an emotion decoding head. The process of processing the audio segment to determine the segment recognition result includes: the shared encoder determining the frame-level hidden representation sequence of the audio segment; the modal particle detection head identifying the soft probability of each frame belonging to a modal particle based on the frame-level hidden representation sequence; learning the first feature of the modal particle region and suppressing the second feature of the non-modal particle region based on the soft probability of each frame, determining the semantic modal particle feature of each frame; obtaining the modal particle context feature based on the semantic modal particle feature of each frame, and determining the expression state based on the semantic modal particle feature and the modal particle context feature; the emotion decoding head obtaining the semantic acoustic feature of the audio segment and combining it with the modal particle context feature to determine the continuous emotion vector; and the speech-to-text decoding head performing speech-to-text recognition on the audio segment to determine the text recognition result.

[0009] Optionally, the training steps of the modal particle detection head include: acquiring first training data and first annotation labels, wherein the first annotation labels include modal particle classification results; inputting the first training data into the modal particle detection head, determining a first prediction result, and training the modal particle detection head based on the difference between the first prediction result and the modal particle classification result.

[0010] Optionally, the emotion decoding head acquires the semantic acoustic features of the audio segment and combines them with the context features of the interjections to determine a continuous emotion vector, including: acquiring the semantic acoustic features of the audio segment, the semantic acoustic features including: the frame-level hidden representation sequence of the shared encoder, the frame-level hidden representation sequence including the shared acoustic representation corresponding to each frame; determining the conditional emotion state representation based on the shared acoustic representation and the context features of the interjections, and after performing a nonlinear transformation, determining the continuous emotion vector of each frame to determine the continuous emotion sequence; wherein, the continuous emotion vector includes: a three-dimensional continuous emotion representation of each frame, the three-dimensional continuous emotion representation including semantic consistency strength, acoustic activity modulation, and discourse dominance tendency; semantic consistency strength is used to describe the coherence between the current expression content and the context topic; acoustic activity modulation is used to characterize speech energy changes, prosodic fluctuations, and the degree of expressive engagement; discourse dominance tendency is used to characterize the speaker's information output intensity and expressive initiative at the current moment.

[0011] Optionally, a shared encoder connects a speech-to-text decoder and an emotion decoder; the training steps for the emotion decoder and the speech-to-text decoder include: the speech-to-text decoder and the emotion decoder process the second training data to determine the second prediction result; and calculate the interjection detection loss, speech recognition loss, and emotion regression loss based on the second prediction result to jointly train the emotion decoder and the speech-to-text decoder.

[0012] Optionally, the step of structuring the fragment recognition results to determine structured data includes: obtaining text units corresponding to the text content and determining the time information of the text units; aligning the text content, time information, and continuous emotion vectors according to the time information of the text units; and marking the text units as functional expression units and semantic expression units according to their expression state to form structured data.

[0013] Optionally, the step of analyzing semantic relevance, emotional expression intensity, expression fluency, and audio quality based on the target text and structured data to determine the evaluation information of multiple audio segments related to the target text includes: determining multiple audio segments corresponding to the target text and obtaining structured data corresponding to the text units of the audio segments; extracting features from the target text to obtain a first semantic vector; determining a second semantic vector based on the text content in the structured data; determining semantic relevance based on the first and second semantic vectors; determining emotional expression intensity based on the continuous emotional vector in the structured data; analyzing the density of interjections based on the functional markers of interjections in the structured data to determine expression fluency; analyzing the signal quality features, semantic integrity features, expression effectiveness features, and interaction coordination features of the target audio segments to determine audio quality; and determining the evaluation information of multiple audio segments based on semantic relevance, emotional expression intensity, expression fluency, audio quality, and corresponding weight parameters. The method further includes: establishing a temporal correspondence between the target text and multiple audio segments and obtaining adjustment information of the target text to edit the audio segments.

[0014] Secondly, this application provides an audio data processing device based on a large voice model. The device includes: an audio data acquisition module for acquiring audio data and extracting audio segments of a target person through voiceprint recognition; a recognition result acquisition module for processing the audio segments based on the large voice model to determine the segment recognition result, the segment recognition result including: expression state and continuous emotion vector determined based on interjections, and text recognition result of speech-to-text conversion, the text recognition result including text content and time information; a structured data acquisition module for performing structured processing on the segment recognition result to determine structured data, the structured data including: aligned text content, time information, continuous emotion vector, and interjection function markers, the interjection function markers being determined based on the expression state; and an audio segment filtering module for acquiring the target text corresponding to the audio data, analyzing semantic relevance, emotional expression intensity, expression fluency, and audio quality based on the target text and structured data, determining evaluation information of multiple audio segments related to the target text, and sorting and filtering the multiple audio segments.

[0015] Thirdly, this application provides an electronic device, comprising: a memory and at least one processor; the memory being used to store computer execution instructions; the at least one processor being used to execute the computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in the first aspect.

[0016] Fourthly, this application provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0017] This application can be applied to audio data filtering scenarios, specifically in the process of filtering synchronous sound in news scenarios. This solution can perform voiceprint recognition on audio data to extract audio segments of target individuals. It uses a large voice model to analyze interjections to determine the expression state and continuous emotion vectors, and uses the text recognition results from speech-to-text conversion to align the data and form structured data. Then, using the structured data and the target text input by the user, multiple audio segments are filtered out. The structured data of the audio segments is then used to analyze semantic relevance, emotional expression intensity, fluency, and audio quality to select multiple high-quality audio segments. This solution can analyze the semantics and audio quality of audio segments, and can also combine interjections and contextual information to analyze the user's continuous emotions and expression, thereby extracting high-quality audio segments. Furthermore, since this solution is applied in news scenarios, the audio data package contains the voiceprints of both the reporter and the interviewee. This solution can pre-set the reporter's voiceprint and use the reporter's voiceprint to determine the time period, so as to more accurately extract the interviewee's audio segments according to the end time of the reporter's voiceprint. This solution can also analyze interjections in audio segments to learn their features, and analyze the features of the context of the interjections to combine the features of the interjections and the features of the context of the interjections to analyze the expression state. It can also combine the semantic acoustic features of the audio segments and the features of the context of the interjections to analyze continuous emotional variables, so as to comprehensively evaluate the audio segments from multiple dimensions and identify higher quality audio segments. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0019] Figure 1 This is a flowchart illustrating an embodiment of an audio data processing method based on a large sound model according to this application.

[0020] Figure 2 This is a schematic diagram of the steps of an audio data processing method based on a large sound model according to an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of an audio data processing device based on a large sound model according to an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms, while “a plurality” refers to two or more, and other quantifiers are similarly understood. It should be further understood that the word “comprising” as used in this application’s specification means the presence of the stated feature, integer, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The word “and / or” as used herein describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.

[0024] This application can be applied to audio data filtering scenarios, specifically in the process of filtering synchronous sound in news scenarios. The audio data package contains the voiceprints of the reporter and the interviewee. This solution can pre-set the reporter's voiceprint and use it to determine the interviewee's audio segment. It also uses a voice model to analyze interjections to determine the expression state and continuous emotion vector, and uses the text recognition results of speech-to-text conversion for data alignment to form structured data. Then, using the structured data and the target text input by the user, multiple audio segments are filtered out. The structured data of the audio segments is then used to analyze semantic relevance, emotional expression intensity, fluency, and audio quality to filter out multiple high-quality audio segments. This solution can analyze the semantics and audio quality of audio segments, and can also combine interjections and contextual information to analyze the user's continuous emotions and expression, thereby extracting high-quality audio segments. After determining the high-quality audio segments corresponding to the target text, a time correspondence can be established between the audio segments and the target text, allowing for adjustments to the audio segments by adjusting the target text. For example, the audio segments can be further edited by deleting interjections from the target text.

[0025] The concept behind this solution is as follows: Currently, most news audio editing relies on manual selection. Although pre-written audio scripts exist, the sheer volume of audio files and the excessive length of individual clips make manual selection inefficient and hinders the rapid identification of valuable segments. During editing, reporter questions are rarely used in the final news footage; therefore, it's crucial to prioritize separating the reporter's and interviewer's voices and removing footage of the reporter asking questions. Secondly, during interviews, meaningless elements such as interjections and reduplicated words inevitably get captured by the camera, requiring secondary processing using the aforementioned method. By converting the audio using an ASR model, only the text of the audio needs to be removed to select the corresponding segments. This method improves editing efficiency while maintaining selection accuracy.

[0026] This solution provides a rapid selection method for news synchronous sound based on a large sound model. It introduces data preprocessing methods to filter and screen the original synchronous sound materials, extracting key personnel and semantic information through a large model. Then, based on dimensions such as text semantics and voice emotion, the most effective synchronous sound segments are further selected. Finally, the sound quality of the returned results is evaluated using the large sound model, resulting in a comprehensive score for each candidate synchronous sound segment. This further improves the selection criteria for synchronous sound, allowing for better integration into application scenarios and solving the problem of low efficiency in manual selection.

[0027] like Figure 1 As shown in the embodiments of this application, the method includes the following steps: First, data preprocessing: the original interview audio is input and voiceprint recognition is used to extract the interviewer's voice segments; then, a unified acoustic semantic representation is extracted by the shared speech encoder (or shared encoder) in the large sound model. Subsequently, three tasks are executed in parallel within the model: semantic modeling of interjections, conditional emotion decoding, and ASR speech recognition. The interjection module generates conditional features related to the expression state, the emotion module outputs continuous emotion vectors, and the ASR module outputs text transcription results with time information. Based on this, a unified structured transcription result is constructed, which organizes the text content, timestamp, continuous emotion vector, and interjection function markers into the same temporal structure to form a structured transcription. Finally, combined with the reporter's press release, a synchronous sound value evaluation model is established through four dimensions: semantic relevance, emotional expression intensity, expression fluency, and audio quality. All candidate segments are comprehensively scored and ranked, and the top 5 synchronous sound segments with the highest value are output for news broadcasting or subsequent editing.

[0028] The input for data preprocessing is sampled audio from reporters and interviewees, which is defined in this scheme as... Where xn represents the speaker, t is the time index, and T1 is the number of sampling points. The audio file has a sampling rate of 16kHz, is mono, has a sampling duration of 5 seconds, and is in WAV format to minimize the loss of high-frequency features. The CAM++ model generates a 192-dimensional normalized speaker embedding vector Ei for each audio segment; voiceprint recognition is performed on each synchronous sound material, with the input being a set of synchronous sound materials Mn, Mn={M1,…,Mi,…,MN}, where Mi represents the i-th synchronous sound material in the set. Finally, the cosine similarity is calculated to determine whether the pre-recorded voiceprint vector and the vector in the material belong to the same person. The final output synchronous sound segments are the reporter's segment and the segment of the interviewer (or interviewee) marked. The following will only process the interviewer's segment.

[0029] In one embodiment, audio is extracted from interview segments obtained through data preprocessing, and after encoding using a large sound model, a frame-level hidden representation sequence is obtained. Where T represents the total number of frames. Let represent the hidden representation corresponding to frame t. To achieve refined modeling of interjections, in the hidden representation sequence... A lightweight modal particle detection head is constructed on top of this. For frame t, the soft probability that it belongs to the modal particle region is defined as... ,in This represents the confidence level that the t-th frame belongs to the modal particle region.

[0030] Unlike traditional binary classification methods for interjections, this invention does not directly output discrete labels. Instead, it utilizes the soft probability to construct semantic interjection features. Within the large sound model, for frame t, its semantic interjection features are defined as follows: ,in Indicating semantic modal features, The gradient stopping operator is identified. Through the above construction, the features corresponding to the modal particle region can retain the complete gradient to participate in subsequent optimization, while the gradient propagation of non-modal particle regions is suppressed, thus enabling the model to focus on learning the expressive functional information carried by the modal particles.

[0031] In one embodiment, to capture local modal particle context information, within a time window Built-in contextual feature representation of modal particles: ,in Indicates the length of the context window; This represents the contextual features of the modal particles corresponding to frame t. Among these, the modal particle features... It is not used for emotion category determination, but rather to characterize changes in the expressive state of speakers in the vicinity of the current moment. The location of modal particles usually corresponds to expressive behaviors such as thinking, pausing, self-correction, emphasis, or hesitation, and therefore can serve as auxiliary conditional information reflecting changes in expressive state. The emotion decoding head combines semantic acoustic features output by the shared encoder. Contextual features of modal particles Joint modeling is used to learn the mapping relationship between the two and continuous emotion representation from the training data, rather than directly inferring the emotion category using preset rules or a dictionary of interjections.

[0032] The contextual features of modal particles are not used as the final classification result output, but rather participate as conditional variables in the subsequent emotion decoding and expression state modeling process, realizing the transformation of modal particles from discrete labels to continuous semantic representations. During the training phase, binary classification supervision is used to optimize the modal particle detection head. ;in This indicates that the labels were manually marked. This indicates that the target frame belongs to the modal particle region; This indicates that the target frame belongs to the non-interjection region. Through the above modeling process, interjections are encoded as continuous learnable feature representations, which, together with subsequent emotion representations, form a unified semantic structure, providing a basic representation for the value assessment of synchronous interview sounds.

[0033] After obtaining the semantic interjection feature sequence, this scheme further constructs a conditional emotion decoding module within the large-scale voice model. Unlike traditional emotion recognition methods that directly classify based on acoustic features, this invention injects the contextual features of interjections as conditional variables into the emotion decoding process, making emotion state prediction simultaneously dependent on speech content, prosodic patterns, and the expressive function of interjections. The hidden representation sequence output by the shared speech encoder is defined as... in , representing the shared acoustic representation corresponding to frame t. Utilizing the contextual features of preceding modal particles... The emotion decoding head uses a shared acoustic representation ( ) and contextual features of modal particles ( Using this as a joint input, construct a conditional emotion state representation: ;in, Represents the acoustic feature mapping matrix; Represents the conditional mapping matrix of modal particles; This indicates the bias term.

[0034] The latent emotional state is obtained after nonlinear transformation. Further mapping to a continuous emotion vector: ,in: This represents the three-dimensional continuous emotion representation corresponding to frame t. This represents the dimension of semantic consistency strength; Indicates the acoustic activity modulation dimension; The discourse dominance dimension is indicated by semantic consistency strength, which describes the coherence between the current content and the contextual topic; acoustic activity modulation, which characterizes changes in speech energy, prosodic fluctuations, and the degree of expressive engagement; and discourse dominance, which characterizes the intensity of the speaker's information output and expressive initiative at the current moment. For the entire interview transcript, a continuous emotion sequence was obtained: During the training phase, continuous value regression was used to optimize the emotion decoding head.

[0035] To achieve joint learning of the speech-to-text transcription and emotion modeling tasks, a shared encoder is connected to both the ASR decoder head and the emotion decoder head. The ASR decoder head outputs the text transcription result. The emotion decoding head outputs a continuous emotion sequence as follows: The joint training objective is defined as follows: ;in This indicates the loss in modal particle detection; Indicates speech recognition loss; This indicates a loss in emotional regression; The corresponding loss weights are represented. Through a joint optimization process, the tasks of interjection detection, speech recognition, and emotion modeling share a unified acoustic representation space, enabling the model to learn a unified temporal representation that simultaneously contains semantic, expressive, and emotional information, providing a foundation for the subsequent construction of structured transcription results.

[0036] After completing speech recognition decoding and conditional emotion decoding, this scheme further constructs a unified structured transcription result. Unlike traditional ASR systems that only output text sequences, this invention maps text content, time alignment information, continuous emotion representation, and modal particle function information into the same temporal structure, forming a multi-dimensional structured representation that can be directly used for news synchronous sound retrieval and value assessment. Regarding the recognition results output by the ASR decoding head: ,in This represents the text content corresponding to the nth recognition unit. The corresponding time boundary is obtained through the forced alignment module. Based on temporal boundaries, frame-level continuous sentiment vectors are aggregated to the text unit level: ; indicates the semantic consistency strength, acoustic activity modulation, and discourse dominance tendency corresponding to the text unit.

[0037] In one embodiment, the modal function tag of the corresponding text unit is calculated based on the modal detection result: ,when If the condition is met, the text unit is marked as a functional expression unit; otherwise, it is marked as a semantic expression unit.

[0038] Finally, a unified structured transcription result is constructed: ,in This corresponds to the text content, start time, end time, continuous emotion vector, and modal particle function markers; any element in the unified structured transcription result can be accurately located via the timeline. Therefore, each candidate segment simultaneously possesses: a text content sequence. Continuous emotion sequence Functional sequence of modal particles And complete time boundary information.

[0039] The unified structured transcription result (structured data) serves as the sole input data structure for both the news release semantic retrieval module and the synchronous sound value assessment module, enabling unified representation and joint calculation of semantic, emotional, expressive, and temporal information. After obtaining the unified structured transcription result, this solution further constructs a synchronous sound value assessment module to achieve automatic retrieval, value quantification, and candidate ranking of synchronous sound driven by news releases. Unlike existing methods that rely on single text matching or emotion classification for filtering, this invention directly utilizes the unified structured transcription result output by the large-scale sound model as input, jointly assessing the communicative value of synchronous sound through four dimensions: semantic relevance, emotional expression intensity, expressive fluency, and audio quality.

[0040] First, obtain the completed press release text from the reporter, and then construct a semantic representation of the press release using a semantic encoding model. ,in This represents the global semantic vector corresponding to the press release. For any candidate interval in the structured transcription result... The corresponding text sequence is defined as follows: Construct candidate interval semantic vectors Finally, the semantic relevance between the candidate intervals and the press release is calculated. The higher the semantic relevance, the more consistent the interview content is with the news topic.

[0041] Calculate the sentiment intensity of the candidate interval. For a continuous sentiment vector within the interval: ,in Define the emotion expression score as shown in Formula 1:

[0042] Formula 1

[0043] in, The emotion expression score is used to measure the interviewee's level of information input, expressive power, and persuasiveness within the candidate range. A fluency score is further calculated. For the functional sequence of interjections corresponding to the candidate range: The density of modal particles is defined as follows: Fluency score is defined as follows: The lower the density of interjections, the smoother the expression and the higher the score. To ensure that the candidate synchronous sound meets broadcast requirements, this invention further introduces an audio quality assessment module. For the audio corresponding to the candidate interval... Extracting signal quality features from the hidden representation of a large sound model Semantic integrity features , expression of effectiveness characteristics and interactive coordination features Construct an audio quality score: ;in, Finally, the comprehensive value function of synchronous sound is constructed: ;in, , representing the weight parameters for semantic relevance, emotional expression intensity, fluency, and audio quality, respectively. A value score is calculated for each candidate interval, and the intervals are sorted from highest to lowest score. The top 5 candidate synchronous sound segments with the highest value are then obtained: .

[0044] This invention achieves a unified computational framework for news release semantic understanding, sentiment expression modeling, fluency assessment, and audio quality analysis by directly utilizing the unified structured transcription results output by the large sound model. It can automatically select the best synchronous sound without additional manual rules or independent sentiment analysis systems, thereby significantly improving the efficiency and quality of synchronous sound screening in the news production process.

[0045] Based on the above embodiments, this application also provides an audio data processing method based on a large sound model, such as... Figure 2 As shown, the method includes:

[0046] Step 102: Acquire audio data and extract audio segments of the target person using voiceprint recognition. Step 104: Process the audio segments based on a large voice model to determine the segment recognition results. These results include: the expression state and continuous emotion vector determined based on interjections, and the text recognition results from speech-to-text conversion. The text recognition results include text content and time information. Step 106: Perform structured processing on the segment recognition results to determine structured data. This structured data includes: aligned text content, time information, continuous emotion vector, and interjection function markers. The interjection function markers are determined based on the expression state. Step 108: Acquire the target text corresponding to the audio data. Based on the target text and structured data, analyze semantic relevance, emotional expression intensity, expression fluency, and audio quality to determine the evaluation information of multiple audio segments related to the target text, and then sort and filter these audio segments.

[0047] The data processing flow in this embodiment is similar to that in the above embodiments. For specific implementation details, please refer to the specific implementation details in the above embodiments. These details will not be repeated here.

[0048] This application can be applied to audio data filtering scenarios. This solution performs voiceprint recognition on audio data to extract audio segments of the target person. It then uses a large voice model to analyze interjections to determine the expression state and continuous emotion vectors. Finally, it uses the text recognition results from speech-to-text conversion to align the data and form structured data. Next, using the structured data and the target text input by the user, multiple audio segments are filtered out. The structured data of these audio segments is then used to analyze semantic relevance, emotional expression intensity, fluency, and audio quality to select multiple high-quality audio segments. This solution can analyze the semantics and audio quality of audio segments, and can also combine interjections and contextual information to analyze the user's continuous emotions and expression, thereby extracting high-quality audio segments.

[0049] This solution can be applied to news scenarios. The audio data package contains the voiceprints of both the reporter and the interviewee. This solution can pre-set the reporter's voiceprint and use it to determine time periods, allowing for more accurate extraction of the interviewee's audio segment based on the reporter's voiceprint end time. Specifically, as an optional embodiment, extracting the target person's audio segment through voiceprint recognition includes: obtaining a preset first voiceprint vector for the first person and identifying the first segment information corresponding to the first person in the audio data. The first segment information includes a first start time and a first end time. Based on the first segment information, determining the second voiceprint vector for the second person and extracting the second person's audio segment as the target person's audio segment. Due to the complex environment and the potential presence of noise from other people, this solution can use the voiceprint of the main speaker (which can be categorized by audio length and quality) within a preset duration after the first end time as the target person's voiceprint.

[0050] This scheme can also analyze interjections in audio segments to learn their features, and analyze the context of the interjections to analyze their features and contextual features. It combines interjection features and contextual features to analyze the expression state, and combines the semantic acoustic features of the audio segment and contextual features to analyze continuous emotional variables. This allows for a comprehensive evaluation of the audio segment from multiple dimensions, identifying higher-quality audio segments. Specifically, as an optional embodiment, the large sound model includes a shared encoder, an interjection detection head, and an emotion decoding head. The processing of the audio segment to determine the segment recognition result includes: the shared encoder determining the frame-level hidden features of the audio segment. The sequence is represented by a hidden representation head; the interjection detection head identifies the soft probability of each frame belonging to an interjection based on the frame-level hidden representation sequence; based on the soft probability of each frame, it learns the first feature of the interjection region and suppresses the second feature of the non-interjection region to determine the semantic interjection feature of each frame; based on the semantic interjection feature of each frame, it obtains the interjection context feature, and determines the expression state based on the semantic interjection feature and the interjection context feature; the emotion decoding head obtains the semantic acoustic feature of the audio segment and combines it with the interjection context feature to determine the continuous emotion vector; the speech-to-text decoding head performs speech-to-text recognition on the audio segment and determines the text recognition result.

[0051] This scheme utilizes soft probability to analyze the features of modal particles. During the training process of the modal particle detection head, the training can be completed using the binary classification results of modal particles (belonging to modal particles or not belonging to modal particles). Specifically, as an optional embodiment, the training steps of the modal particle detection head include: acquiring first training data and first annotation labels, wherein the first annotation labels include the modal particle classification results; inputting the first training data into the modal particle detection head, determining the first prediction result, and training the modal particle detection head based on the difference between the first prediction result and the modal particle classification result.

[0052] The emotions in the target person's audio are continuous. Therefore, when analyzing emotions, the continuous changes in emotions can be considered. Accordingly, the emotion decoding head can process semantic interjection features, interjection context features, and semantic acoustic features to form a continuous emotion sequence for audio evaluation. Specifically, as an optional embodiment, the emotion decoding head obtains the semantic acoustic features of the audio segment and combines them with the interjection context features to determine the continuous emotion vector. This includes: obtaining the semantic acoustic features of the audio segment, which include: the frame-level hidden representation sequence of the shared encoder, which includes the shared acoustic representation corresponding to each frame. Based on shared acoustic representations and contextual features of modal particles, conditional emotion state representations are determined. After nonlinear transformation, continuous emotion vectors for each frame are determined to establish a continuous emotion sequence. These continuous emotion vectors include: three-dimensional continuous emotion representations for each frame, comprising semantic consistency strength, acoustic activity modulation, and discourse dominance tendency. Semantic consistency strength describes the coherence between the current expression and the contextual topic; acoustic activity modulation characterizes changes in speech energy, prosodic fluctuations, and the degree of expressive engagement; and discourse dominance tendency characterizes the speaker's information output intensity and expressive initiative at the current moment.

[0053] During the training of the large-scale audio model, this scheme can comprehensively analyze the outputs of the speech-to-text decoder, the emotion decoder, and the shared encoder to determine the interjection detection loss, speech recognition loss, and emotion regression loss, thereby determining the joint loss. The emotion decoder and the speech-to-text decoder are then jointly trained. Specifically, as an optional embodiment, the shared encoder connects the speech-to-text decoder and the emotion decoder. The training steps for the emotion decoder and the speech-to-text decoder include: processing the second training data to determine the second prediction result; calculating the interjection detection loss, speech recognition loss, and emotion regression loss based on the second prediction result to jointly train the emotion decoder and the speech-to-text decoder. Loss analysis is performed based on the second prediction result and the second annotation result, whereby the second annotation result includes the annotated interjection detection result, speech-to-text recognition result, and continuous emotion vector.

[0054] To facilitate audio data processing, this solution can use text units as the basic alignment unit to align multiple data sets. Specifically, as an optional embodiment, the structuring processing of the segment recognition results to determine structured data includes: obtaining text units corresponding to the text content and determining the time information of the text units; aligning the text content, time information, and continuous emotion vectors according to the time information of the text units; and marking the text units as functional expression units and semantic expression units based on their expression state to form structured data. Each text unit corresponds to a time, and the continuous emotion vectors and modal particle functional tags are aligned based on the time information. The modal particle functional tags include functional expression unit tags and semantic expression unit tags.

[0055] After determining the structured data, the input target text can be used to select multiple corresponding structured data for evaluation. Specifically, as an optional embodiment, the analysis of semantic relevance, emotional expression intensity, fluency, and audio quality based on the target text and structured data to determine the evaluation information of multiple audio segments related to the target text includes: determining multiple audio segments corresponding to the target text and obtaining the structured data corresponding to the text units of the audio segments; extracting features from the target text to obtain a first semantic vector; determining a second semantic vector based on the text content in the structured data; and determining the second semantic vector based on the first semantic vector and the third semantic vector. The method uses two semantic vectors to determine semantic relevance; it uses continuous emotion vectors in structured data to determine the intensity of emotion expression; it analyzes the density of modal particles based on the functional markers of modal particles in structured data to determine the fluency of expression; it analyzes the signal quality characteristics, semantic integrity characteristics, expression effectiveness characteristics, and interaction coordination characteristics of the target audio segment to determine the audio quality; and it determines the evaluation information of multiple audio segments based on semantic relevance, emotion expression intensity, fluency of expression, audio quality, and corresponding weight parameters. The method also includes establishing a temporal correspondence between the target text and multiple audio segments, and obtaining adjustment information for the target text to edit the audio segments. Furthermore, after determining multiple audio segments (synchronous sound) corresponding to the target text, this solution can delete one or more text units in the target text to remove the corresponding audio, facilitating editing.

[0056] Based on the above embodiments, this application also provides an audio data processing device based on a large sound model, such as... Figure 3As shown, the device includes: an audio data acquisition module 202, used to acquire audio data and extract audio segments of the target person through voiceprint recognition; a recognition result acquisition module 204, used to process the audio segments based on a large voice model to determine the segment recognition result, which includes: the expression state and continuous emotion vector determined based on interjections, and the text recognition result of speech-to-text conversion, which includes text content and time information; a structured data acquisition module 206, used to perform structured processing on the segment recognition result to determine structured data, which includes: aligned text content, time information, continuous emotion vector, and interjection function markers, where the interjection function markers are determined based on the expression state; and an audio segment filtering module 208, used to acquire the target text corresponding to the audio data, and based on the target text and structured data, analyze semantic relevance, emotional expression intensity, expression fluency, and audio quality to determine the evaluation information of multiple audio segments related to the target text, so as to sort and filter the multiple audio segments.

[0057] The data processing flow in this embodiment is similar to that in the above method embodiment. For specific implementation details, please refer to the specific implementation details in the above method embodiment. These details will not be repeated here.

[0058] This application can be applied to audio data filtering scenarios. This solution performs voiceprint recognition on audio data to extract audio segments of the target person. It then uses a large voice model to analyze interjections to determine the expression state and continuous emotion vectors. Finally, it uses the text recognition results from speech-to-text conversion to align the data and form structured data. Next, using the structured data and the target text input by the user, multiple audio segments are filtered out. The structured data of these audio segments is then used to analyze semantic relevance, emotional expression intensity, fluency, and audio quality to select multiple high-quality audio segments. This solution can analyze the semantics and audio quality of audio segments, and can also combine interjections and contextual information to analyze the user's continuous emotions and expression, thereby extracting high-quality audio segments.

[0059] Based on the above embodiments, this application also provides an electronic device, including: a memory and at least one processor; the memory is used to store computer execution instructions; the at least one processor is used to execute the computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in the above embodiments.

[0060] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described data processing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0061] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0063] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0064] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.

[0065] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0066] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0067] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0068] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0069] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0070] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. An audio data processing method based on a large sound model, characterized in that, The method includes: Acquire audio data and extract audio segments of the target person through voiceprint recognition; The audio segment is processed based on the large sound model to determine the segment recognition result. The segment recognition result includes: the expression state and continuous emotion vector determined based on the interjection, and the text recognition result of speech-to-text conversion. The text recognition result includes text content and time information. The fragment recognition results are processed in a structured manner to determine the structured data, which includes: aligned text content, time information, continuous emotion vectors and modal function tags, with the modal function tags determined according to the expression state; Obtain the target text corresponding to the audio data. Based on the target text and structured data, analyze semantic relevance, emotional expression intensity, expression fluency, and audio quality to determine the evaluation information of multiple audio segments related to the target text, so as to sort and filter the multiple audio segments.

2. The method according to claim 1, characterized in that, The method of extracting audio segments of the target person through voiceprint recognition includes: Obtain the first voiceprint vector preset by the first person and identify the first segment information corresponding to the first person in the audio data. The first segment information includes the first start time and the first end time. Based on the information in the first segment, the second voiceprint vector of the second person is determined, and the audio segment of the second person is extracted as the audio segment of the target person.

3. The method according to claim 1, characterized in that, The large voice model includes a shared encoder, a modal particle detector, and an emotion decoder; The process of processing the audio segment to determine the segment recognition result includes: A shared encoder determines the sequence of frame-level hidden representations for an audio segment; The modal particle detection head identifies the soft probability of each frame belonging to a modal particle based on the frame-level hidden representation sequence; Based on the soft probabilities of each frame, the first feature of the modal particle region is learned, and the second feature of the non-modal particle region is suppressed to determine the semantic modal particle features of each frame. Based on the semantic modal particle features of each frame, the context features of the modal particles are obtained, and the expression state is determined based on the semantic modal particle features and the context features of the modal particles. The emotion decoder acquires the semantic acoustic features of audio segments and combines them with the contextual features of interjections to determine continuous emotion vectors; The speech-to-text decoder performs speech-to-text recognition on audio segments and determines the text recognition result.

4. The method according to claim 3, characterized in that, The training steps for the modal particle detection head include: Obtain the first training data and the first annotation label, wherein the first annotation label includes the classification result of the modal particle; The first training data is input into the modal particle detection head to determine the first prediction result. The modal particle detection head is then trained based on the difference between the first prediction result and the modal particle classification result.

5. The method according to claim 3, characterized in that, The emotion decoding head acquires the semantic acoustic features of the audio segment and combines them with the contextual features of interjections to determine a continuous emotion vector, including: Obtain the semantic acoustic features of the audio segment. The semantic acoustic features include: the frame-level hidden representation sequence of the shared encoder, which includes the shared acoustic representation corresponding to each frame. Based on shared acoustic representations and contextual features of modal particles, conditional emotion state representations are determined. After nonlinear transformation, continuous emotion vectors for each frame are determined to establish a continuous emotion sequence. These continuous emotion vectors include: three-dimensional continuous emotion representations for each frame, comprising semantic consistency strength, acoustic activity modulation, and discourse dominance tendency. Semantic consistency strength describes the coherence between the current expression and the contextual topic; acoustic activity modulation characterizes changes in speech energy, prosodic fluctuations, and the degree of expressive engagement; and discourse dominance tendency characterizes the speaker's information output intensity and expressive initiative at the current moment.

6. The method according to claim 5, characterized in that, The shared encoder connects a speech-to-text decoder and an emotion decoder. The training steps for the emotion decoder and speech-to-text decoder include: The speech-to-text decoder and the emotion decoder process the second training data to determine the second prediction result; Based on the second prediction results, the loss for interjection detection, speech recognition, and emotion regression are calculated to jointly train the emotion decoder and the speech-to-text decoder.

7. The method according to claim 1, characterized in that, The step of structuring the fragment recognition results to determine structured data includes: Obtain the text unit corresponding to the text content and determine the time information of the text unit; The text content, time information, and continuous emotion vectors are aligned according to the time information of the text units; and the text units are marked as functional expression units and semantic expression units according to their expression status to form structured data.

8. The method according to claim 7, characterized in that, Based on the target text and structured data, the semantic relevance, emotional expression intensity, fluency, and audio quality are analyzed to determine the evaluation information of multiple audio segments related to the target text, including: Identify the multiple audio segments corresponding to the target text and obtain the structured data corresponding to the text units of the audio segments; Feature extraction is performed on the target text to obtain a first semantic vector; a second semantic vector is determined based on the text content in the structured data; semantic relevance is determined based on the first and second semantic vectors. Determine the intensity of emotion expression based on continuous emotion vectors in structured data; Based on the functional markers of modal particles in structured data, the density of modal particles is analyzed to determine the fluency of expression; The signal quality characteristics, semantic integrity characteristics, expressive effectiveness characteristics, and interactive coordination characteristics of the target audio segment are analyzed to determine the audio quality. Evaluation information for multiple audio segments is determined based on semantic relevance, intensity of emotional expression, fluency of expression, audio quality, and corresponding weight parameters. The method further includes: Establish a time correspondence between the target text and multiple audio segments, and obtain adjustment information for the target text in order to edit the audio segments.

9. An audio data processing device based on a large sound model, characterized in that, The device includes: The audio data acquisition module is used to acquire audio data and extract audio segments of the target person through voiceprint recognition. The recognition result acquisition module is used to process audio segments based on a large sound model and determine the segment recognition result. The segment recognition result includes: the expression state and continuous emotion vector determined based on interjections, and the text recognition result of speech-to-text conversion. The text recognition result includes text content and time information. The structured data acquisition module is used to perform structured processing on the fragment recognition results and determine the structured data. The structured data includes: aligned text content, time information, continuous emotion vectors and modal function tags. The modal function tags are determined according to the expression state. The audio segment filtering module is used to obtain the target text corresponding to the audio data. Based on the target text and structured data, it analyzes semantic relevance, emotional expression intensity, expression fluency, and audio quality to determine the evaluation information of multiple audio segments related to the target text, so as to sort and filter the multiple audio segments.

10. An electronic device, characterized in that, include: Memory and at least one processor; The memory is used to store computer-executed instructions; The at least one processor is configured to execute computer execution instructions stored in the memory, such that the at least one processor performs the method as described in any one of claims 1-8.