A method and system for attribution and structuring of multimedia segments for live streaming control events

CN122845857APending Publication Date: 2026-09-29CSC FINANCIAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611048396.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

第一,现有切片技术与直播流控制记录完全脱节

Benefits of technology

[0022]借由上述技术方案,本申请提供的面向直播流控制事件的多媒体片段归因与结构化方法及系统,该方法通过实时监测观众反馈指标的异常波动,识别包含“观众反馈波动→流控制干预/主播行为调整→观众反馈恢复”完整闭环的直播流控制事件区间,从中提取具有叙事完整性和结构化价值的多媒体片段,解决现有切片技术与直播流控制记录脱节的问题;通过构建因果推断模型,量化不同主播行为特征和不同干预策略对观众反馈恢复的平均处理效应,为每个多媒体片段赋予“为什么有效”的因果归因元数据,解决切片内容缺少策略有效性归因的问题;以及,将直播流控制事件切片为带归因标签的多媒体知识片段,明确标注触发波动的情境特征、采用的干预策略、以及策略有效性证据,形成可检索、可对比的结构化多媒体数据库,解决直播经验无法沉淀为结构化多媒体资产的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845857A_ABST
    Figure CN122845857A_ABST
Patent Text Reader

Abstract

The application provides a multimedia segment attribution and structuring method and system for live streaming control events, and relates to the technical field of data processing. The method identifies a live streaming control event interval containing a complete closed loop of "audience feedback fluctuation -> streaming control intervention / anchor behavior adjustment -> audience feedback recovery" by monitoring abnormal fluctuations in audience feedback indicators in real time, extracts multimedia segments with narrative integrity and structured value from the interval, and solves the problem that existing slicing technology is disconnected from live streaming control records. By building a causal inference model, the average treatment effect of different anchor behavior characteristics and different intervention strategies on audience feedback recovery is quantified, and the causal attribution metadata of "why is it effective" is given to each multimedia segment, solving the problem that sliced content lacks strategy effectiveness attribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and system for attribution and structuring of multimedia segments for live stream control events. Background Technology

[0002] The large amount of recorded data generated after a live stream contains a wealth of reusable structured content assets, especially the host's responses to audience feedback, adjustments to explanations of professional questions, and the system's technical interventions during the live stream. However, existing live stream recording processing methods mainly rely on manual playback or physical segmentation based on fixed rules (such as duration, keywords, and interaction density). This is not only inefficient, but the resulting segments also lack in-depth attribution explanations of "why the host did what they did" and "whether this response was effective," making it difficult to effectively reuse them as structured multimedia assets with causal evidence.

[0003] At the same time, existing real-time live streaming processing systems (such as real-time compliance control systems and real-time content assistance systems) trigger various technical interventions based on audience feedback during the live stream, such as automatic bullet screen injection, video masking rendering, and audio mixing prompts. However, these intervention records are often just isolated log data, and they are not further systematically analyzed, attributed, or assetized after the live stream ends.

[0004] In existing technologies, the main methods for slicing live video recordings include the following: (1) Automatic slicing based on fixed duration Existing systems typically physically cut live stream recordings into short video segments according to preset time intervals (such as every 5 or 10 minutes). This method completely ignores the semantics of the content and the recording of technical interventions during the live stream. The cut points often fall when the host is halfway through speaking or before the system intervention is complete, resulting in abrupt segment boundaries.

[0005] (2) Automatic slicing based on topic similarity Some systems first perform topic clustering on the live stream subtitle text, grouping consecutive content under the same topic into a single segment. The segmentation logic is based on text similarity or keyword density, which cannot determine whether the segment contains a complete causal process of "fluctuations in audience feedback → technical intervention or anchor adjustment → recovery of audience feedback," nor can it explain why this segment is worth reusing.

[0006] (3) Automatic editing based on highlight clips Some live streaming platforms offer an automatic "highlight moment" editing function based on bullet screen density, number of gifts, or peak volume. This method selects segments based on "entertainment" or "marketing" criteria, completely ignoring the technical intervention records of the live stream control system and the audience's feedback and response process, and failing to identify structured content with strategic attribution value.

[0007] The existing technology mainly has the following problems: First, existing slicing technology is completely disconnected from live stream control recording. During a live stream, the system may trigger multiple stream modulation interventions (such as masking rendering, audio mixing, and bullet screen injection), but these intervention records are not used for boundary identification and causal attribution in post-event recordings, resulting in highly valuable technical intervention processes being buried in lengthy recordings.

[0008] Second, existing technology cannot provide attribution for the effectiveness of segmented content. It can only tell users "this segment is from 23 minutes 15 seconds to 23 minutes 50 seconds," but it cannot explain "why the host adjusted the speaking speed in this segment," "what role did the system's automatic bullet comment injection play," or "what would have happened if no intervention had been made." Segments without causal attribution are merely physical transfers of content.

[0009] Third, existing technologies cannot accumulate structured knowledge on "what response strategy should be adopted in what situation". Different broadcasters or systems may adopt completely different responses to the same fluctuations in audience feedback, some effective and some ineffective. Existing systems cannot make causal comparisons of the effects of these response strategies, resulting in the inability to identify, replicate, and systematically manage the most effective experiences.

[0010] Fourth, existing video clips lack structured coupling with content reuse systems. The cut video segments are often thrown into ordinary media libraries without accompanying attribution tags, applicable context descriptions, or search guidelines, leaving content operators overwhelmed by the sheer volume of clips.

[0011] Therefore, the aforementioned technical problems urgently need to be solved. Summary of the Invention

[0012] In view of the above problems, this application is proposed to provide a multimedia segment attribution and structuring method and system for live stream control events that overcomes or at least partially solves the above problems. The technical solution is as follows: Firstly, a multimedia segment attribution and structuring method for live stream control events is provided, the method comprising: Acquire multimodal data during the live broadcast, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows; Calculate the audience feedback metrics for each time window based on the window data; Based on audience feedback metrics and stream modulation intervention records in window data, identify live stream control event intervals, where each live stream control event interval has a start time point and an end time point; Within the live stream control event interval, extract the intervention strategy from the stream modulation intervention record; Extract the feature vector of the streamer's behavior before the occurrence of the live stream control event interval; Based on the intervention strategy and the anchor's behavioral feature vector, the average treatment effect of the intervention strategy on the audience feedback recovery is estimated by a pre-trained causal inference model. Based on the average treatment effect, effective intervention events are identified from the live stream control event interval; Multimedia clips are generated based on effective intervention events, and causal attribution metadata is generated for the multimedia clips; Multimedia clips and causal attribution metadata are stored in a structured multimedia database.

[0013] In one possible implementation, the multimodal data includes audio and video stream data, speech recognition text data, audience interaction text data, broadcaster screen sharing data, and stream modulation intervention record data; the window data includes broadcaster speech recognition text, broadcaster screen optical character recognition results, audience interaction text set, and stream modulation intervention record.

[0014] In one possible implementation, audience feedback metrics include at least one of the following: The sentiment density of bullet comments is calculated based on the probability of negative sentiment in each viewer interaction text within a time window. The concentration of bullet comments topics is calculated based on the semantic distribution of audience interaction text topics within a time window. Audience churn rate, which is calculated based on the number of viewers entering and leaving within a time window; Interaction deviation is calculated based on the cosine distance between the topic vector of the anchor's speech recognition text and the center of the topic vector of the audience's interactive text set.

[0015] In one possible implementation, the audience feedback metric is a weighted sum of the emotional density of the bullet comments, the concentration of bullet comment topics, the audience churn rate, and the degree of interaction deviation.

[0016] In one possible implementation, live stream control event intervals are identified based on audience feedback metrics and stream modulation intervention records in the window data, including: In multiple time windows, the time period in which consecutive abnormal time windows occur is identified as the live stream control event interval. Among them, the abnormal time window is the time window in which the audience feedback index is greater than the preset abnormal threshold. When M consecutive time windows are all abnormal time windows, the starting time point of the first time window in the M consecutive time windows is marked as the starting time point, where M is an integer greater than or equal to 2; After the starting time point, when K consecutive time windows are not abnormal time windows, the end time point of the last abnormal time window before the K consecutive time windows is marked as the end time point, where K is an integer greater than or equal to 1.

[0017] In one possible implementation, the intervention strategy includes systemic technical intervention and the anchor's self-adjustment; System technical interventions include one or more of the following: video masking rendering, audio mixing prompts, automatic bullet screen injection, and video segmentation marking; The anchor's self-adjustment is automatically identified through semantic analysis of the anchor's speech text and audio-visual feature detection. The anchor's self-adjustment includes one or more of the following: reduced speech rate, repeated explanation of concepts, use of analogy, and proactive questioning for confirmation.

[0018] In one possible implementation, the anchor's behavioral feature vector includes the following dimensions: The speech rate standardization value is the ratio of the broadcaster's average speech rate before the occurrence of the live stream control event interval to the average speech rate of the entire live stream. Terminology density, where terminology density is the proportion of terminology in the broadcaster's speech recognition text to the total number of words before the occurrence of the live stream control event interval; Visual assistance coverage rate, which is determined by the proportion of the graphic and text areas in the shared screen of the anchor that are semantically related to the spoken content to the total screen area. The graphic and text areas are identified by weighted semantic similarity between the optical character recognition results of the anchor screen and the text recognized by the anchor's speech. The topic shifting speed is determined based on the rate of change in the semantic similarity of topics in adjacent time windows before the occurrence of the live stream control event interval. The broadcaster's confidence level is calculated based on the frequency of certain words and intonation features in the broadcaster's speech.

[0019] In one possible implementation, audience feedback recovery is quantified through recovery metrics, which are calculated as follows: The recovery metric is equal to the difference between the maximum value of the audience feedback metric within the live stream control event interval and the audience feedback metric within a preset number of time windows after the event ends, divided by the difference between the maximum value of the audience feedback metric and the preset abnormal threshold.

[0020] In one possible implementation, the causal inference model is a causal forest model. The input features of the causal forest model include the anchor behavior feature vector, event intensity, and live broadcast topic. The event intensity is calculated based on the maximum value of the audience feedback index, the number of continuous windows, and the average audience churn rate within the live broadcast control event interval. The training data for the causal forest model consists of samples corresponding to multiple live stream control events accumulated in history. Each sample includes input features, intervention strategies, and recovery indicators. The training process of the causal forest model also includes propensity score matching, which is used to balance the samples before training to control selection bias caused by intelligent triggering of system technology intervention.

[0021] Secondly, a multimedia segment attribution and structuring system for live stream control events is provided, the system comprising: The live streaming multimodal data acquisition and storage module is used to acquire multimodal data during the live streaming process, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows. The audience feedback fluctuation monitoring module is used to calculate the audience feedback index corresponding to each time window based on the window data; The live stream control event recognition module is used to identify live stream control event intervals based on audience feedback indicators and stream modulation intervention records in window data. The live stream control event interval has a start time point and an end time point. The intervention strategy recording module is used to extract intervention strategies from the stream modulation intervention records within the live stream control event interval. The anchor behavior feature extraction module is used to extract the anchor behavior feature vector before the occurrence of the live stream control event interval; The causal inference modeling module is used to estimate the average treatment effect of the intervention strategy on the recovery of audience feedback based on the intervention strategy and the anchor's behavioral feature vector through a pre-trained causal inference model; and to determine the effective intervention events from the live stream control event interval based on the average treatment effect. The multimedia segment generation and structured database management module is used to generate multimedia segments based on effective intervention events and generate causal attribution metadata for the multimedia segments; and to store the multimedia segments and causal attribution metadata into a structured multimedia database.

[0022] By employing the aforementioned technical solutions, this application provides a multimedia segment attribution and structuring method and system for live stream control events. This method identifies live stream control event intervals containing a complete closed loop of "viewer feedback fluctuation → stream control intervention / anchor behavior adjustment → viewer feedback recovery" by real-time monitoring of abnormal fluctuations in viewer feedback indicators. It extracts multimedia segments with narrative integrity and structural value from these segments, addressing the disconnect between existing slicing techniques and live stream control recording. Furthermore, by constructing a causal inference model, it quantifies the average processing effect of different anchor behavior characteristics and different intervention strategies on viewer feedback recovery, assigning causal attribution metadata ("why it works") to each multimedia segment, addressing the lack of strategy effectiveness attribution in the sliced ​​content. Finally, it slices live stream control events into multimedia knowledge segments with attribution labels, clearly marking the contextual characteristics triggering the fluctuations, the intervention strategies employed, and evidence of strategy effectiveness, forming a searchable and comparable structured multimedia database, thus solving the problem of live stream experience not being able to be precipitated as structured multimedia assets. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0024] Figure 1 A flowchart of a multimedia segment attribution and structuring method for live stream control events provided in an embodiment of this application is shown; Figure 2 A structural diagram of a multimedia segment attribution and structuring system for live stream control events provided in an embodiment of this application is shown. Detailed Implementation

[0025] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."

[0027] To address the aforementioned technical problems, embodiments of this application provide a multimedia segment attribution and structuring method for live stream control events, such as... Figure 1 As shown, the multimedia segment attribution and structuring method for live stream control events may include the following steps S101 to S107: Step S101: Obtain multimodal data during the live broadcast, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows.

[0028] In this step, all multimodal data of the entire live broadcast can be collected, which may include audio and video stream data, automatic speech recognition (ASR) text data, audience interaction text data, broadcaster screen sharing data, and stream modulation intervention record data, etc. This embodiment does not limit this.

[0029] The entire multimodal data is stored in a time-aligned manner with a fixed time window granularity. The time window granularity can be set to 5 seconds or 10 seconds, which can be flexibly adjusted according to the live streaming scenario. The window data corresponding to each time window can include: the anchor's ASR text in that window, the anchor's screen optical character recognition (OCR) results, the audience interaction text set, and the record of the streaming modulation intervention performed within that window.

[0030] This step provides a unified time reference and a complete data foundation for subsequent fluctuation monitoring, event identification, and causal analysis by storing multi-dimensional and multi-modal data in a time-series aligned manner. This ensures that the data in each dimension corresponds accurately in the time dimension and avoids attribution bias caused by data misalignment.

[0031] Step S102: Calculate the audience feedback index corresponding to each time window based on the window data.

[0032] In this step, audience feedback metrics may include at least one of the following: The sentiment density of bullet comments is calculated based on the probability of negative sentiment in each viewer interaction text within a time window. The concentration of bullet comments topics is calculated based on the semantic distribution of audience interaction text topics within a time window. Audience churn rate, which is calculated based on the number of viewers entering and leaving within a time window; Interaction deviation is calculated based on the cosine distance between the topic vector of the anchor's speech recognition text and the center of the topic vector of the audience's interactive text set.

[0033] This step characterizes audience feedback from four dimensions: emotion, topic focus, user retention, and content relevance. It can comprehensively and multidimensionally identify abnormal fluctuations in the audience, avoid omissions or misjudgments based on a single indicator, and improve the accuracy of event identification.

[0034] In an optional embodiment, the audience feedback metric is a weighted sum of bullet screen sentiment density, bullet screen topic concentration, audience churn rate, and interaction deviation. Here, a comprehensive audience feedback metric is generated through a weighted combination, which allows for flexible adjustment of the weights of each dimension according to the live streaming scenario and domain characteristics, adapting to the feedback fluctuation characteristics of different types of live streams and improving the adaptability of anomaly detection.

[0035] Step S103: Identify the live stream control event interval based on the audience feedback indicators and the stream modulation intervention records in the window data, wherein the live stream control event interval has a start time point and an end time point.

[0036] Step S104: Extract the intervention strategy from the stream modulation intervention record within the live stream control event interval.

[0037] Step S105: Extract the feature vector of the anchor's behavior before the occurrence of the live stream control event interval.

[0038] Step S106: Based on the intervention strategy and the anchor's behavioral feature vector, estimate the average treatment effect of the intervention strategy on the recovery of audience feedback using a pre-trained causal inference model; based on the average treatment effect, determine the effective intervention events from the live stream control event interval.

[0039] Step S107: Generate multimedia segments based on effective intervention events, and generate causal attribution metadata for the multimedia segments; store the multimedia segments and causal attribution metadata in a structured multimedia database.

[0040] In this step, multimedia segments are generated based on effective intervention events. Specifically, the starting point of the multimedia segment can be set to a preset duration position before the start time point, and the ending point of the multimedia segment can be set to the time point after the end time point when the audience feedback indicates that the situation has returned to a stable state. The main video track, subtitle track, bullet screen track, auxiliary screen track, and intervention track within the multimedia segment are extracted.

[0041] Here, causal attribution metadata may include: Feedback fluctuation type label, wherein the feedback fluctuation type label is determined based on the fluctuation characteristics of the audience feedback index; The specific numerical values ​​of the anchor's behavioral feature vector; Types and parameters of intervention strategies; Estimates of the average treatment effect; Effectiveness rating, which is determined based on average treatment effect and recovery indicators; The structured suggestion text is automatically generated based on feedback fluctuation type tags, anchor behavior feature vectors, and intervention strategies.

[0042] Thus, the generated multimedia clips contain complete causal loops and multi-track auxiliary information, along with multi-dimensional causal attribution metadata, which can clearly explain the reuse value and applicable scenarios of the clips, solving the problems of traditional clips lacking attribution and being difficult to reuse.

[0043] This embodiment identifies live stream control event intervals containing a complete closed loop of "viewer feedback fluctuation → flow control intervention / anchor behavior adjustment → viewer feedback recovery" by real-time monitoring of abnormal fluctuations in viewer feedback indicators. Multimedia segments with narrative integrity and structured value are extracted from these intervals, addressing the disconnect between existing slicing technology and live stream control recording. A causal inference model is constructed to quantify the average processing effect of different anchor behavior characteristics and intervention strategies on viewer feedback recovery. Each multimedia segment is assigned causal attribution metadata explaining "why it works," addressing the lack of strategy effectiveness attribution in the sliced ​​content. Furthermore, live stream control events are sliced ​​into multimedia knowledge segments with attribution labels, clearly annotating the contextual characteristics triggering the fluctuations, the intervention strategies employed, and evidence of strategy effectiveness, forming a searchable and comparable structured multimedia database, solving the problem of live streaming experience not being able to be precipitated as structured multimedia assets.

[0044] This application embodiment provides a possible implementation method. In step S103, the live stream control event interval is identified based on the audience feedback indicators and the stream modulation intervention records in the window data. Specifically, this may include the following steps A1 to A3: Step A1: In multiple time windows, identify the time period in which consecutive abnormal time windows occur as the live stream control event interval, where the abnormal time window is the time window in which the audience feedback index is greater than the preset abnormal threshold. Step A2: When M consecutive time windows are all abnormal time windows, mark the starting time point of the first time window among the M consecutive time windows as the starting time point, where M is an integer greater than or equal to 2. Step A3: After the starting time point, when K consecutive time windows are not abnormal time windows, mark the end time point of the last abnormal time window before the K consecutive time windows as the end time point, where K is an integer greater than or equal to 1.

[0045] This embodiment, through the judgment logic of continuous abnormal windows, can accurately identify the complete interval from the occurrence of an anomaly to its recovery in viewer feedback, ensuring that the identified live stream control events have temporal continuity and completeness, and avoiding false triggering caused by isolated instantaneous anomalies.

[0046] This application provides a possible implementation method, in which the intervention strategy mentioned in step S104 may include system technical intervention and anchor self-adjustment; System technical interventions may include one or more of the following: video masking rendering, audio mixing prompts, automatic bullet comment injection, and video segmentation marking; The anchor's self-adjustment is automatically identified through semantic analysis of the anchor's speech text and audio-visual feature detection. The anchor's self-adjustment can include one or more of the following: slowing down the speech rate, repeating the explanation of concepts, using analogy rhetoric, and actively asking questions for confirmation.

[0047] This embodiment covers both automatic system intervention and anchor self-adjustment strategies, fully reproducing all response actions in the live stream control event, providing comprehensive processing variable data for subsequent causal attribution, and ensuring the integrity of the attribution analysis.

[0048] This application embodiment provides a possible implementation method, and the anchor behavior feature vector mentioned in step S105 may include the following dimensions: The speech rate standardization value is the ratio of the broadcaster's average speech rate before the occurrence of the live stream control event interval to the average speech rate of the entire live stream. Terminology density, where terminology density is the proportion of terminology in the broadcaster's speech recognition text to the total number of words before the occurrence of the live stream control event interval; Visual assistance coverage rate, which is determined by the proportion of the graphic and text areas in the shared screen of the anchor that are semantically related to the spoken content to the total screen area. The graphic and text areas are identified by weighted semantic similarity between the optical character recognition results of the anchor screen and the text recognized by the anchor's speech. The topic shifting speed is determined based on the rate of change in the semantic similarity of topics in adjacent time windows before the occurrence of the live stream control event interval. The broadcaster's confidence level is calculated based on the frequency of certain words and intonation features in the broadcaster's speech.

[0049] This embodiment systematically extracts the behavioral characteristics of the anchor before the event from five dimensions: speech rate, terminology density, visual aids, topic shift speed, and anchor confidence. It comprehensively describes "under what types of anchor behavior are abnormal fluctuations in audience feedback", providing rich covariate information for causal inference models and improving the accuracy of average treatment effect estimation.

[0050] This application embodiment provides a possible implementation method, in which the audience feedback recovery mentioned in step S106 is quantified by a recovery index, which is calculated as follows: The recovery metric is equal to the difference between the maximum value of the audience feedback metric within the live stream control event interval and the audience feedback metric within a preset number of time windows after the event ends, divided by the difference between the maximum value of the audience feedback metric and the preset abnormal threshold.

[0051] This embodiment quantifies the relative extent to which intervention strategies restore audience feedback from peak levels to normal levels by using recovery indicators, providing a standardized measurement method for evaluating intervention effectiveness and making the effects of different events comparable.

[0052] This application provides a possible implementation method. The causal inference model mentioned in step S106 is a causal forest model. The input features of the causal forest model include the anchor behavior feature vector, event intensity and live broadcast topic. The event intensity is calculated based on the maximum value of the audience feedback index, the number of continuous windows and the average audience churn rate within the live broadcast control event interval. The training data for the causal forest model consists of samples corresponding to multiple live stream control events accumulated in history. Each sample includes input features, intervention strategies, and recovery indicators. The training process of the causal forest model also includes propensity score matching, which is used to balance the samples before training to control selection bias caused by intelligent triggering of system technology intervention.

[0053] This embodiment employs a causal forest model to estimate the average treatment effect of different intervention strategies under different feature combinations. By incorporating event intensity, propensity score matching, and the honest estimation and local moment estimation techniques inherent in causal forests, it effectively mitigates selection bias caused by intelligent triggering of system technical interventions, making the estimation of the average treatment effect more accurate and reliable. With the accumulation of live streaming data, the estimation accuracy of the causal inference model will continuously improve, and the number of effective multimedia cases identified will become increasingly rich.

[0054] This application provides a possible implementation method, which may further include the following steps B1 and B2: Step B1: Receive a search request from the user terminal. The search request includes search conditions, which include at least one of the following: feedback fluctuation type label, anchor behavior characteristic dimension, intervention strategy type, effectiveness rating, and average treatment effect value range. Step B2: Based on the search criteria, match and return the corresponding multimedia segments from the structured multimedia database.

[0055] This embodiment equips multimedia segments with multi-dimensional structured causal attribution metadata and provides multi-dimensional combined search capabilities, enabling content operators, broadcaster trainers, or system decision-makers to accurately locate multimedia cases that meet specific contexts and strategy effectiveness criteria. Users can directly access high-quality knowledge segments—determined by "under what type of feedback fluctuation, for what broadcaster behavioral characteristics, what intervention strategy was adopted, and what effectiveness rating was achieved"—without needing to replay lengthy recordings or rely on vague keywords or duration guesses. This significantly lowers the barrier to reusing structured multimedia assets, transforming implicit coping experiences from massive amounts of recordings into an explicit knowledge base that can be precisely searched and compared, effectively supporting the rapid replication of best practices and strategy optimization, and greatly improving the efficiency of live streaming experience accumulation and reuse.

[0056] The above introduces Figure 1 The embodiments shown have various implementation methods for each stage. The following will further explain the multimedia segment attribution and structuring method for live stream control events according to the embodiments of this application through specific examples.

[0057] In specific embodiments, the process may include live multimodal data acquisition and storage, audience feedback fluctuation monitoring, live stream control event identification, intervention strategy recording, anchor behavior feature extraction, causal inference modeling, and multimedia segment generation and structured database management. These steps will be described in detail below.

[0058] 1) Live streaming multimodal data acquisition and storage The audio and video streams, ASR-transcribed text, audience comments, broadcaster screen sharing, and action records of interventions triggered by the live stream feedback control gate are stored in a time-aligned manner throughout the entire live stream. The timeline is sliced ​​with a fixed time window granularity (e.g., 5 seconds or 10 seconds), and each window contains at least: the broadcaster's ASR text, the broadcaster's screen OCR results, the audience comment text set, and the record of the stream modulation interventions performed within that window.

[0059] 2) Monitoring of fluctuations in audience feedback Continuous monitoring of audience behavior data is conducted within a fixed time window to calculate a sequence of audience feedback indicators. These indicators characterize the overall response of the audience to the current content or technical interventions, and can be calculated based on one or more data sources, including audience interaction text, broadcaster output, and audience size data.

[0060] In practice, this audience feedback metric can be obtained from one or more of the following dimensions: (1) Bullet screen sentiment density sequence s_t: Classify the sentiment of each comment text within the window and output the probability of negative sentiment. Then: s_t=(Σprobability of negative emotions_i) / N Where N is the total number of bullet comments. When s_t suddenly increases, it indicates that the audience has experienced significant emotional fluctuations.

[0061] (2) The sequence of the concentration of bullet comments topic c_t: Calculate the semantic concentration of the topic in the bullet screen text. The lower the concentration, the more scattered the audience discussion; a sudden drop in concentration often means that the streamer's content has failed to capture the audience's attention.

[0062] (3) Real-time audience churn rate sequence l_t: The net churn rate is calculated by obtaining viewer entry / exit data within the window through the live streaming platform's API (Application Programming Interface).

[0063] (4) Interaction deviation sequence d_t: Map the anchor's text and the bullet screen text to the topic vector space respectively, and calculate the topic center deviation angle d_t. The closer d_t is to 1, the more the audience's discussion deviates from the anchor's explanation topic.

[0064] The comprehensive audience feedback anomaly index A_t can be represented as a weighted combination of one or more of the above components: A_t=w_s×s_t+w_c×(1-c_t)+w_l×l_t+w_d×d_t; Where w_s, w_c, w_l, and w_d are weighting coefficients, and w_s+w_c+w_l+w_d=1.

[0065] 3) Live Stream Control Event Recognition A live stream control event is defined as a continuous time interval in which audience feedback metrics show significant abnormal fluctuations and the system or the streamer takes appropriate intervention measures.

[0066] Specific identification steps: (1) Set the threshold for abnormal audience feedback θ_A (e.g., 0.6). When A_t>θ_A is satisfied for M consecutive windows (M≥2, corresponding to at least 10 seconds), mark the starting point of this continuous interval as the point of generation of the live stream control event t_start.

[0067] (2) Continuously monitor subsequent windows until A_t drops below θ_A again and remains stable for K consecutive windows (K≥1), and mark the last abnormal window as the event end point t_end.

[0068] (3) For the identified live stream control events, further calculate the event intensity: EventStrength=max(A_t)×(t_end-t_start+1)×avg(l_t) Where max(A_t) is the maximum outlier during the event, (t_end-t_start+1) is the number of event duration windows, and avg(l_t) is the average audience churn rate during the event.

[0069] 4) Recording and classifying intervention strategies Within the live stream control event interval [t_start, t_end], the system records all triggered intervention strategies. Intervention strategies can be divided into two categories: (1) System technical intervention: Technical actions automatically executed by the live stream feedback control system, such as video masking rendering, audio mixing prompts, automatic bullet screen injection, video segmentation marking, etc. The system records the type of intervention strategy, the trigger time, and the specific stream modulation parameters.

[0070] (2) Host self-adjustment: The host makes self-adjustments after seeing system prompts or audience comments. Automatically identified through ASR text analysis and audio-visual feature detection, such as: significantly reduced speaking speed, repeated explanation of a concept, use of more analogies, and proactively asking questions to confirm audience feedback.

[0071] For each intervention strategy T, the system records its implementation timestamp t_T and the order of implementation in the event.

[0072] 5) Extraction of anchor behavior features For each live stream control event prior to its occurrence, the broadcaster's behavior feature vector X is extracted (with the analysis interval being L windows prior to the event, where L is typically 2-4, i.e., 10-20 seconds before the event). X=[x1,x2,x3,x4,x5] in: (1) x1 = Standardized speech rate value: the average speech rate of the anchor before the event, divided by the average speech rate of the entire live broadcast; (2) x2 = Terminology Density: The proportion of specialized terms in the anchor's ASR text before the event occurred to the total number of words. The terminology list is pre-constructed based on the domain knowledge graph entity dictionary; (3) x3 = Visual Auxiliary Coverage Rate: The proportion of the total screen area in the shared screen view of the anchor that is semantically related to the current spoken content. Screen text is extracted using OCR, and the semantically similar weighted coverage rate with the anchor's ASR text is calculated. (4) x4 = Topic shift speed: Measured by the rate of change in the semantic similarity of the topics between the window before the event and the window immediately before. A sharp drop in similarity indicates that the broadcaster has quickly switched topics; (5) Host Confidence: A confidence score calculated based on the frequency of certainty words and intonation features in the host's voice.

[0073] These feature vectors X describe "under what type of host behavior, abnormal fluctuations in audience feedback are likely to occur".

[0074] 6) Intervention Effect Evaluation For each live stream control event, evaluate the effect of the intervention strategy. Define the audience feedback recovery indicator as: Recovery=(max(A over [t_start,t_end])-A_{t_end+N}) / (max(A over [t_start,t_end])-θ_A) wherein max(A over [t_start,t_end]) is the maximum value of the audience feedback indicator within the interval of the live stream control event; A_{t_end+N} is the audience feedback indicator of the N-th window after the end of the event (e.g., N=2 or N=3), and this indicator measures the relative degree to which the intervention strategy restores audience feedback from the peak to the normal level.

[0075] When Recovery>0.3, the live stream control event is determined as an "effective intervention event".

[0076] 7) Average Treatment Effect Estimation Based on Causal Inference Using the historically accumulated live stream control event dataset, train a causal inference model to quantify the ATE (Average Treatment Effect) of different intervention strategies T on Recovery under different host behavior feature vectors X.

[0077] Data organization method: Unit: each live stream control event corresponds to one sample; Features: host behavior feature vector X, event strength EventStrength, live stream theme when the event occurs, etc.; Treatment variable: intervention strategy T; Outcome variable: Recovery.

[0078] Model selection: Causal Forest model. The model training objective is: τ(X)=E[Recovery|X,T=1]-E[Recovery|X,T=0] wherein τ(X) represents the average treatment effect of adopting the intervention strategy T relative to the control group under the condition of the host behavior feature vector X.

[0079] Selection bias that needs to be controlled: Since system interventions are typically triggered intelligently based on event severity, selection bias may occur. To mitigate this issue: (1) Add EventStrength and max(A_t) to X so that the model can control the impact of event severity; (2) Use PSM (Propensity Score Matching) to balance the samples before training; (3) Causal forests themselves further reduce the estimation bias caused by model overfitting through honest estimation and local moment estimation techniques.

[0080] 8) Generation and structuring of multimedia segments For events requiring effective intervention, the system automatically generates multimedia clips with attribution metadata. The generation steps are as follows: Step 1: Determine the segment boundaries The multimedia segment starts K seconds before the live stream control event occurs (e.g., 5 seconds before t_start) and ends when the audience feedback returns to a stable state after the event (usually 1-2 windows after t_end). The segment cut in this way contains a complete closed loop of "the original behavior of the anchor → fluctuations in audience feedback → technical intervention / anchor adjustment → recovery of audience feedback".

[0081] Step 2: Extract multitrack content from the fragment In addition to the main video track, the system also extracts the following auxiliary tracks: Subtitle track: ASR text and key concepts highlighted; Bullet screen track: Filter and display key audience feedback bullet screen comments; Auxiliary screen track: Keyframes of the broadcaster's shared screen; Intervention Track: The timing and specific parameters of system technical interventions and anchor self-adjustments are marked in timeline form at the bottom or side of the screen.

[0082] Step 3: Generate causal attribution metadata For each multimedia segment, the system generates structured causal attribution metadata, including at least: (1) Feedback fluctuation type labels: emotional outburst type, theme deviation type, audience loss type, comprehensive fluctuation type; (2) The specific value of the trigger feature X; (3) Intervention strategies adopted: System technical intervention (specific type and parameters), anchor self-adjustment (specific description); (4) Estimate ATE value: The estimated value of the average treatment effect τ(X) of this strategy T relative to other strategies under the current feature X; (5) Effectiveness rating: If τ(X)>0.5 and Recovery>0.5, it is rated as “highly effective”; if 0.3<τ(X)≤0.5, it is rated as “moderately effective”; otherwise, it is rated as “weakly effective or situation-dependent”. (6) Structured suggestion copy: Based on the feedback fluctuation type, trigger characteristics and intervention strategy, structured suggestions for content operation personnel are automatically generated.

[0083] Step 4: Store in a structured multimedia database Multimedia clips and their causal attribution metadata are stored in a structured multimedia database. The database supports retrieval by dimensions such as feedback fluctuation type, broadcaster behavior characteristics, intervention strategy type, effectiveness rating, and ATE value range.

[0084] 9) Content recommendation application of structured multimedia assets Based on the user's skill profile or content needs, the system retrieves the most relevant multimedia clips from a structured multimedia database. For example, if a user needs to find a case of "a broadcaster speaking too fast, causing an emotional outburst in the audience, and successfully recovering through dimensionality reduction explanation," the system can directly retrieve the corresponding tags and return the clip with the highest ATE value.

[0085] This embodiment can achieve the following technical effects: (1) Upgrade live stream control records from post-event logs to structured multimedia assets. Existing technologies store stream modulation intervention records only as isolated logs. This application is the first to combine these intervention records with the recovery process of anchor behavior and audience feedback, and extract them into complete multimedia segments with attribution information.

[0086] (2) Assign causal attribution metadata to multimedia segments By estimating the ATE (Aspect-Oriented Termination) of different intervention strategies using a causal forest model, the system can answer questions such as "Why is this segment worth reusing?" and "In what context is this strategy most effective?"

[0087] (3) Transform implicit coping experiences into explicit structured multimedia assets The effectiveness of outstanding anchors' on-the-spot adjustment experience and systematic intervention strategies is structurally identified, verified, and accumulated through causal attribution analysis.

[0088] (4) Construct a continuously evolving structured multimedia database As live streaming data accumulates, the estimation accuracy of the causal inference model will continue to improve, and the number of effective multimedia cases identified will become increasingly rich.

[0089] It should be noted that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In practical applications, all the above possible implementation methods can be arbitrarily combined in a combined manner to form possible embodiments of this application, which will not be described in detail here.

[0090] Based on the multimedia segment attribution and structuring methods for live stream control events provided in the above embodiments, and based on the same inventive concept, this application also provides a multimedia segment attribution and structuring system for live stream control events.

[0091] Figure 2 This is a structural diagram of a multimedia segment attribution and structuring system for live stream control events provided in an embodiment of this application. Figure 2 As shown, the multimedia segment attribution and structuring system for live stream control events may specifically include a live multimodal data acquisition and storage module 210, a viewer feedback fluctuation monitoring module 220, a live stream control event identification module 230, an intervention strategy recording module 240, a broadcaster behavior feature extraction module 250, a causal inference modeling module 260, and a multimedia segment generation and structured database management module 270.

[0092] The live streaming multimodal data acquisition and storage module 210 is used to acquire multimodal data during the live streaming process, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows. The audience feedback fluctuation monitoring module 220 is used to calculate the audience feedback index corresponding to each time window based on the window data. The live stream control event recognition module 230 is used to identify live stream control event intervals based on audience feedback indicators and stream modulation intervention records in window data. The live stream control event interval has a start time point and an end time point. The intervention strategy recording module 240 is used to extract the intervention strategy from the stream modulation intervention record within the live stream control event interval. The anchor behavior feature extraction module 250 is used to extract the anchor behavior feature vector before the occurrence of the live stream control event interval; The causal inference modeling module 260 is used to estimate the average treatment effect of the intervention strategy on the audience feedback recovery based on the intervention strategy and the anchor's behavioral feature vector through a pre-trained causal inference model; and to determine the effective intervention events from the live stream control event interval based on the average treatment effect. The multimedia segment generation and structured database management module 270 is used to generate multimedia segments based on effective intervention events and generate causal attribution metadata for the multimedia segments; and to store the multimedia segments and causal attribution metadata into a structured multimedia database.

[0093] This application provides a possible implementation method, in which multimodal data includes audio and video stream data, speech recognition text data, audience interaction text data, anchor screen sharing image data, and stream modulation intervention record data; window data includes anchor speech recognition text, anchor screen optical character recognition results, audience interaction text set, and stream modulation intervention record.

[0094] This application provides a possible implementation method, where the audience feedback indicators include at least one of the following: The sentiment density of bullet comments is calculated based on the probability of negative sentiment in each viewer interaction text within a time window. The concentration of bullet comments topics is calculated based on the semantic distribution of audience interaction text topics within a time window. Audience churn rate, which is calculated based on the number of viewers entering and leaving within a time window; Interaction deviation is calculated based on the cosine distance between the topic vector of the anchor's speech recognition text and the center of the topic vector of the audience's interactive text set.

[0095] This application provides a possible implementation method in which the audience feedback index is a weighted sum of the emotional density of bullet comments, the concentration of bullet comment topics, the audience churn rate, and the degree of interaction deviation.

[0096] This application embodiment provides a possible implementation method, wherein the live stream control event recognition module 230 is further used for: In multiple time windows, the time period in which consecutive abnormal time windows occur is identified as the live stream control event interval. Among them, the abnormal time window is the time window in which the audience feedback index is greater than the preset abnormal threshold. When M consecutive time windows are all abnormal time windows, the starting time point of the first time window in the M consecutive time windows is marked as the starting time point, where M is an integer greater than or equal to 2; After the starting time point, when K consecutive time windows are not abnormal time windows, the end time point of the last abnormal time window before the K consecutive time windows is marked as the end time point, where K is an integer greater than or equal to 1.

[0097] This application provides a possible implementation method, and the intervention strategy includes system technical intervention and anchor self-adjustment; System technical interventions include one or more of the following: video masking rendering, audio mixing prompts, automatic bullet screen injection, and video segmentation marking; The anchor's self-adjustment is automatically identified through semantic analysis of the anchor's speech text and audio-visual feature detection. The anchor's self-adjustment includes one or more of the following: reduced speech rate, repeated explanation of concepts, use of analogy, and proactive questioning for confirmation.

[0098] This application provides a possible implementation method, where the anchor behavior feature vector includes the following dimensions: The speech rate standardization value is the ratio of the broadcaster's average speech rate before the occurrence of the live stream control event interval to the average speech rate of the entire live stream. Terminology density, where terminology density is the proportion of terminology in the broadcaster's speech recognition text to the total number of words before the occurrence of the live stream control event interval; Visual assistance coverage rate, which is determined by the proportion of the graphic and text areas in the shared screen of the anchor that are semantically related to the spoken content to the total screen area. The graphic and text areas are identified by weighted semantic similarity between the optical character recognition results of the anchor screen and the text recognized by the anchor's speech. The topic shifting speed is determined based on the rate of change in the semantic similarity of topics in adjacent time windows before the occurrence of the live stream control event interval. The broadcaster's confidence level is calculated based on the frequency of certain words and intonation features in the broadcaster's speech.

[0099] This application provides a possible implementation method in which audience feedback recovery is quantified through recovery indicators, which are calculated as follows: The recovery metric is equal to the difference between the maximum value of the audience feedback metric within the live stream control event interval and the audience feedback metric within a preset number of time windows after the event ends, divided by the difference between the maximum value of the audience feedback metric and the preset abnormal threshold.

[0100] This application provides a possible implementation method in which the causal inference model is a causal forest model. The input features of the causal forest model include the anchor behavior feature vector, event intensity and live broadcast topic. The event intensity is calculated based on the maximum value of the audience feedback index, the number of continuous windows and the average audience churn rate within the live broadcast control event interval. The training data for the causal forest model consists of samples corresponding to multiple live stream control events accumulated in history. Each sample includes input features, intervention strategies, and recovery indicators. The training process of the causal forest model also includes propensity score matching, which is used to balance the samples before training to control selection bias caused by intelligent triggering of system technology intervention.

[0101] Those skilled in the art will clearly understand that the specific working process of the systems, devices, and modules described above can be referred to the corresponding process in the foregoing method embodiments. For the sake of brevity, it will not be repeated here.

[0102] Those skilled in the art will understand that the technical solution of this application, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several program instructions to cause an electronic device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of this application when running the program instructions. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0103] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as electronic devices like personal computers, servers, or network devices) associated with program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the embodiments of this application.

[0104] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of this application, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to leave the protection scope of this application.

Claims

1. A multimedia segment attribution and structuring method for live stream control events, characterized in that, The method includes: Acquire multimodal data during the live broadcast, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows; Calculate the audience feedback metrics for each time window based on the window data; Based on audience feedback metrics and stream modulation intervention records in window data, identify live stream control event intervals, where each live stream control event interval has a start time point and an end time point; Within the live stream control event interval, extract the intervention strategy from the stream modulation intervention record; Extract the feature vector of the streamer's behavior before the occurrence of the live stream control event interval; Based on the intervention strategy and the anchor's behavioral feature vector, the average treatment effect of the intervention strategy on the recovery of audience feedback is estimated by a pre-trained causal inference model; based on the average treatment effect, effective intervention events are identified from the live stream control event interval. Multimedia clips are generated based on effective intervention events, and causal attribution metadata is generated for the multimedia clips; the multimedia clips and causal attribution metadata are stored in a structured multimedia database.

2. The method according to claim 1, characterized in that, Multimodal data includes audio and video stream data, speech recognition text data, audience interaction text data, anchor screen sharing data, and stream modulation intervention record data; window data includes anchor speech recognition text, anchor screen optical character recognition results, audience interaction text set, and stream modulation intervention record.

3. The method according to claim 2, characterized in that, Audience feedback metrics include at least one of the following: The sentiment density of bullet comments is calculated based on the probability of negative sentiment in each viewer interaction text within a time window. The concentration of bullet comments topics is calculated based on the semantic distribution of audience interaction text topics within a time window. Audience churn rate, which is calculated based on the number of viewers entering and leaving within a time window; Interaction deviation is calculated based on the cosine distance between the topic vector of the anchor's speech recognition text and the center of the topic vector of the audience's interactive text set.

4. The method according to claim 3, characterized in that, The audience feedback metrics are a weighted sum of the emotional density of bullet comments, the concentration of bullet comment themes, the audience churn rate, and the degree of interaction deviation.

5. The method according to claim 1, characterized in that, Based on audience feedback metrics and stream modulation intervention records in the window data, identify live stream control event intervals, including: In multiple time windows, the time period in which consecutive abnormal time windows occur is identified as the live stream control event interval. Among them, the abnormal time window is the time window in which the audience feedback index is greater than the preset abnormal threshold. When M consecutive time windows are all abnormal time windows, the starting time point of the first time window in the M consecutive time windows is marked as the starting time point, where M is an integer greater than or equal to 2; After the starting time point, when K consecutive time windows are not abnormal time windows, the end time point of the last abnormal time window before the K consecutive time windows is marked as the end time point, where K is an integer greater than or equal to 1.

6. The method according to claim 2, characterized in that, Intervention strategies include systemic technical intervention and self-adjustment by the broadcaster; System technical interventions include one or more of the following: video masking rendering, audio mixing prompts, automatic bullet screen injection, and video segmentation marking; The anchor's self-adjustment is automatically identified through semantic analysis of the anchor's speech text and audio-visual feature detection. The anchor's self-adjustment includes one or more of the following: reduced speech rate, repeated explanation of concepts, use of analogy, and proactive questioning for confirmation.

7. The method according to claim 2, characterized in that, The anchor's behavioral feature vector includes the following dimensions: The speech rate standardization value is the ratio of the broadcaster's average speech rate before the occurrence of the live stream control event interval to the average speech rate of the entire live stream. Terminology density, where terminology density is the proportion of terminology in the broadcaster's speech recognition text to the total number of words before the occurrence of the live stream control event interval; Visual assistance coverage rate, which is determined by the proportion of the graphic and text areas in the shared screen of the anchor that are semantically related to the spoken content to the total screen area. The graphic and text areas are identified by weighted semantic similarity between the optical character recognition results of the anchor screen and the text recognized by the anchor's speech. The topic shifting speed is determined based on the rate of change in the semantic similarity of topics in adjacent time windows before the occurrence of the live stream control event interval. The broadcaster's confidence level is calculated based on the frequency of certain words and intonation features in the broadcaster's speech.

8. The method according to claim 1, characterized in that, Audience feedback recovery is quantified using recovery metrics, which are calculated as follows: The recovery metric is equal to the difference between the maximum value of the audience feedback metric within the live stream control event interval and the audience feedback metric within a preset number of time windows after the event ends, divided by the difference between the maximum value of the audience feedback metric and the preset abnormal threshold.

9. The method according to claim 1, characterized in that, The causal inference model is a causal forest model. The input features of the causal forest model include the anchor behavior feature vector, event intensity, and live broadcast topic. Among them, the event intensity is calculated based on the maximum value of the audience feedback index, the number of continuous windows, and the average audience churn rate within the live broadcast control event interval. The training data for the causal forest model consists of samples corresponding to multiple live stream control events accumulated in history. Each sample includes input features, intervention strategies, and recovery indicators. The training process of the causal forest model also includes propensity score matching, which is used to balance the samples before training to control selection bias caused by intelligent triggering of system technology intervention.

10. A multimedia segment attribution and structuring system for live stream control events, characterized in that, The system includes: The live streaming multimodal data acquisition and storage module is used to acquire multimodal data during the live streaming process, perform time-series alignment on the multimodal data, and obtain window data corresponding to each of the multiple time windows. The audience feedback fluctuation monitoring module is used to calculate the audience feedback index corresponding to each time window based on the window data; The live stream control event recognition module is used to identify live stream control event intervals based on audience feedback indicators and stream modulation intervention records in window data. The live stream control event interval has a start time point and an end time point. The intervention strategy recording module is used to extract intervention strategies from the stream modulation intervention records within the live stream control event interval. The anchor behavior feature extraction module is used to extract the anchor behavior feature vector before the occurrence of the live stream control event interval; The causal inference modeling module is used to estimate the average treatment effect of the intervention strategy on the recovery of audience feedback based on the intervention strategy and the anchor's behavioral feature vector through a pre-trained causal inference model; and to determine the effective intervention events from the live stream control event interval based on the average treatment effect. The multimedia segment generation and structured database management module is used to generate multimedia segments based on effective intervention events and generate causal attribution metadata for the multimedia segments; and to store the multimedia segments and causal attribution metadata into a structured multimedia database.