AI director assistant method and system capable of implementing intelligent interaction effect
By using an AI-powered director assistant method, the spatiotemporal alignment and semantic feature extraction of multiple audio and video signals are achieved, key scenes are identified, candidate shot schemes are generated and prioritized, solving the problem of response delay in the traditional director process and improving the accuracy of capturing exciting moments and the continuity of presentation.
Patent Information
- Application Number
- CN202511637058.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-06
AI Technical Summary
Traditional broadcasting processes lack autonomous understanding and cannot complete a closed-loop response from signal input to decision output within milliseconds. This results in delayed responses from human broadcasters, making it difficult to capture exciting moments and affecting the program's dissemination and audience experience.
The AI-powered broadcast assistant method is adopted to generate a synchronized audio and video stream set and spatial mapping parameters by synchronously accessing multiple audio and video signals and performing spatiotemporal alignment. Layered analysis and semantic feature extraction are then performed to identify individual behavioral profiles, predict behavioral intentions, construct a scene salience scoring system, generate candidate shot schemes and prioritize them, and execute shot switching and dynamic composition adjustment.
It achieves an automated closed loop from signal perception to behavior understanding to image output, significantly shortening event response time, improving the accuracy of capturing wonderful moments and the continuity of presentation, and reducing reliance on human experience.
Smart Images

Figure CN121486599A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of director, and particularly relates to an AI director assistant method and system capable of implementing intelligent interactive effect. BACKGROUND
[0002] In the current multi-camera video production environment, director work generally relies on manual operation. The director needs to observe multiple picture signals in real time, manually selects the main picture according to the on-site speech, the emotional changes of the characters and the audience reaction, and controls the camera switching. In this mode, the operator needs to process a large amount of audio-visual information at the same time, and the judgment process is easily affected by subjective experience, and it is difficult to take into account all key details in the scene of multi-person interaction and rapid rhythm. Especially when important speeches or emotional climax appear, due to the delay of manual response, the wonderful moment is often not captured in time, the best shot presentation opportunity is missed, and the transmission effect of the final program and the audience experience are affected.
[0003] The core of the above problem is that the traditional director process lacks the ability to understand the content of the scene independently, and cannot complete the closed-loop response from signal input to decision output within milliseconds. Although the existing technology has realized the functions of multi-channel signal access and basic switching, it has not established a deep analysis chain from raw audio and video to semantic scene, so that the system cannot "understand" what is happening in the scene like a human being, and then make a forward-looking judgment. SUMMARY
[0004] The purpose of the present application is to provide an AI director assistant method and system capable of implementing intelligent interactive effect, which improves the capture accuracy and presentation continuity of wonderful moments, and effectively solves the problems of response delay and one-sided judgment of manual director.
[0005] To achieve the above purpose, the application adopts the following technical scheme: an AI director assistant method capable of implementing intelligent interactive effect, comprising the following steps: Synchronously access multiple audio and video signals and perform time and space alignment to generate a synchronized audio and video stream set and a space mapping parameter; based on the synchronized audio and video stream set and the space mapping parameter, perform layered analysis and semantic feature extraction on the audio-visual content to form individual behavior portraits; based on the individual behavior portraits, identify social function roles and predict behavior intentions; according to the social function roles and behavior intentions, build a scene saliency scoring system to identify key scenes; based on the key scenes, generate candidate shot schemes and perform priority sorting, and execute camera switching and dynamic composition adjustment through the shot schemes sorted by priority.
[0006] Preferably, the synchronous access of multiple audio and video signals and the time and space alignment comprise: Access audio and video streams from N independent devices and record original time stamps; Mapping the original time stamp to the global standard time coordinate system according to the master clock source; Aligning the collection time of each video frame, making all video streams reach frame-level synchronization, and then establishing a spatial registration relationship and solving a projection transformation matrix by using preset marker points.
[0007] Preferably, the layered analysis and semantic feature extraction of the audio-visual content comprises: Detecting a face region in each frame of video and calculating a face area proportion; Performing expression state discrimination on the detected face region; Performing sound source separation and speech activity detection on the audio stream, and then fusing visual and auditory features to establish an individual behavior portrait.
[0008] Preferably, the identification of social function roles and the prediction of behavior intention comprises: Statistically calculating the cumulative speech duration and average speech intensity of an individual within continuous 5 minutes, and calculating a dominance index based on the cumulative speech duration and average speech intensity; Monitoring the mutation of individual behavior, and defining a state change rate based on the change trend of the dominance index; Analyzing the attention distribution of a group, and statistically calculating the proportion of the gaze direction pointing to a specific person based on the state change rate and the gaze direction vector; Establishing a role state transition graph, and recording the state conversion frequency and duration based on the gaze direction proportion and the state change rate.
[0009] Preferably, the identification of a key scene comprises: Calculating a time continuity factor based on the time length of the continuously high active state of a social function role; Calculating a spatial concentration factor based on the position coordinates of an individual with an interactive intention; Calculating an emotional intensity factor based on the individual expression index and the audio energy variance; Calculating a scene saliency score based on the time continuity factor, the spatial concentration factor, the emotional intensity factor and the interactive potential factor, and identifying a key scene.
[0010] Preferably, the generation of a candidate shot scheme and the priority sorting comprise: According to the key scene type, calling a shot logic template library, and matching an initial shot combination based on the key scene type; Calculating coverage effectiveness based on the coverage set and the activity score of each shot in the initial shot combination; Judging whether to impose a delay penalty based on the difference in the angle of view between the last enabled shot and the current candidate shot; Comprehensively scoring and sorting the candidate shots based on the coverage effectiveness, the scene saliency score and the delay penalty result.
[0011] Preferably, the execution of camera switching and dynamic composition adjustment includes: Send control commands to the physical camera corresponding to the preferred lens; During camera transitions, a screen transition effect is activated simultaneously. After the switch is completed, the system continuously monitors the positional shift of the main characters in the scene and automatically adjusts the exposure parameters according to changes in ambient lighting.
[0012] Preferably, the method further includes: collecting bullet comments and likes data from the live streaming platform, aligning them based on the timestamp of the main video stream, and analyzing the trend of audience sentiment, specifically including: Real-time capture of newly added bullet screen text, and calculation of sentiment index based on the frequency of positive and negative words; Monitor the time series of likes, calculate the instantaneous growth rate based on the change in likes within the corresponding time period of the sentiment index, and determine the surge in interaction popularity; Based on the comment content during the periods of surge in interactive popularity, the focus of audience discussion is identified, and the social media comments are clustered by topic. An audience resonance index is constructed based on the aforementioned sentiment tendency index, instantaneous growth rate, and frequency of mentioning the current speaker's name.
[0013] Preferably, the method further includes: associating the contextual information of key scenes with the emotional trends of the audience and writing it into the event log, specifically including: writing the complete contextual information of key scenes into the event log; performing pattern mining on historical logs to find frequently co-occurring scene-shot combinations; establishing individual performance profiles, recording historical behavioral characteristics, and generating a director's quality report.
[0014] On the other hand, this invention proposes an AI-powered broadcast assistant system capable of implementing intelligent interactive effects, comprising: The spatiotemporal alignment module is used to perform synchronous access to multiple audio and video signals and spatiotemporal alignment, generating a synchronized audio and video stream set and spatial mapping parameters; The semantic feature extraction module is used to perform hierarchical analysis and semantic feature extraction of audiovisual content based on the synchronized audio and video stream set and spatial mapping parameters, and form an individual behavior profile. The behavioral intent inference module is used to identify social functional roles and predict behavioral intent based on the individual's behavioral profile. The key scenario determination module is used to construct a scenario salience scoring system based on the social functional roles and behavioral intentions to identify key scenarios; The priority sorting module is used to generate candidate shot schemes based on the key scenes and sort them by priority; The dynamic adjustment module is used to execute the shot scheme sorted by priority, and to perform shot switching and dynamic composition adjustment.
[0015] Technical effects and advantages of the present invention: The AI-powered live broadcast assistant method and system capable of implementing intelligent interactive effects proposed in this invention have the following advantages compared with the prior art: This invention generates a synchronized audio and video stream set and spatial mapping parameters under a unified benchmark by synchronously accessing multiple audio and video signals and performing spatiotemporal alignment, providing a foundation for cross-perspective analysis. Based on this foundation, the audiovisual content is analyzed in layers and semantic features are extracted to form individual behavioral profiles including dimensions such as location, expression, and voice. Furthermore, social functional roles are identified and behavioral intentions are predicted based on these individual behavioral profiles, enabling the system to judge the identity of individuals and the trend of their actions. A scene salience scoring system is constructed based on social functional roles and behavioral intentions to automatically identify key scenes with dissemination value. Finally, candidate shot schemes are generated and ranked based on the key scenes, and shot switching and dynamic composition adjustments are performed. This method achieves an automated closed loop from signal perception to behavioral understanding to image output, significantly shortening the time difference from event occurrence to shot response, and improving the accuracy of capturing exciting moments and the continuity of presentation. Attached Figure Description
[0016] Figure 1 This is a flowchart of the AI broadcast assistant method for implementing intelligent interactive effects according to the present invention; Figure 2 This is a block diagram of the AI broadcast assistant system of the present invention, which can implement intelligent interactive effects. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] This invention provides, for example Figure 1This paper presents an AI-powered live broadcast assistant method capable of implementing intelligent interactive effects. In scenarios with parallel input from multiple audio and video signal sources, it enables autonomous analysis and understanding of live content, identifying key scenes, speaker identities, and memorable moments with broadcast value, thereby driving automated switching and presentation of broadcast actions. This method constructs a spatiotemporally coordinated perception structure, combined with a semantic-level dynamic evaluation system, to gradually form a comprehensive judgment of live events. Based on the judgment results, it generates output instructions with interactive guidance capabilities, effectively reducing reliance on human experience and improving the coherence and timeliness of switching responses. Specifically, it includes the following steps: The system synchronously accesses multiple audio and video signals and performs spatiotemporal alignment to generate a synchronized audio and video stream set and spatial mapping parameters. Specifically, this includes: accessing audio and video streams from N independent devices and recording the original timestamps; generating a global standard time based on the master clock source and mapping the original timestamps to the global standard time coordinate system; aligning the acquisition times of each video frame to achieve frame-level synchronization of all video streams; then establishing a spatial registration relationship and solving the projection transformation matrix using preset marker points.
[0019] This step ensures the consistency of multi-source signals in both time and space, eliminating information misalignment caused by asynchronous devices, transmission delays, or differences in viewing angle. Temporal alignment makes audio and video events comparable, supporting precise temporal analysis across different viewpoints; spatial registration enables the mapping of the position of the same target in different frames, providing a geometric basis for subsequent cross-camera tracking of individual movement and determination of the relative positional relationships of people, ensuring the physical accuracy of the analysis results.
[0020] Based on the synchronized audio and video stream set and spatial mapping parameters, the audiovisual content is analyzed in layers and semantic features are extracted to form an individual behavior profile. Specifically, this includes: detecting the face region in each frame of video and calculating the face area ratio; judging the expression state of the detected face region; separating the sound source and detecting speech activity in the audio stream, and then fusing visual and auditory features to establish an individual behavior profile.
[0021] This step transforms raw sensory signals into comprehensible behavioral states, accurately depicting an individual's outward performance through multimodal information fusion. The collaborative analysis of visual and auditory features effectively distinguishes between active expression and passive presence, enhancing the system's ability to identify truly active individuals. The resulting individual behavioral profile includes dynamic facial expressions, vocal activity trends, and visual salience, providing fine-grained data support for subsequent judgments of their social roles and behavioral intentions, thus improving the system's depth of understanding of complex interactive scenarios.
[0022] Based on individual behavioral profiles, the system identifies social functional roles and predicts behavioral intentions. Specifically, this includes: calculating the cumulative speaking time and average voice intensity of an individual over a continuous 5-minute period, and calculating a dominance index based on the cumulative speaking time and average voice intensity; monitoring sudden changes in individual behavior and defining a state change rate based on the trend of the dominance index; analyzing the distribution of group attention and calculating the proportion of gazes directed at specific individuals based on the state change rate and gaze direction vector; and establishing a role state transition map and recording the frequency and duration of state transitions based on the gaze direction ratio and the state change rate.
[0023] This step enables dynamic identification of an individual's functional role within a group, distinguishing between speakers, interaction initiators, and ordinary participants, moving beyond simple judgments based solely on vocalization. By combining behavioral stability and abrupt change characteristics, the system can predict an individual's upcoming action trends, such as speaking preparation or attention shifts. Quantitative analysis of group attention introduces social cues, enhancing consensus on the focus of the moment. Recording role state transitions gives the system short-term memory capabilities, enabling it to understand the continuity of behavior and contextual relationships, providing a more in-depth situational basis for scene judgment.
[0024] Based on social functional roles and behavioral intentions, a scene salience scoring system is constructed to identify key scenes. Specifically, this includes: calculating a time continuity factor based on the duration of sustained high activity in social functional roles; calculating a spatial concentration factor based on the location coordinates of individuals with interaction intentions; calculating an emotion intensity factor based on individual facial expression index and audio energy variance; and calculating a scene salience score based on the time continuity factor, spatial concentration factor, emotion intensity factor, and interaction potential factor to identify key scenes.
[0025] This step, through the fusion of multi-dimensional indicators, achieves a comprehensive assessment of the importance of on-site events, accurately capturing high-profile moments with dissemination value. The combination of temporal continuity and spatial concentration effectively identifies typical scenarios such as sustained speeches and group interactions, avoiding false triggering of brief actions. The emotional intensity factor introduces non-verbal emotional signals, enabling the system to perceive collective emotional fluctuations such as applause and laughter. The overall scoring system considers content depth, interaction breadth, and emotional intensity, improving the comprehensiveness and robustness of key scene identification and ensuring that the director's response focuses on segments with genuine narrative value.
[0026] Based on key scenarios, candidate shot schemes are generated and prioritized. Specifically, this includes: calling the shot logic template library according to the key scenario type, matching the initial shot combination based on the key scenario type; calculating the coverage effectiveness based on the coverage set and activity score of each shot in the initial shot combination; determining whether to apply a delay penalty based on the difference in perspective between the last activated shot number and the current candidate shot; and comprehensively scoring and ranking the candidate shots based on the coverage effectiveness, scene salience score, and delay penalty results.
[0027] This step enables intelligent shot selection and contextual adaptation, ensuring that the visual presentation both highlights the core content and conforms to visual narrative principles. By matching scene types with shot templates, the system can adopt optimal composition strategies for different events, enhancing the professionalism of the expression. The assessment of coverage effectiveness guarantees the information density of the image, avoiding transitions to ineffective perspectives. Visual fluidity constraints effectively suppress abrupt perspective jumps, maintaining viewing continuity and comfort. The comprehensive scoring mechanism strikes a balance between content importance and viewing experience, ensuring that shot decisions are not only accurate but also natural and fluid, enhancing the narrative quality and watchability of the final output video.
[0028] The system employs a priority-based shot selection mechanism to execute shot switching and dynamic composition adjustments. Specifically, this includes: sending control commands to the physical camera corresponding to the preferred shot; activating screen transition effects during shot switching; and continuously monitoring the positional shift of the main characters in the frame after switching, while automatically adjusting exposure parameters based on changes in ambient lighting.
[0029] This process ensures the precise implementation of the director's decisions and the consistent stability of image quality. Reliable execution of instructions achieves low-latency response from judgment to presentation, guaranteeing the timely capture of crucial moments. Automatic matching of transition effects enhances the naturalness of switching and strengthens the coherence of the visual narrative. Dynamic composition adjustment effectively addresses image imbalance caused by character movement, maintaining the subject's reasonable position at the visual center of gravity. Automatic exposure adjustment adapts to changes in ambient lighting, avoiding impacts on the viewing experience due to sudden brightness fluctuations. The overall output image maintains professional-grade stability and clarity, significantly improving the final video's production quality and viewer immersion.
[0030] Furthermore, the above method also includes: collecting bullet comments and likes data from the live streaming platform, aligning them based on the timestamp of the main video stream, and analyzing the trend of audience sentiment. Specifically, this includes: capturing newly added bullet comment text in real time, calculating a sentiment tendency index based on the frequency of positive and negative words; monitoring the time series of likes, calculating the instantaneous growth rate based on the change in likes during the corresponding time period of the sentiment tendency index, and determining a surge in interaction popularity; identifying the focus of audience discussion based on the comment content during the period of surge in interaction popularity and performing topic clustering of social media comments; and constructing an audience resonance index based on the sentiment tendency index, the instantaneous growth rate, and the frequency of mentioning the current speaker's name.
[0031] This step establishes a perceptual channel between the directing system and external audience feedback, achieving a leap from one-way output to two-way interaction. By quantitatively analyzing real-time audience reactions, the system can verify the dissemination effect of the current visual content and identify high-value segments that truly resonate. The combination of emotional inclination and interaction intensity effectively distinguishes between short-lived activity and deep identification, avoiding misjudging popularity based solely on quantity. The extraction of discussion focus reveals the intrinsic dimensions of content influence, providing a basis for optimizing subsequent shot strategies. The construction of an audience resonance index enables the system to perceive social influence, incorporating external feedback into the evaluation loop and improving the public relevance and dissemination effectiveness of directing decisions.
[0032] Furthermore, the above method also includes: associating the contextual information of key scenes with the emotional trends of the audience and writing it into the event log, specifically including: writing the complete contextual information of key scenes into the event log; performing pattern mining on historical logs to find frequently co-occurring scene-shot combinations; establishing individual performance profiles, recording historical behavioral characteristics, and then generating a director's quality report.
[0033] This step enables the structured accumulation and continuous optimization of directing experience, giving the system long-term learning and adaptability. By associating and archiving scene features with audience feedback, a traceable decision-making knowledge base is formed, providing a reference for similar situations in the future. Mining frequent patterns can uncover highly effective shot strategies, improving the accuracy and efficiency of future responses. The establishment of individual performance profiles enhances the personalized understanding of regular participants, accelerating role recognition and behavior prediction. The generation of quality reports not only supports manual debriefing but also provides quantitative indicators for overall system performance evaluation, driving the continuous evolution of directing logic in practice and achieving a leap from passive response to proactive optimization.
[0034] On the other hand, this invention proposes an AI-powered broadcast assistant system capable of implementing intelligent interactive effects, such as... Figure 2 As shown, it includes: The spatiotemporal alignment module is used to perform synchronous access to multiple audio and video signals and spatiotemporal alignment, generating a synchronized audio and video stream set and spatial mapping parameters; The semantic feature extraction module is used to perform hierarchical parsing and semantic feature extraction of audiovisual content based on the synchronized audio and video stream set and spatial mapping parameters, forming an individual behavior profile. The behavioral intent inference module is used to identify social functional roles and predict behavioral intent based on individual behavioral profiles. The key scenario identification module is used to construct a scenario salience scoring system based on social functional roles and behavioral intentions to identify key scenarios; The priority sorting module is used to generate candidate shot schemes based on key scenes and sort them by priority. The dynamic adjustment module is used to execute shot schemes sorted by priority, and to perform shot switching and dynamic composition adjustments.
[0035] In addition, the modules in the above system also perform other steps in implementing the aforementioned AI-powered live broadcast assistant method capable of intelligent interactive effects, as follows: Step 1: Synchronous access and spatiotemporal alignment of multiple audio and video signals This step aims to integrate audio and video signals from different physical locations or devices into the processing environment and perform consistency calibration between the time axis and the spatial reference frame, establishing a unified perceptual foundation for subsequent joint analysis. Since each signal source may use an independent clock, experience transmission delays, or have different frame rates, failure to perform alignment processing will lead to temporal misalignment or spatial mismatch in subsequent analysis, affecting the accuracy of judgments.
[0036] 1.1 Access audio and video streams from N independent devices, where N≥3, with each device deployed in a different location within the broadcast space, covering the speaking area, audience area, and close-up area; record the original timestamp of each signal during the access process. (i is the signal source number), and an independent data channel is established to ensure that the original information is preserved without loss during transmission; the timestamp comes from the device's local clock, and its accuracy is affected by network jitter and hardware differences, so it cannot be directly used for cross-source comparison.
[0037] 1.2 Generate global standard time based on the master clock source The time source is provided by a highly stable time synchronization server, with an error controlled within ±1ms; the original timestamp of each signal is... Mapped to Coordinate system, to obtain the corrected timestamp ,in The delay compensation for the i-th signal is estimated by measuring the network round-trip time delay and using the sliding window averaging method; this mapping process makes all signals comparable in the time dimension.
[0038] 1.3 Align the acquisition times of each video frame, selecting the main view camera as the reference frame source, with a frame rate of The remaining video streams are arranged as follows: Resampling is performed, and frames are inserted or deleted on the timeline using bilinear interpolation to achieve frame-level synchronization across all video streams; the synchronized video sequence is denoted as... Where M is the number of frames per unit time, and all i correspond to All represent visual images at the same physical moment.
[0039] 1.4 Establish spatial registration relationships and extract the pixel coordinates of corresponding markers in each video stream using multiple pre-set fixed markers (such as corner reflectors or QR code stickers) within the studio. (j is the marker point number), and compared with the known three-dimensional spatial coordinates. Perform matching and solve for the projection transformation matrix of each camera. ,satisfy This allows for spatial mapping between different viewpoints, enabling the same object in different scenes to be located in a unified coordinate system, laying the foundation for subsequent cross-view tracking.
[0040] This step addresses the heterogeneity of multi-source signals across time and space by introducing a global time reference and spatial reference system. Time alignment ensures event synchronization between sound and image, and between different viewpoints, preventing misjudgments due to delays. Spatial alignment enables target position mapping across viewpoints, allowing the system to track character movement paths in three-dimensional space. For example, when a speaker moves out of the frame from the main viewpoint, the system can predict their position in the side viewpoint based on the spatial transformation matrix, thus maintaining tracking continuity. The output of this stage—the synchronized audio and video stream set and the spatial mapping parameters—forms the input basis for the next stage of content analysis.
[0041] Step Two: Layered Analysis and Semantic Feature Extraction of Audiovisual Content After achieving spatiotemporal unification of the signals, this step involves hierarchical content deconstruction of each audio and video stream, extracting multimodal feature vectors encompassing speech, facial expressions, body movements, and environmental atmosphere to form a preliminary semantic description of the on-site activities. This process does not rely on preset rules but rather gradually builds an understanding of individual behavior and group state through continuous perceptual operations.
[0042] 2.1 In each frame of video To detect the presence of face regions, a face localization method based on a cascaded classifier is used to scan the image and identify rectangular regions that conform to the geometric structure of a face. And calculate its relative area percentage in the image. Where W and H are the image width and height, respectively; if (If the threshold is set to 0.05), then the person is determined to be in close-up view and has high potential for attracting attention; this operation is performed frame by frame to form a sequence of facial presence. This is used for subsequent behavioral trend analysis.
[0043] 2.2 Expression state determination is performed on the detected facial regions, and the coordinate positions of key facial points (such as the corners of the eyes, corners of the mouth, and center of the eyebrows) are extracted. The rate of change of Euclidean distance between each point is calculated. Combined with eyelid opening and closing With mouth opening Constructing a dynamic expression index The weight Based on historical data statistics; when When the value exceeds a preset range, it is marked as an emotional state such as "excitement", "surprise" or "focus". This mark is linked to the corresponding timestamp to form an emotional fluctuation record.
[0044] 2.3 The audio stream is processed by source separation. Blind source separation technology is used to decompose the mixed audio into several independent channels and identify the dominant speech signal. The system performs voice activity detection (VAD) to determine whether there is valid speech at a given moment; it then extracts the fundamental frequency of the speech. Speech rate (Number of syllables per unit time) and energy intensity , forming speech feature triples ;when When the speaker is active, the system determines that the speaker is in an active speaking state. This state is then matched spatiotemporally with the facial area in the video to confirm the identity of the current speaker.
[0045] 2.4 The above visual and auditory features are fused together, and for each identified individual, a time window is established. Internal behavioral profile ,in Indicates spatial location, Indicates the presence of the face. Indicates facial expression status. This indicates the level of voice activity; the profile is updated every 50ms, forming a continuous stream of the person's state, which serves as the original basis for the next stage of role recognition and behavior inference.
[0046] This step transforms the raw signal into semantic features, elevating low-level pixel and waveform data into a readable description of a person's state. Through layered processing, the system can distinguish between background noise and valid information; for example, in a multi-person conversation scenario, only those speaking and facing the audience are assigned a high activity score. (Expression Index) The design reflects the importance of nonverbal communication; when someone suddenly widens their eyes and raises their voice, their... A significant increase in the value suggests a potential emotional climax. Matching voice and vision ensures accurate attribution of the speaker, avoiding misattribution of sound to stationary figures in the image. This stage outputs an individual behavioral profile. This becomes the basis for subsequent judgments about their social roles and behavioral intentions.
[0047] Step 3: Dynamic Role Recognition and Behavioral Intent Inference Based on the individual behavioral profiles generated in the previous step, this step further analyzes the social functional roles of each participant (such as speaker, questioner, audience member, etc.) and predicts their short-term behavioral intentions (such as about to speak, preparing to leave, initiating interaction, etc.), thereby providing contextual basis for the director's decisions. This process relies on a combination of long-term observation of behavioral patterns and short-term mutation detection.
[0048] 3.1 Calculate the cumulative speaking time of each individual within a continuous 5-minute period. With average speech intensity Calculate its dominance index ,in and These are empirical proportionality coefficients, assigned values of 0.6 and 0.4 respectively; when If the value is set to 0.7 and lasts for more than 2 minutes, it is marked as the "main speaker". This speaker is usually located in the center area and has a stable vocal behavior.
[0049] 3.2 Monitor sudden changes in individual behavior and define the rate of state change. This indicates the degree of drastic change in the behavioral profile within the time window T; when A sudden increase (more than twice the standard deviation of the mean over the past 10 seconds), accompanied by a change in facial orientation (from sideways to facing the camera) and vocal activation, indicates that the individual has the intention to "initiate interaction"; this type of behavior often occurs when an audience member raises their hand to ask a question or a guest suddenly interrupts.
[0050] 3.3 Analyze the group attention distribution and calculate the gaze direction vector of all individuals at the current moment. And to count the proportion that points to a specific person. ;like If the person is speaking, their "focus character" attribute is strengthened; conversely, if someone is speaking but no one is watching, the priority of switching to a close-up is reduced, reflecting actual influence rather than simply speaking time.
[0051] 3.4 Establish a role state transition diagram, record the frequency and duration of each individual's transitions between the four states of "speaking," "interacting," "listening," and "irrelevant," and use a Markov chain model to predict the probability of their next state. For example, a person who is currently in a "listening" state but whose gaze frequently sweeps towards the podium is significantly more likely to switch to an "interactive" state than someone who is constantly looking down to take notes; this prediction result is used to prepare contingency plans for camera transitions in advance.
[0052] This step goes beyond simply capturing surface behavior and delves into understanding social roles and underlying motivations. Dominance Index It comprehensively considers both the volume of speech and the intensity of voice to avoid misjudging the speaker solely based on volume (such as the high volume of applause from the audience). Rate of State Change As a mutation detection tool, it can capture individuals who have made preparatory actions but have not yet spoken, enabling "predictive" broadcasting. Group attention analysis introduces the "collective gaze" phenomenon from social psychology, where the convergence point of most people's gaze often represents the current information focus; this principle is used to verify the importance of individual behaviors. Role transfer prediction endows the system with forward-looking capabilities, enabling it not only to respond to current events but also to anticipate upcoming interaction transitions. The role labels and intent predictions output at this stage become the core basis for identifying key scenarios in the next stage.
[0053] Step 4: Multi-dimensional comprehensive judgment of key scenarios Based on the understanding of individual roles and behavioral intentions, this step integrates information from multiple dimensions, including temporal continuity, spatial concentration, emotional intensity, and interaction potential, to construct a scene salience scoring system. This system identifies key segments worthy of the director's attention, such as important speeches, emotional outbursts, or moments that resonate with the audience. This judgment process employs a weighted fusion strategy to ensure that different types of memorable moments are effectively captured.
[0054] 4.1 Define the time continuity factor This is used to measure the length of time a character remains in a highly active state; let's say the character was in a highly active state recently. The cumulative duration within seconds that meets the criteria of active voice and facial expression index above the baseline is ,but ;when At that time, it was believed that his speech was coherent and complete, and suitable for being presented as the main screen for a long time.
[0055] 4.2 Calculation of Spatial Concentration Factor This reflects the degree of aggregation of multiple active individuals in physical space; let the set of individuals currently identified as having "interaction intent" be... Its position coordinates are Calculate its centroid Find the average distance from each point to the centroid. ;like (Assuming it's 1.5 meters), then ,otherwise High spatial concentration usually corresponds to group discussions or close-up exchanges, making it a good time to switch to a wide-angle or overhead perspective.
[0056] 4.3 Introducing an Emotion Intensity Factor Quantify the overall emotional fluctuation level at the scene; and the expression index of all individuals. Perform a weighted average, with the weights determined by the percentage of the image displayed. ,have to Simultaneously monitor the short-time variance of audio energy. This reflects the intensity of sound fluctuations; ultimately ,in The largest historical variance; when At times, these are marked as "high-emotional scenes," such as applause, laughter, or heated debates.
[0057] 4.4 Constructing a Scenario Saliency Score ,in The interaction potential factor (derived from the mutation detection results in step 3.2; if ≥1 person shows an intention to interact, then...) (otherwise it is 0), each weight ;when When the value is set to 0.75, the current time period is determined to be a "critical scene", triggering the director's decision engine to enter the preparation state and record the start and end times of the scene and the list of people involved.
[0058] This step represents a leap from individual behavior to the overall scenario, capturing different types of important moments through a multi-dimensional scoring mechanism. (Time continuity factor) This ensures respect for complete speaking segments and avoids frequent switching during the speech; spatial concentration This identifies the physical form of group interaction, providing a basis for shot composition. Emotional intensity. Combining visual and auditory cues allows us to detect nonverbal emotional climaxes; for example, a sudden burst of laughter from the audience, even without a clear speaker, can still be detected as a surge in emotional intensity. The final saliency score... Using a linear weighting method makes it easy to adjust the priority of various scenarios; for example, it can increase the priority in interview programs. Weighting increases in debates. Impact. The key scene markers output at this stage include not only "when" it occurs, but also "who" is involved and "what type" of event, providing structured input for subsequent automatic directing instruction generation. This determination will directly drive the shot selection and switching logic in the fifth step.
[0059] Step 5: Adaptive Generation and Prioritization of Lens Strategies After identifying key scenes, this step generates a set of candidate shot schemes based on scene type, spatial distribution of participants, and behavioral characteristics. These schemes are then dynamically prioritized to ensure that the output footage conforms to visual narrative logic while highlighting the most relevant content. This process combines spatial relationship analysis with behavioral context understanding to create a context-aware director's decision-making sequence.
[0060] 5.1 Based on the key scene types marked in step four, call the preset shot logic template library to match the corresponding initial shot combination; for example, when The value is mainly composed of and When the system is activated, if it is determined to be a "critical climax of a key speech," the "speaker close-up + audience reaction" dual-track mode will be enabled; if and If the result is "Group Interaction Startup", it will switch to "Wide Angle Overhead Shot + Focus Person Tracking" mode; each template includes the number of shots, perspective type, duration suggestions and switching transition methods, which serve as the basic framework for candidate solutions.
[0061] 5.2 Calculate the coverage effectiveness of each candidate shot, defined as a comprehensive index of the number of key figures that the shot can present and the image clarity; for the k-th candidate shot... Its coverage set Includes all individuals visible and active from this perspective, and calculates the effective coverage. ,in The area of person i in the image. Rate its activity level (from step three) );like If set to 0.4, then the shot will be eliminated to avoid selecting empty shots or perspectives with only a static background.
[0062] 5.3 Introduce visual smoothness constraints to prevent frequent switching or abrupt changes in perspective within a short period of time; record the last used shot number. Calculate the current candidate shot Difference in perspective ,in It is the horizontal deflection angle. The vertical pitch angle; if If the angle is set to 60° and the duration of the previous shot is less than 3 seconds, a delay penalty is applied to the candidate shot, and its priority is reduced by one level to maintain viewing comfort.
[0063] 5.4 The remaining candidate shots are evaluated based on their overall scores. Sort, where The attenuation term (value 0 or 0.5) is introduced due to the sudden change in perspective; the shot with the highest score is marked as the "preferred output", and the rest are listed as backup options in order; the sorting result, along with its activation conditions (such as duration, exit trigger point), is written into the director's execution queue for use in the next stage.
[0064] This step transforms the process from "recognition" to "decision-making," converting abstract scene judgments into concrete visual presentation instructions. The introduction of shot templates ensures appropriate visual expression in different situations; for example, it automatically adds audience reaction shots during emotional outbursts, enhancing the communicative impact. (Coverage effectiveness) The design avoids the mistake of "form over content" in transitions; for example, it avoids choosing to only film the ceiling from a low angle during multi-person discussions. Visual fluidity constraints simulate the smooth transitions of a human director, preventing machine-like jumps that create a disjointed viewing experience. The final overall score... The system considers not only the current information density but also the viewing experience, ensuring a stable pace while pursuing rich content. The sorted shot list output at this stage becomes the direct source of instructions for execution control in the sixth step.
[0065] Step Six: Real-time Camera Switching and Dynamic Composition Adjustment After determining the preferred camera setup, this step involves performing the actual signal switching and dynamically fine-tuning the camera parameters based on the subject's movements and the content of the shot to ensure the target is always in the optimal visual position. This process integrates remote control commands with local feedback adjustments, forming a closed-loop control mechanism.
[0066] 6.1 Send control commands to the physical camera corresponding to the preferred lens, including the target focal length. Gimbal rotation angle and aperture value The instructions are encapsulated and encrypted before transmission via the RTSP protocol; upon receiving the instructions, the camera performs adjustments and simultaneously sends back a confirmation signal. If not received within 200ms If so, the backup lens will be activated and the device malfunction log will be recorded.
[0067] 6.2 During the shot transition, the screen transition effect is activated simultaneously, and the transition method is selected according to the semantic relationship between the shots: if the two shots focus on the same person (such as from a wide shot to a close-up), a "fade in" fade out gradient is used; if the perspective jump is involved (such as from the podium to the audience), a "swipe" or "push" effect is used; the transition duration is fixed at 0.8 seconds, and the inter-frame interpolation is completed by the video compositor in the background to ensure that the screen transition is natural and tear-free.
[0068] 6.3 After the switch is completed, continuously monitor the positional shift of the main character in the image; let the position of the target character's facial center in the image coordinate system be... Define offset ,in The ideal composition center (usually the golden ratio point); when When the image is set to 15% of the diagonal, a fine-tuning command is generated to drive the gimbal to slowly compensate, so that the subject returns to the visual center of gravity.
[0069] 6.4 Automatically adjust exposure parameters based on changes in ambient light and capture the brightness histogram of the current image. Calculate the average brightness ;like (Dark light) or If overexposed, adjust the gain proportionally. With shutter speed ,make Approaching 120; this adjustment is performed every 2 seconds to avoid frequent flickering and ensure screen stability.
[0070] This step completes the implementation from decision-making to execution, realizing the physical presentation of automated broadcasting. Reliable transmission of instructions ensures timely system response, and the ACK mechanism enhances fault tolerance. The selection of transition effects is based on the relevance of the preceding and following content, making the switching not only a technical operation but also part of the narrative language; for example, "fade in" implies continuation, and "swipe" indicates a scene change. Dynamic composition adjustment solves the problem of characters going out of frame due to movement, maintaining professional-grade image quality without manual intervention. Automatic exposure adjustment handles changes in studio lighting (such as spotlight movement), preventing sudden brightness changes from affecting the viewing experience. The output of this stage—a stable, clear, and well-composed real-time video stream—constitutes the main signal for the final external release and also provides an input source for subsequent interactive feedback analysis.
[0071] Step Seven: Collecting Audience Feedback Signals and Analyzing Emotional Resonance While the main video stream is being output, this step collects external audience reaction data through multiple channels, including the bullet comments on the live streaming platform, the rate of change in the number of likes, and real-time comments on social media. This analysis helps to understand public sentiment trends and determine whether the current video content resonates widely, providing external validation for optimizing the directing strategy.
[0072] 7.1 Access the open interface of the live streaming platform to capture newly added bullet screen text every 5 seconds. The text was analyzed for sentiment polarity; a dictionary-based matching method was used to count the frequency of positive words (such as "wonderful", "agree", "moved") and negative words (such as "boring", "can't hear", "too fast"), and the sentiment tendency index was calculated. ,in This represents the total number of bullet comments; if This is marked as the "positive feedback concentration period".
[0073] 7.2 Time series of monitoring likes Calculate its instantaneous growth rate ,in ;when When the interaction intensity exceeds 1.8 times the average of the past minute, it is judged as a "surge in interaction intensity". The time is then aligned with the timestamp of the main video stream, and the dominant shot types and scene characteristics in the previous 30 seconds are searched in reverse to establish a "high resonance content profile".
[0074] 7.3 Perform topic clustering on public comments on social media, extract high-frequency noun and verb combinations, and identify the focus of audience discussion; for example, if "expert," "opinion," and "profound" frequently co-occur, it indicates that the depth of the content is recognized; if "visual," "switching," and "chaos" appear in a concentrated manner, it suggests that there may be a problem with the director's pacing; this type of information is used for long-term strategy optimization and does not directly affect real-time switching.
[0075] 7.4 Constructing an Audience Resonance Index ,in The frequency with which the current speaker's name is mentioned in the comments. ;when When the value is set to 0.7, the system considers the current content to have successfully resonated with the group, marks this segment as a "segment with high dissemination potential", stores it in the highlights library, and prioritizes its recommendation in subsequent replays.
[0076] This step breaks away from the traditional one-way output model of broadcasting, introducing external feedback to form a closed-loop optimization. Bullet comments and likes serve as real-time indicators, quickly reflecting audience emotional fluctuations. For example, when a guest utters a memorable quote, the bullet comments instantly flood the screen with "famous quotes," and the like rate surges, allowing the system to confirm that moment as a true "highlight." Emotional Tendency Index It provides quantitative sentiment assessments, avoiding the reliance on view counts alone to judge content quality. Resonance Index The system's design enables it to identify not only "technically critical scenarios" but also "emotionally impactful moments." For example, an impromptu speech might not be initially identified as the speaker's climax but could be re-evaluated due to its strong emotional resonance. The output of this stage provides the basis for learning and adaptation in the eighth step.
[0077] Step 8: Memory Storage and Experience Transfer of Directing Behavior The final step in this embodiment is to structurally archive the entire process and extract reusable experience patterns to optimize processing efficiency for similar scenarios in the future. This process does not rely on parameter updates, but rather achieves knowledge accumulation and transfer through the summarization and correlation analysis of event logs.
[0078] 8.1 Write complete context information for each key scenario to the event log, including timestamp, involved personnel, and scenario type. Value composition, selected shot sequence, audience resonance index And the final dissemination effect (such as the number of times it is viewed); each log is stored in structured JSON format to form a searchable historical database.
[0079] 8.2 Perform pattern mining on historical logs to find frequently co-occurring "scene-shot" combinations; for example, if a "sudden increase in emotional intensity" occurs, using a "close-up + slow motion" strategy yields high results in 80% of cases. If the value is positive, this combination is marked as a "high confidence strategy" and will be called first in subsequent similar scenarios to reduce decision delay.
[0080] 8.3 Establish individual performance profiles to record the historical behavioral characteristics of each resident participant, such as average speaking pace, common gestures, and emotional expression. When the person reappears, the system can preload their behavioral model to speed up role recognition and predict their possible behavioral paths, thereby improving the foresight of the response.
[0081] 8.4 After the daily tasks are completed, a director quality report is generated, which summarizes the key scene coverage, camera transition smoothness, and number of peak audience resonances for the day, for the operations staff to review; at the same time, suggested adjustments are output, such as "the response latency of a certain type of interactive scene is too high, and it is recommended to add a side view preset position" to achieve human-machine collaborative optimization.
[0082] This step endows the system with the potential for continuous evolution, enabling it to not only perform well in single tasks but also continuously improve over long-term operation. The structured storage of event logs makes each directorial action a traceable experience asset, rather than a one-time consumption. The extraction of frequent patterns achieves a leap "from data to patterns," for example, discovering that the "silence followed by an outburst" speaking pattern is often accompanied by high resonance; the system can prepare close-up shots as soon as initial silence is detected. The creation of individual profiles enhances personalized service capabilities, especially suitable for series programs with fixed guests. The final quality report not only serves machine optimization but also provides data support for human review, forming a virtuous cycle.
[0083] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for an AI-powered live broadcast assistant capable of implementing intelligent interactive effects, characterized in that, Includes the following steps: Simultaneously access multiple audio and video signals and perform spatiotemporal alignment to generate a synchronized audio and video stream set and spatial mapping parameters; Based on the synchronized audio and video stream set and spatial mapping parameters, the audiovisual content is subjected to hierarchical analysis and semantic feature extraction to form an individual behavior profile; Based on the individual behavioral profile, social functional roles are identified and behavioral intentions are predicted. Based on the aforementioned social functional roles and behavioral intentions, a scenario salience scoring system is constructed to identify key scenarios; Based on the key scenarios, candidate shot schemes are generated and prioritized. Shot switching and dynamic composition adjustment are then performed using the prioritized shot schemes.
2. The AI-powered broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The synchronous access to multiple audio and video signals and spatiotemporal alignment includes: Receive audio and video streams from N independent devices and record the original timestamps; Generate a global standard time based on the master clock source, and map the original timestamp to the global standard time coordinate system; Align the acquisition times of each video frame to achieve frame-level synchronization of all video streams, then establish spatial registration relationships and use preset marker points to solve the projection transformation matrix.
3. The AI-powered live broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The layered analysis and semantic feature extraction of audiovisual content includes: Detect face regions in each frame of video and calculate the proportion of face area; Perform facial expression determination on the detected facial regions; The audio stream is subjected to sound source separation and speech activity detection, and then visual and auditory features are fused to create an individual behavior profile.
4. The AI-powered broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The identification of social functional roles and prediction of behavioral intentions includes: The cumulative speaking time and average speech intensity of an individual within a continuous 5-minute period are statistically analyzed, and a dominance index is calculated based on the cumulative speaking time and average speech intensity. Monitor abrupt changes in individual behavior and define the rate of state change based on the changing trend of the dominant index; Analyze the distribution of group attention and statistically determine the proportion of gazes directed at a specific person based on the rate of state change and gaze direction vector. Establish a character state transition graph and record the frequency and duration of state transitions based on the gaze direction ratio and state change rate.
5. The AI-powered broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The key scenarios identified include: Calculate the time continuity factor based on the duration of sustained high activity in social functional roles; Calculate the spatial concentration factor based on the location coordinates of individuals with interactive intentions; Emotion intensity factor is calculated based on individual facial expression index and audio energy variance. The saliency score of a scenario is calculated based on factors such as temporal continuity, spatial concentration, emotional intensity, and interaction potential to identify key scenarios.
6. The AI-powered broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The generation of candidate shot schemes and their priority ranking include: Based on the key scene type, call the shot logic template library to match the initial shot combination based on the key scene type; Coverage effectiveness is calculated based on the coverage set and activity score of each shot in the initial shot set; Whether to apply a delay penalty is determined based on the difference in perspective between the previously used lens number and the current candidate lens; Candidate shots are comprehensively scored and ranked based on coverage effectiveness, scene salience score, and delay penalty results.
7. The AI-powered live broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The execution of camera switching and dynamic composition adjustment includes: Send control commands to the physical camera corresponding to the preferred lens; During camera transitions, a screen transition effect is activated simultaneously. After the switch is completed, the system continuously monitors the positional shift of the main characters in the scene and automatically adjusts the exposure parameters according to changes in ambient lighting.
8. The AI-powered broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The method further includes: collecting bullet comments and likes data from the live streaming platform, aligning them based on the timestamp of the main video stream, and analyzing the trend of audience sentiment, specifically including: Real-time capture of newly added bullet screen text, and calculation of sentiment index based on the frequency of positive and negative words; Monitor the time series of likes, calculate the instantaneous growth rate based on the change in likes within the corresponding time period of the sentiment index, and determine the surge in interaction popularity; Based on the comment content during the periods of surge in interactive popularity, the focus of audience discussion is identified, and the social media comments are clustered by topic. An audience resonance index is constructed based on the aforementioned sentiment tendency index, instantaneous growth rate, and frequency of mentioning the current speaker's name.
9. The AI-powered live broadcast assistant method for implementing intelligent interactive effects according to claim 1, characterized in that: The method further includes: associating the contextual information of key scenes with the emotional trends of the audience and writing it into the event log, specifically including: writing the complete contextual information of key scenes into the event log; performing pattern mining on historical logs to find frequently co-occurring scene-shot combinations; establishing individual performance profiles, recording historical behavioral characteristics, and generating a director's quality report.
10. An AI-powered broadcast assistant system for implementing the method as described in any one of claims 1-9, characterized in that, include: The spatiotemporal alignment module is used to perform synchronous access to multiple audio and video signals and spatiotemporal alignment, generating a synchronized audio and video stream set and spatial mapping parameters; The semantic feature extraction module is used to perform hierarchical analysis and semantic feature extraction of audiovisual content based on the synchronized audio and video stream set and spatial mapping parameters, and form an individual behavior profile. The behavioral intent inference module is used to identify social functional roles and predict behavioral intent based on the individual's behavioral profile. The key scenario determination module is used to construct a scenario salience scoring system based on the social functional roles and behavioral intentions to identify key scenarios; The priority sorting module is used to generate candidate shot schemes based on the key scenes and sort them by priority; The dynamic adjustment module is used to execute the shot scheme sorted by priority, and to perform shot switching and dynamic composition adjustment.