A method for producing and releasing a fusion media product

CN122845819APending Publication Date: 2026-09-29BEIJING HUAYI ZHONGXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611049277.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0002]在气象融媒体与灾害直播领域,当前技术难以满足实时、精准、沉浸式的信息呈现与交互需求

Benefits of technology

通过多模态数据感知与智能决策,实现了增强内容的自动化、精准化生产与发布,能实时识别关键瞬间与观众兴趣,自动生成并融合数据可视化、虚拟播报等增强信息,极大提升了直播的信息维度、互动性与沉浸感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845819A_ABST
    Figure CN122845819A_ABST
Patent Text Reader

Abstract

This application discloses a method for producing and publishing converged media products. After the live stream begins, the method simultaneously accesses and processes the main live stream signal stream, external business data stream, and audience interaction data stream, extracts visual, audio, and text features, performs initial classification of the live stream content, and determines the core points of audience attention. Based on preset rules and a real-time analysis model, monitoring is performed, and when business data exceeds a threshold, specific keywords are identified, or a high-attention moment is detected, an enhanced content generation instruction is triggered. The content generation module is called in parallel to fill relevant data into templates to generate infographics or drive 3D models and virtual anchors to generate enhanced content. Through low-latency fusion technology, the enhanced content is rendered and published in real-time using traditional video stream overlay or augmented reality space registration, achieving automated and precise production and presentation of enhanced content, significantly improving the interactivity, immersion, and information delivery effect of the live stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of converged communication technology, and in particular to a method for producing and publishing converged media products. Background Technology

[0002] In the field of meteorological media convergence and disaster live streaming, current technologies are insufficient to meet the demands for real-time, accurate, and immersive information presentation and interaction. Traditional meteorological live streams or disaster reports still primarily rely on presenters' narration and static text, images, and video interludes, failing to intelligently and dynamically integrate real-time dynamic meteorological data with actual disaster site videos. Viewers find it difficult to intuitively understand the evolution of complex meteorological systems, the scope of disaster impact, and risk levels from traditional live streams. Existing technologies also cannot accurately and stably anchor and dynamically track key meteorological entities such as typhoon eyes, rainband boundaries, and flood inundation lines in live stream footage. Furthermore, they cannot generate and overlay enhanced information such as evacuation routes, shelter locations, and real-time risk warnings in real time based on disaster development and public interaction feedback, resulting in low information transmission efficiency and limited public perception and decision-making support.

[0003] Furthermore, while augmented reality (AR) technology offers new avenues for the three-dimensional visualization of meteorological information and the simulation of disaster scenarios, its application in real-world, large-scale, and dynamically changing disaster live broadcasts faces significant challenges. First, meteorological elements and disaster entities are dynamically evolving, and existing AR spatial registration technologies struggle to achieve stable and accurate matching between virtual information and the real geographical environment in complex natural scenes, easily resulting in drift and misalignment. Second, the production of augmented content heavily relies on manual offline creation, making it impossible to automatically drive the generation and updating of three-dimensional meteorological fields based on real-time collected observation data, forecast model outputs, and live video streams. It also makes it difficult to achieve intelligent information prompts based on entity behavior predictions. These shortcomings make it difficult for existing technologies to support a new generation of meteorological converged media live broadcast platform with real-time data-driven, intelligent interactive, and immersive presentation capabilities, facilitating public science education, emergency command, and media broadcasting. Summary of the Invention

[0004] This application provides a method for producing and publishing converged media products, which enables seamless continuation of information association and adaptive adjustment of visual presentation in complex and dynamic scenarios.

[0005] This application provides a method for producing and publishing converged media products, including: S1: At the same time as the live broadcast starts, the main live broadcast signal stream is accessed, its visual and audio features are extracted in real time, and the live broadcast content is initially classified based on these features. S2 synchronously accesses external business data streams and audience interaction data streams, performs sentiment analysis, topic clustering, and keyword extraction on interactive texts, and combines the scene classification results of S1 to comprehensively determine the audience's core concerns and the core theme of the live broadcast. S3 monitors based on preset rules and real-time analysis models. When business data exceeds the threshold, a specific keyword combination is identified, or a high-attention moment is detected, corresponding enhanced content generation instructions are triggered. S4, based on the enhanced instructions generated by S3, calls the content generation module in parallel, retrieves the corresponding resources from the template library and resource library, fills the relevant data into the visualization template to generate information charts, or drives the 3D model and virtual anchor to generate the corresponding broadcast content. S5, for augmented reality deployment, calculates the precise spatial registration parameters of virtual augmented content based on spatial perception data and spatial anchor point information in the instructions, for AR clients to perform localized rendering and overlay; S6 will overlay and blend enhancement layers on the live video screen in real time according to the screen position and time specified by the instruction. The final video stream after blending is encoded and pushed to the content delivery network. For traditional video publishing, S7 uses the various enhanced content generated by S4 as layers, which are then overlaid onto the main live video stream in real time according to the command parameters. After encoding, the content is pushed through the content delivery network.

[0006] Preferably, the extraction of visual and audio features specifically includes: visual features refer to scene and object recognition, person detection and recognition, text information extraction and image attribute analysis; audio features refer to speech-to-text conversion, keyword and topic detection, voiceprint and speaker recognition and audio event detection.

[0007] Preferably, the localized rendering and overlay for the AR client includes: the AR client using its own environmental tracking capabilities to maintain a continuous understanding of the real-world space; accurately drawing virtual content on a specified three-dimensional spatial location based on the received spatial transformation matrix; and fusing and overlaying the drawn virtual content layer with real-world video frames captured by the device's camera in real time.

[0008] Preferably, the triggering of the corresponding enhanced content generation instruction specifically includes: S31 processes the live data stream in real time to complete instance perception, pixel-level segmentation and temporary ID tracking, and extracts entity instances and their outlines from the live scene and the surrounding shooting environment and uploads them to the cloud. S32 refines the identification and 3D reconstruction of entities in the live broadcast by using data uploaded from the cloud, and constructs and continuously updates a dynamic spatiotemporal map describing entities and their relationships. S33, the enhanced command is matched and parsed with the spatiotemporal graph to derive the target entity, and an enhanced anchoring command describing the relative spatial relationship is generated and bound to the entity; where the target entity is a specific object instance that is uniquely identified and continuously tracked in the constructed and maintained scene entity spatiotemporal graph; S34: The client dynamically calculates the pose of the virtual content for each frame based on the real-time status of the entity synchronized from the cloud and the anchoring instructions, and performs stable following and rendering of the dynamic entity based on the virtual content.

[0009] Preferably, the extraction of entity instances and their outlines from the live streaming scene and the surrounding shooting environment specifically includes: receiving the original video frame sequence from the live streaming main signal stream and preprocessing the video frames; inputting the preprocessed video frames into a lightweight neural network model deployed on the edge; the model performs instance perception and pixel-level segmentation in parallel to identify and distinguish each independent object instance in the image, assigns a temporary ID to each instance, and outputs a pixel-level mask for each instance to outline its contour boundary; cross-frame association between the recognition result of the current frame and the results of historical frames, maintaining the continuity of the temporary ID of the same instance through appearance feature and motion trajectory matching, and handling the appearance of new entities and the disappearance of old entities; and combining the edge-side spatial perception data to associate the two-dimensional position of the entity in the image with its approximate position in the three-dimensional scene.

[0010] Preferably, the nodes in the dynamic spatiotemporal graph represent each tracked entity instance, and the edges represent the dynamic relationships between entities.

[0011] Preferably, the step of parsing the enhanced instructions with the spatiotemporal graph to extract the target entity includes: S331: Real-time acquisition of high-frequency motion sequences of target entities, fine-grained analysis of the video region where the target entity is located, and extraction of refined behavioral features beyond the bounding box; S332, based on the spatiotemporal graph at each moment, constructs a dynamic graph for neural network processing. The dynamic graph is input into the spatiotemporal graph neural network to simulate the transmission and coordination of intentions between entities, and outputs a set of probabilistic trajectories for each key entity in the future short period of time, as well as the discrete behavior label most likely to be executed at the corresponding moment. S333: Based on the prediction results, the client predefines dynamic parameterized anchoring rules and generation logic for the enhanced content. The client performs advance rendering based on the prediction sequence and smoothly corrects the predicted content through real-time observation.

[0012] Preferably, the behavioral features specifically involve: extracting video data of the target entity's location from the main live video stream based on the target entity's image range; running a pose estimation model, an action recognition model, and an appearance feature extraction model in parallel on the video data; wherein, the pose estimation model is used to extract the pose sequence of the target entity, the action recognition model is used to determine the type of micro-action it performs, and the appearance feature extraction model is used to encode its refined texture and appearance changes; and fusing the results of the pose estimation, action recognition, and appearance feature extraction to form refined features characterizing the dynamic behavior of the target entity.

[0013] Preferably, the step of outputting a set of probabilistic trajectories for each key entity in the short future includes: inputting a dynamic graph reflecting the entity and its relationships into a pre-trained spatiotemporal graph neural network; calculating and outputting a hidden state representing the overall situation of each key entity node through graph inference of the neural network; inputting the hidden state into a trajectory decoder; the trajectory decoder outputting a parameterized probability distribution defining the range of possible spatial locations of the entity at various discrete time points in the future; and generating multiple trajectory hypotheses representing the future movement possibilities of the entity based on the probability distribution.

[0014] Preferably, the real-time acquisition of the high-frequency motion sequence of the target entity specifically includes: real-time monitoring of the state continuity of the target entity; when a physical or visual split is detected, identifying and tracking each generated sub-entity, and establishing the derivation relationship between the sub-entity and the target entity; after the split event is confirmed, converting the type of the target entity in the spatiotemporal graph into a logical entity container, adding each sub-entity as a member entity of the container, and dynamically aggregating the states of all members to calculate the logical center and spatial range of the container in real time; automatically rebinding the enhancement content originally bound to the target entity to the logical container or its designated member entity; setting and monitoring the resolution conditions of the logical container, and when the conditions are met, smoothly migrating the enhancement content to the designated member entity or fading it out, and cleaning up the container structure in the graph.

[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages: Through multimodal data perception and intelligent decision-making, the production and release of enhanced content have been automated and precise. It can identify key moments and audience interests in real time, and automatically generate and integrate enhanced information such as data visualization and virtual broadcasting, which greatly improves the information dimension, interactivity and immersion of live broadcasts.

[0016] By constructing and updating the spatiotemporal map of scene entities in real time, augmented commands can be precisely anchored to specific entities, enabling virtual information to stably follow the movement of entities. This fundamentally solves the problem of content drift and misalignment in dynamic scenes, significantly improving the realism and accuracy of AR overlay.

[0017] Furthermore, the system incorporates short-term behavioral intent prediction capabilities, enabling the proactive generation of enhanced content that points to future actions. Delays are eliminated through advanced rendering and smoothing correction mechanisms, ensuring smooth and consistent display. For morphological changes such as entity decomposition and separation, an innovative logical container mechanism is employed to dynamically manage split sub-entities, ensuring intelligent rebinding of enhanced content. This achieves seamless continuation of information association and adaptive adjustment of visual presentation in complex dynamic scenarios. Attached Figure Description

[0018] Figure 1This is a flowchart illustrating a method for producing and publishing converged media products according to an embodiment of the present invention. Detailed Implementation

[0019] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terminology used herein includes any and all combinations of one or more of the associated listed items.

[0021] Example 1: Figure 1 This is a flowchart illustrating a method for producing and publishing converged media products according to an embodiment of the present invention.

[0022] like Figure 1 As shown, a method for producing and releasing converged media products includes the following steps: S1: At the same time as the live broadcast starts, the main live broadcast signal stream is accessed, and its visual and audio features are extracted in real time. Based on this, the live broadcast content is initially classified into scenes.

[0023] The visual features include scene and object recognition, identifying the main scenes and prominent objects in the image; person detection and recognition, detecting people appearing in the image and performing face recognition or identity association; text information extraction, using OCR technology to identify and extract text information such as overlaid subtitles, titles, slide text, or product labels in the image; and image attribute analysis, analyzing camera movement, image switching, and overall color and composition.

[0024] Audio features include: speech-to-text conversion, which converts live audio conversations into transcripts in real time; keyword and topic detection, which detects predefined or dynamically discovered keywords, technical terms, and core topics from the audio text in real time; voiceprint and speaker recognition, which distinguishes different speakers and may identify the identity of a specific speaker; and audio event detection, which identifies non-speech audio events, such as applause, cheers, specific background music, or sound effects.

[0025] S2 synchronously accesses external business data streams and audience interaction data streams, performs sentiment analysis, topic clustering, and keyword extraction on interactive texts, and combines the scene classification results from S1 to comprehensively determine the audience's core concerns and the core theme of the live broadcast.

[0026] Specifically, structured operational data from meteorological departments is acquired in real time via interfaces, such as satellite cloud imagery, radar echoes, automatic weather station observations (wind speed, rainfall), early warning signals, forecast conclusions, and disaster risk alerts. The data is accessed in JSON and other formats, and field parsing and timestamp alignment are performed. Interactive data from live streaming platforms, including bullet comments, likes, requests for help, and location reports, are also collected in real time. Text comments (bullet comments / reviews) are cleaned, segmented, and standardized, with particular attention paid to statements containing geographical locations, disaster descriptions, and requests for help.

[0027] Each cleaned interactive text is analyzed in real-time using a pre-trained sentiment analysis model, categorizing it into types such as concern, panic, seeking verification, neutrality, or request for help, and calculating the overall public sentiment distribution and evolution trend. A short text clustering algorithm dynamically groups incoming comments, automatically identifying multiple topic clusters currently being discussed by the public, such as areas with severe flooding, power outages, and shelter searches. The real-time popularity of each topic cluster is calculated and ranked based on the number of comments and interaction frequency. High-frequency words, core keywords, and named entities (specific place names, typhoon names, disaster types, infrastructure names) are extracted from the interactive text to quantify the public's focus and information needs.

[0028] The initial scenario classification results output from step S1 (live broadcast of typhoon landfall, on-site coverage of urban flooding) are used as the core context, and weights are assigned to business data and interactive analysis based on the scenario. For example, in the typhoon landfall scenario, the weights of descriptive entities extracted from interactions, such as roofs being blown off and windows being broken, as well as the typhoon eye location and maximum wind speed in the business data, will significantly increase. By comprehensively analyzing the interactive topic clustering results, extracted key entities, and changes in key fields in the business data (sudden surge in rainfall in a certain area, new warnings issued), the core public concerns and most urgent information needs at the current moment are determined. This core concern is then compared and verified with the initial classification from S1. For example, S1 initially classified it as the impact of heavy rain based on visuals and audio, while S2, through analysis of interactions and business data, confirms that the core theme is traffic disruption on XX Road in the urban area due to flooding, and the emergency drainage situation in surrounding residential areas. This achieves refinement of the initial classification and scenario focus. The final output is a structured description of the core theme of the live broadcast, which integrates the disaster scenario, the core focus (Typhoon Muifa, XX dam, specific affected communities), and the core issues (path changes, risk level, and rescue progress). S3 monitors based on preset rules and real-time analysis models. When business data exceeds a threshold, a specific keyword combination is identified, or a high-attention moment is detected, corresponding enhanced content generation instructions are triggered.

[0029] Among these, the pre-defined rules are developed in collaboration with meteorological and emergency management departments, solidifying key operational nodes for disaster prevention and mitigation into rules. For example, when a typhoon warning signal is upgraded, the cumulative rainfall in a certain area exceeds the flood threshold, a storm cell of a specific intensity appears in radar echoes, or an official disaster report is received, these are considered clearly defined nodes that must be triggered. Data from past live broadcasts of similar disasters is analyzed to identify the correlation between sudden surges in public interaction (comments, requests for help) and disaster events (the appearance of footage of dam breaches, landslides), and these highly correlated moments are also transformed into rules.

[0030] Real-time analysis models are trained for specific tasks in disaster scenarios. For example, using a large amount of labeled historical interaction data, a classification model is trained to recognize public intentions (disaster reporting, seeking help, seeking safety advice, verifying information) and emotional states (panic, anxiety). Simultaneously, the model can be trained to identify key stages in disaster evolution (rapid rise in floodwaters, sudden increase in wind force). The model integrates visual features from S1 (the extent of floodwaters in the image, the visible effect of wind) and analytical results from S2 (public focus, emotion) as input, and continuously optimizes and iterates based on actual results.

[0031] Triggering the corresponding enhanced content generation command occurs during the live stream planning phase. A series of enhanced content types are designed in advance, and a command template is created for each type. The template defines the parameters required for that type of enhanced content, including chart type, data source fields, virtual human broadcast text template, screen display position, and duration. When one (or more) preset rules are triggered, or when the real-time analysis model outputs a specific judgment, the specific command template to be invoked, along with the real-time data from the current live stream as parameters, is entered into that template, thereby instantiating a concrete and executable enhanced command.

[0032] S4, based on the enhanced instructions generated by S3, calls the content generation module in parallel, retrieves corresponding resources from the template library and resource library, fills relevant data into the visualization template to generate infographics, or drives the 3D model and virtual anchor to generate corresponding broadcast content.

[0033] S5, for augmented reality deployment, calculates precise spatial registration parameters for virtual augmented content based on spatial perception data and spatial anchor point information in instructions, for AR clients to perform localized rendering and overlay.

[0034] Specifically, real-time data streams from dedicated sensors (depth cameras, LiDAR, SLAM cameras) are accessed, including depth maps, point clouds, feature points, or pre-constructed sparse spatial maps of the current scene, accompanied by precise timestamps and the device's own pose information. Parameters related to augmented reality deployment are parsed from the augmented content generation instructions issued in step S3. These parameters must explicitly specify the spatial anchor points where the virtual content should be placed; anchor point information may be defined in various forms.

[0035] The coordinate system of the sensor containing the spatial perception data, the world coordinate system of the AR device, and the coordinate system referenced by the anchor point in the command are aligned. If the anchor point is a preset marker, the marker is detected in real time in the current sensor data, and the precise position (translation vector) and orientation (rotation matrix) of the marker relative to the AR device's camera are calculated. If the anchor point is a semantic location, it needs to be combined with environmental understanding. If the anchor point is scene coordinates, the transformation relationship from the virtual content coordinates to the device camera coordinates is calculated directly using the real-time position of the device itself in the scene map provided by SLAM technology. Combining the above calculations, a 4x4 transformation matrix is ​​finally obtained, which accurately describes the position, rotation (orientation), and scaling of the virtual content in real-world space.

[0036] The calculated transformation matrix is ​​filtered in the time domain (Kalman filtering) to smooth its changes, thereby eliminating jitter caused by sensor noise and making the virtual content appear stable and locked in space.

[0037] Occlusion Handling Data Preparation: Based on the depth map information, determine the geometric relationship between the intended placement location of the virtual content and the real scene, and prepare the necessary depth information so that the AR client can correctly handle the effect of virtual objects being occluded by real objects (or occluding real objects) during rendering. The calculated and optimized results are encapsulated into a lightweight spatial registration parameter package. This parameter package typically includes: a unique identifier for the virtual content, the final transformation matrix in the world coordinate system, the necessary scaling ratio, a timestamp synchronized with the main video stream, and optional depth collision information. This parameter package, along with metadata such as the virtual content's resource links and rendering logic, is synchronously sent to the AR client via a low-latency channel.

[0038] After receiving the video stream, virtual content resources, and registration parameter package for this space, the AR client completes the final rendering locally. Utilizing the device's own tracking capabilities, it maintains a continuous understanding of the world space and, based on the issued transformation matrix, accurately renders the virtual content at the specified 3D spatial location. Combined with real-time camera footage, it achieves the fusion and overlay of virtual content with the real-world scene, presenting it to the end user.

[0039] S6 will overlay and blend enhancement layers on the live video screen in real time according to the screen position and time specified by the instruction. The final video stream after blending is encoded and pushed to the content delivery network.

[0040] For traditional video publishing, S7 uses the various enhanced content generated by S4 as layers, which are then overlaid onto the main live video stream in real time according to the command parameters. After encoding, the content is pushed through the content delivery network.

[0041] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: Through real-time perception and intelligent decision-making using multimodal data, the system enables automated and precise production and distribution of enhanced content. It can intelligently identify key moments in a live stream and core points of interest for viewers, automatically generating enhanced information such as data visualizations and virtual broadcasts tailored to the scene. This is then presented in real-time via low-latency fusion technology, either in the form of traditional video streams or AR. This not only greatly enriches the information dimensions and expressiveness of live streams but also allows for real-time content adjustments based on viewer interaction, significantly improving the interactivity, immersion, and information delivery effectiveness of the live stream.

[0042] Example 2: In Example 1, augmented reality (AR) deployment calculates the registration parameters of virtual content using spatial anchor point information and static spatial coordinates, achieving fixed-position overlay of virtual content in real space. However, when the target entity in the live stream continuously moves, the virtual content, fixed at its initial spatial coordinates, cannot move accordingly, leading to misalignment between the augmented information and the entity, thus compromising immersion and accuracy. The movement patterns and trajectories of entities vary significantly across different dynamic scenarios, inevitably resulting in content drift when using static spatial anchor points. Due to the lack of strong correlation between virtual content and real entities, the tracking ability for moving entities is insufficient. To achieve stable and accurate AR overlay in dynamic scenes, further optimization and improvement are needed, shifting from static spatial registration to real-time perception and binding of dynamic entities. In some embodiments, triggering the corresponding enhanced content generation instruction, step S3 further includes: S31 processes the live data stream in real time, performing instance perception, pixel-level segmentation, and temporary ID tracking, extracting entity instances and their outlines from the live scene and surrounding shooting environment, and uploading them to the cloud.

[0043] Specifically, the system receives the original video frame sequence from the live stream's main signal stream. For each frame or keyframe sampled at a certain frequency, necessary preprocessing is performed, and the preprocessed video frames are input into a lightweight neural network model deployed on the edge (broadcast device or edge server).

[0044] The quantized neural network model performs two core tasks in parallel: first, instance perception, which identifies and distinguishes each independent object instance in the image, assigning each instance a temporary ID valid only within the current tracking session; and second, pixel-level segmentation, where the model outputs a precise pixel-level mask for each identified entity instance, clearly outlining the entity's contour boundaries.

[0045] The instance perception and segmentation results of the current frame are correlated with the detection results of the previous frame (or several previous frames). Matching is performed using appearance features (color, texture) and motion trajectory prediction to maintain the continuity of the temporary ID for the same entity across different frames. If a new entity enters the frame, a new temporary ID is assigned; if the entity disappears, its ID is marked as inactive.

[0046] By combining the spatially aware data stream available at the edge (coarse camera pose and sparse point cloud from SLAM), the 2D position of each entity in the image is associated with its approximate position in the coarse 3D scene. For each tracked entity instance, the following key metadata is extracted and packaged: temporary ID, category label, pixel-level segmentation mask, 2D bounding box, center point coordinates in the image, associated coarse 3D position (if available), and timestamp of the current frame.

[0047] The metadata of all entity instances obtained from processing this frame (or this batch) is efficiently compressed together with the corresponding keyframe image index. The compressed data is then uploaded to the cloud processing center in real time via a low-latency network link.

[0048] S32 uses data uploaded from the cloud to perform refined identification and 3D reconstruction of entities in the live stream, and builds and continuously updates a dynamic spatiotemporal map describing entities and their relationships.

[0049] Among them, 3D reconstruction estimates the approximate location, orientation, and simplified 3D shape of an entity in 3D space. Nodes in the dynamic spatiotemporal graph represent each tracked entity instance, and node attributes include unique ID, category, identity, real-time 3D bounding volume, historical trajectory, etc.; edges represent dynamic relationships between entities.

[0050] S33 matches and parses the enhanced instructions with the spatiotemporal map to identify the target entity, and generates an enhanced anchoring instruction that is bound to the entity and describes the relative spatial relationship.

[0051] The target entity is a specific object instance that is uniquely identified and continuously tracked in the scene entity spatiotemporal graph constructed and maintained in steps S31 and S32.

[0052] S34: The client dynamically calculates the pose of the virtual content for each frame based on the real-time status of the entity synchronized from the cloud and the anchoring instructions, and performs stable following and rendering of the dynamic entity based on the virtual content.

[0053] Specifically, the client continuously receives and parses two core data streams from the cloud: the first is the latest status of the relevant target entities in the spatiotemporal graph of the scene entities, including their unique identifiers, positions and orientations in three-dimensional space; the second is the enhanced anchoring instruction, which defines the target entity identifier that the virtual content needs to be bound to, as well as the position, rotation and scaling relationship of the virtual content relative to that entity.

[0054] Before each frame of rendering begins, the client searches for the target entity identifier specified in the anchoring command in the entity status list synchronized from the cloud, and obtains the entity's latest 3D position and orientation matrix at the current moment. Then, using the target entity's current 3D pose obtained in the previous step as a reference, the client converts the relative displacement, rotation, and scaling relationships defined in the anchoring command into the absolute position and orientation of the virtual content in the current world coordinate system through matrix operations. This calculation process is executed dynamically frame by frame.

[0055] The calculated virtual content pose is smoothed in the temporal domain. The system checks if the target entity is occluded by other objects or has moved out of the frame, and determines the virtual content's visibility based on predefined rules in the anchoring instructions. The client uses its own rendering engine to draw the virtual content based on the calculated and smoothed pose matrix. Through the device's augmented reality framework, the rendered virtual content layer is fused with real-world video frames captured by the camera, ultimately outputting a complete image combining virtual and real elements on the screen.

[0056] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: By performing real-time 3D perception and digital modeling of the live streaming scene, a spatiotemporal map describing entities and their relationships was constructed, enabling the accurate parsing and anchoring of augmented commands to specific target entities. This allows virtual information to follow the movement and rotation of entities in real time and stably, just as if it were actually attached to an object, effectively solving the problem of augmented content drifting and misalignment in dynamic scenes, and significantly improving the accuracy, realism, and immersive experience of AR overlay.

[0057] Example 3: In Example 2, by constructing a spatiotemporal graph of scene entities and implementing the binding and following of virtual content to dynamic entities, the problem of spatial registration and stable attachment of augmented content in dynamic scenes was solved. However, this method is essentially a follow-up augmentation, and its rendering is based entirely on the current and past states of the entities. When the entity moves quickly, its intent changes suddenly, or the system has inherent processing and transmission delays, the display of augmented content will exhibit a perceptible lag, failing to proactively indicate the entity's upcoming actions or future developments. Using binding rules purely based on the current state to deal with entities with different movement patterns and behavioral intentions inevitably leads to insufficient immediacy and a lack of guidance in information transmission. In high-speed changing scenarios such as sports events and live commentary, it is necessary to introduce the ability to predict short-term entity behavior and drive the generation and rendering of augmented content based on the prediction results, requiring further optimization and improvement.

[0058] In some embodiments, step S33, which involves parsing the enhanced instructions against the spatiotemporal map to extract the target entity, further includes: S331 acquires high-frequency motion sequences of target entities in real time, performs fine-grained analysis on the video region where the target entity is located, and extracts refined behavioral features that exceed the bounding box.

[0059] In this process, behavioral features are extracted from the main live video stream by cropping the high-resolution image region containing the target entity's 2D bounding box provided by the scene entity spatiotemporal graph. Multiple lightweight deep learning models are run in parallel on this region, including: a pose estimation model to extract the coordinates of key skeletal or structural points of the human body or object, forming a pose sequence; an action recognition model to determine the category of the micro-action performed; and an appearance feature extraction network to encode the entity's fine-grained texture and appearance variations. Finally, the structured data (coordinate sequence, action category, feature vector) output by these models are fused to form a unified, refined behavioral feature representation that transcends the bounding box. S332 constructs a dynamic graph for neural network processing based on the spatiotemporal graph at each moment. The dynamic graph is input into the spatiotemporal graph neural network to simulate the transmission and coordination of intentions between entities. The output is a set of probabilistic trajectories for each key entity in the future short period of time, as well as the discrete behavior label most likely to be executed at the corresponding moment.

[0060] In this dynamic graph, nodes represent entities, and their initial feature vectors encode the entity's kinematic state (position, velocity) at the current time (denoted as time t), the refined behavioral features extracted by step S331, and identity semantics. Edges in the graph represent relationships between entities, and their weights or features encode physical distance, kinematic interaction (relative velocity), and semantic relationships. The "future short time" needs to be defined according to the specific application scenario; for example, in live sports broadcasts, predictions are typically made for the next 1 to 3 seconds.

[0061] Specifically, at time t, a dynamic spatiotemporal graph is constructed based on the real-time maintained spatiotemporal graph of scene entities. This dynamic graph is then input into a pre-trained spatiotemporal graph neural network. The network performs graph reasoning through multiple rounds of message passing and node update operations: in each round, each entity node aggregates information from its neighboring nodes (other entities connected by edges); subsequently, each node updates its current hidden state based on its own historical hidden state and the messages aggregated from its neighbors. This hidden state is a high-dimensional vector representing the entity's overall posture after condensing its own history and surrounding context. The aforementioned message passing and node update process occurs across multiple layers of the network, enabling information to propagate across multiple hops in the graph, thereby simulating complex long-distance interactions and collaborations.

[0062] After multiple rounds of graph inference, the final hidden state of each key entity node is fed into a specific decoder. The trajectory decoder maps this hidden state to a prediction of the entity's location at a series of discrete future time points (e.g., from t+0.5 seconds to t+2.0 seconds, with 0.5-second intervals). To achieve probabilistic prediction, the decoder outputs parameters (mean and covariance matrix) of a multivariate Gaussian distribution, which defines the range of possible spatial locations of the entity at each future time point. By sampling from this distribution, multiple possible trajectory hypotheses can be obtained. Simultaneously, a parallel behavior classification decoder predicts the probability distribution of the entity performing various preset behaviors within the corresponding future time intervals based on the same hidden state. The behavior label with the highest probability is output as the most likely behavior to be performed, and its probability value serves as the confidence level of the prediction.

[0063] S333: Based on the prediction results, the client predefines dynamic parameterized anchoring rules and generation logic for the enhanced content. The client performs advance rendering based on the prediction sequence and smoothly corrects the predicted content through real-time observation.

[0064] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: By introducing short-term behavioral intent prediction, it achieves a leap from presenting the present to predicting the future. Based on high-probability predictions of the future trajectories and behaviors of dynamic entities in a scene, it proactively generates and renders enhanced content, naturally directing virtual information towards upcoming actions. Through advanced rendering and a smoothing correction mechanism based on real-time observation, it effectively eliminates display lag caused by network latency, ensuring the stability, accuracy, and visual coherence of enhanced content in dynamic scenes, significantly improving the depth, immersion, and interactive intelligence of live broadcast interpretation.

[0065] Example 4: In Example 3, the proactive presentation of enhanced content through behavior prediction effectively solved the motion latency problem. However, this method assumes that entities maintain the integrity of their form and structure during prediction. When a live-stream entity undergoes dynamic changes in its physical structure, such as disassembly, dispersion, or recombination (e.g., product demonstration disassembly, instantaneous dispersion of tactical formations), the assumption of its single position and posture fails, causing a complete break in all prediction, enhanced content binding, and presentation logic based on that entity, making it unable to adapt to subsequent scenarios. To maintain the continuity of the enhanced information's semantics and the rationality of its presentation even when the entity's form undergoes fundamental changes, it is necessary to introduce real-time perception and management capabilities for structural changes such as entity splitting and aggregation, and to further optimize and improve the method.

[0066] In some embodiments, step S331, which involves acquiring the high-frequency motion sequence of the target entity in real time, further includes: Real-time monitoring of the state continuity of the target entity; when physical or visual splitting of the target entity is detected, identification and tracking of all sub-target entities generated by the split, and establishment of the derivation relationship between the sub-target entities and the target entity.

[0067] When the split event is confirmed, the type of the target entity in the spatiotemporal graph is changed from the parent entity to a logical entity container. All the child target entity nodes that were identified and established with the derivation relationship in step C1 are added as member entities of this logical container, and the state of all its members is dynamically aggregated. The logical center, spatial range and the derivation attributes of the parent entity are calculated in real time.

[0068] Among them, the logical entity container inherits the core attributes such as the unique identifier, semantic category, and name of the parent entity (target entity), but no longer has a single fixed position and posture.

[0069] After an entity splits, the enhancement content associated with the parent entity is automatically rebound to the logical container or its designated member entity. The resolution conditions of the logical container are set and monitored. When the conditions are met, the enhancement content is smoothly migrated to the designated member entity or faded out, and the container structure in the graph is cleaned up.

[0070] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: By employing a logical container mechanism, the core challenge of disrupted content binding relationships and chaotic information presentation when entities in live streaming undergo morphological changes such as disassembly or separation is effectively addressed. It can detect entity splitting events in real time, dynamically create and manage logical containers representing the original entities, ensuring that all enhanced content can be automatically and intelligently rebound to the container or its key components. This achieves seamless continuity of information association and adaptive adjustment of visual presentation, significantly improving the continuity, accuracy, and overall consistency of the user experience in complex and dynamic scenarios such as demonstrations and explanations.

[0071] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for producing and releasing converged media products, characterized in that, include: S1: At the same time as the live broadcast starts, the main live broadcast signal stream is accessed, its visual and audio features are extracted in real time, and the live broadcast content is initially classified based on these features. S2 synchronously accesses external business data streams and audience interaction data streams, performs sentiment analysis, topic clustering, and keyword extraction on interactive texts, and combines the scene classification results of S1 to comprehensively determine the audience's core concerns and the core theme of the live broadcast. S3 monitors based on preset rules and real-time analysis models. When business data exceeds the threshold, a specific keyword combination is identified, or a high-attention moment is detected, corresponding enhanced content generation instructions are triggered. S4, based on the enhanced instructions generated by S3, calls the content generation module in parallel, retrieves the corresponding resources from the template library and resource library, fills the relevant data into the visualization template to generate information charts, or drives the 3D model and virtual anchor to generate the corresponding broadcast content. S5, for augmented reality deployment, calculates the precise spatial registration parameters of virtual augmented content based on spatial perception data and spatial anchor point information in the instructions, for AR clients to perform localized rendering and overlay; S6 will overlay and blend enhancement layers on the live video screen in real time according to the screen position and time specified by the instruction. The final video stream after blending is encoded and pushed to the content delivery network. For traditional video publishing, S7 uses the various enhanced content generated by S4 as layers, which are then overlaid onto the main live video stream in real time according to the command parameters. After encoding, the content is pushed through the content delivery network.

2. The method for producing and releasing converged media products according to claim 1, characterized in that, The extraction of visual and audio features specifically includes: visual features refer to scene and object recognition, person detection and recognition, text information extraction and image attribute analysis; audio features refer to speech-to-text conversion, keyword and topic detection, voiceprint and speaker recognition and audio event detection.

3. The method for producing and releasing converged media products according to claim 1, characterized in that, The localized rendering and overlay for the AR client includes: the AR client using its own environment tracking capabilities to maintain a continuous understanding of the real-world space; accurately drawing virtual content on a specified three-dimensional spatial location based on the received spatial transformation matrix; and fusing and overlaying the drawn virtual content layer with real-world video frames captured by the device's camera in real time.

4. The method for producing and releasing converged media products according to claim 1, characterized in that, The triggering of the corresponding enhanced content generation instruction specifically includes: S31 processes the live data stream in real time to complete instance perception, pixel-level segmentation and temporary ID tracking, and extracts entity instances and their outlines from the live scene and the surrounding shooting environment and uploads them to the cloud. S32 refines the identification and 3D reconstruction of entities in the live broadcast by using data uploaded from the cloud, and constructs and continuously updates a dynamic spatiotemporal map describing entities and their relationships. S33, the enhanced command is matched and parsed with the spatiotemporal graph to derive the target entity, and an enhanced anchoring command describing the relative spatial relationship is generated and bound to the entity; where the target entity is a specific object instance that is uniquely identified and continuously tracked in the constructed and maintained scene entity spatiotemporal graph; S34: The client dynamically calculates the pose of the virtual content for each frame based on the real-time status of the entity synchronized from the cloud and the anchoring instructions, and performs stable following and rendering of the dynamic entity based on the virtual content.

5. The method for producing and releasing converged media products according to claim 4, characterized in that, The extraction of entity instances and their outlines from the live streaming scene and surrounding shooting environment specifically includes: receiving the original video frame sequence from the live streaming main signal stream and preprocessing the video frames; inputting the preprocessed video frames into a lightweight neural network model deployed on the edge; the model performs instance perception and pixel-level segmentation in parallel to identify and distinguish each independent object instance in the image, assigns a temporary ID to each instance, and outputs a pixel-level mask for each instance to outline its contour boundary; cross-frame association between the recognition results of the current frame and the results of historical frames, maintaining the continuity of the temporary ID of the same instance through appearance feature and motion trajectory matching, and handling the appearance of new entities and the disappearance of old entities; and combining edge-side spatial perception data to associate the two-dimensional position of the entity in the image with its approximate position in the three-dimensional scene.

6. The method for producing and releasing converged media products according to claim 5, characterized in that, In the dynamic spatiotemporal graph, nodes represent each tracked entity instance, and edges represent dynamic relationships between entities.

7. The method for producing and releasing converged media products according to claim 5, characterized in that, The step of parsing the enhanced commands with the spatiotemporal graph to extract the target entity includes: S331: Real-time acquisition of high-frequency motion sequences of target entities, fine-grained analysis of the video region where the target entity is located, and extraction of refined behavioral features beyond the bounding box; S332, based on the spatiotemporal graph at each moment, constructs a dynamic graph for neural network processing. The dynamic graph is input into the spatiotemporal graph neural network to simulate the transmission and coordination of intentions between entities, and outputs a set of probabilistic trajectories for each key entity in the future short period of time, as well as the discrete behavior label most likely to be executed at the corresponding moment. S333: Based on the prediction results, the client predefines dynamic parameterized anchoring rules and generation logic for the enhanced content. The client performs advance rendering based on the prediction sequence and smoothly corrects the predicted content through real-time observation.

8. The method for producing and releasing converged media products according to claim 7, characterized in that, The behavioral features are specifically defined as follows: from the main live video stream, video data of the target entity's location is extracted based on its image range; a pose estimation model, an action recognition model, and an appearance feature extraction model are run in parallel on the video data; wherein, the pose estimation model is used to extract the pose sequence of the target entity, the action recognition model is used to determine the type of micro-action it performs, and the appearance feature extraction model is used to encode its refined texture and appearance changes; the results of pose estimation, action recognition, and appearance feature extraction are fused to form refined features representing the dynamic behavior of the target entity.

9. The method for producing and releasing converged media products according to claim 7, characterized in that, The process of outputting a set of probabilistic trajectories for each key entity in the short term includes: inputting a dynamic graph reflecting the entity and its relationships into a pre-trained spatiotemporal graph neural network; calculating and outputting the hidden state of each key entity node representing its overall situation through graph inference of the neural network; inputting the hidden state into the trajectory decoder; the trajectory decoder outputting a parameterized probability distribution defining the range of possible spatial locations of the entity at each discrete time point in the future; and generating multiple trajectory hypotheses representing the future movement possibilities of the entity based on the probability distribution.

10. The method for producing and releasing converged media products according to claim 7, characterized in that, The real-time acquisition of the high-frequency motion sequence of the target entity specifically includes: real-time monitoring of the state continuity of the target entity; when a physical or visual split is detected, identifying and tracking each generated sub-entity, and establishing the derivation relationship between the sub-entity and the target entity; after the split event is confirmed, converting the type of the target entity in the spatiotemporal graph into a logical entity container, adding each sub-entity as a member entity of the container, and dynamically aggregating the states of all members to calculate the logical center and spatial range of the container in real time; automatically rebinding the enhancement content originally bound to the target entity to the logical container or its designated member entity; setting and monitoring the resolution conditions of the logical container, and when the conditions are met, smoothly migrating the enhancement content to the designated member entity or fading it out, and cleaning up the container structure in the graph.