Video Content Generation Method, System, Device and Medium for Multi-Source Material Fusion
Through multi-source material acquisition, multi-modal semantic analysis and terminal adaptation processing, the problem of fixed fusion output mode in the multi-source material live broadcast system is solved, and the synchronization and adaptive display of multi-modal live broadcast data is realized, improving the live broadcast quality and user experience.
Patent Information
- Application Number
- CN202510450251.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing multi-source live broadcast system lacks deep understanding of content semantics and adaptive adjustment capabilities, resulting in a fixed fusion output method, making it difficult to adapt to dynamic changes in complex scenarios, and has poor real-time performance.
Through multi-source material acquisition and preprocessing, multi-modal semantic analysis and importance scoring, multi-source material fusion and layout decision-making, and audience terminal adaptation and layout rendering, the synchronization, intelligent analysis and adaptive display of multi-modal live broadcast data is achieved.
The integration capability and display flexibility of the multi-source material live broadcast system have been improved, so that the live broadcast system can dynamically adapt to changes in complex scenarios, and personalized optimization is implemented for different terminals, significantly improving the quality of live broadcast and user experience.
Smart Images

Figure CN119996786B_ABST
Abstract
Description
Technical Field
[0001] This application relates to data processing technologies, and in particular, to a method, system, device, and medium for generating video content by fusing multi-source materials. Background Art
[0002] With the rapid development of video live broadcast technologies, the live broadcast content has become increasingly rich, and the material sources have become increasingly diverse. Especially in scenarios such as e-commerce live broadcasts, sports event live broadcasts, news reports, and virtual studios, it is often necessary to simultaneously fuse and display videos, audios, texts, images, and sensor data from multiple collection terminals. How to achieve real-time fusion and intelligent display of multi-modal materials while ensuring synchronization has become the core challenge for improving the live broadcast quality and user experience.
[0003] However, existing multi-source material live broadcast systems mainly rely on preset templates or manual intervention for material arrangement, lacking in-depth understanding of content semantics and adaptive adjustment capabilities. As a result, the fusion output method is fixed, making it difficult to adapt to the dynamic changes in complex scenarios, and the real-time performance is poor, affecting the live broadcast experience. In addition, the diversity of terminal devices (such as mobile phones, tablets, VR devices) and the variability of user interaction methods make it difficult for traditional static fusion methods to meet the personalized optimization requirements. Summary of the Invention
[0004] This application provides a method, system, device, and medium for generating video content by fusing multi-source materials, aiming to solve the problems in existing multi-source material live broadcast systems, such as the lack of in-depth understanding of content semantics and adaptive adjustment capabilities, resulting in a fixed fusion output method, difficulty in adapting to the dynamic changes in complex scenarios, and poor real-time performance.
[0005] In a first aspect, this application provides a method for generating video content by fusing multi-source materials, including:
[0006] Collecting and preprocessing multi-source materials, collecting multiple terminal data on the host side in real time, performing modal classification, global timestamp annotation, format standardization processing, and structure unification processing for each piece of the terminal data, and annotating display intention tags for the terminal data based on a display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality;
[0007] Performing multi-modal semantic analysis and importance scoring, inputting the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extracting semantic features, and performing importance scoring;
[0008] Fusing multi-source materials and making layout decisions, constructing a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring, and automatically generating a display decision plan according to the weighted attribute graph using a fusion strategy;
[0009] The viewer terminal adaptation and layout rendering dynamically generate a visual layout and complete the display and rendering of multi-source materials based on the display decision-making scheme in combination with the screen attributes and interaction capabilities of the viewer terminal.
[0010] In a possible design, the multi-source material collection and preprocessing include:
[0011] The server receives terminal data uploaded from multiple collection terminals and labels the modal type tags according to the modal type;
[0012] The server labels a global timestamp for each piece of the terminal data according to the modal timestamp labeling strategy;
[0013] For the terminal data with the global timestamp labeled, perform format standardization processing according to its modal type and convert it into an internally unified data representation format;
[0014] For the terminal data with the format standardized, perform structure unification processing according to its modal type so that the terminal data of the same modality has a unified structure specification;
[0015] Based on the display intention labeling strategy, label the display intention tags for the terminal data with the structure unified, including:
[0016] Extract display intention feature information from the terminal data, and the display intention feature information includes modal type, source identifier of the host side collection terminal, metadata information, data structure fields, and material size and ratio information;
[0017] Compare the display intention feature information with a preset matching rule;
[0018] Label the corresponding display intention tag for the terminal data according to the comparison result.
[0019] In a possible design, the modal timestamp labeling strategy includes:
[0020] The server maintains a unified global timestamp benchmark;
[0021] When the server receives terminal data uploaded from multiple collection terminals, record the reception time of each piece of the terminal data respectively;
[0022] Maintain the one-way network delay of each collection terminal;
[0023] Calculate the global timestamp of each piece of the terminal data based on the one-way network delay;
[0024] Label the global timestamp for each piece of the terminal data according to the stamping rule, and the stamping rule includes:
[0025] When the terminal data is in video mode, a global timestamp is marked for each frame of video data;
[0026] When the terminal data is in audio mode, a global timestamp is marked for the audio data according to a first preset period;
[0027] When the terminal data is in text mode, a global timestamp is marked for each piece of text;
[0028] When the terminal data is in picture mode, a global timestamp is marked for each picture;
[0029] When the terminal data is in sensor mode, a global timestamp is marked at each update.
[0030] In a possible design, the multi-modal semantic analysis and importance scoring includes:
[0031] Construct a modal data sequence, where the modal data sequence includes a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intention label, metadata information, material size and ratio information, and modal-specific content;
[0032] Send the modal data sequence to a multi-modal semantic analysis engine for processing according to the modal type respectively;
[0033] The multi-modal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors;
[0034] Based on the collected terminal data, display intention label, and semantic feature vector, importance scoring is performed.
[0035] In a possible design, the multi-source material fusion and layout decision includes:
[0036] Based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring, construct a weighted attribute graph. The weighted attribute graph includes nodes and associated edges. Each node represents a terminal data unit. The attributes of the node include a node unique identifier, a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intention label, metadata information, material size and ratio information, a semantic feature vector, and importance scoring. The associated edges include semantic association edges, time overlap edges, and conflict detection edges;
[0037] Perform material priority sorting based on the semantic association edges. Use a clustering algorithm to group each node into clusters according to the weights of the semantic association edges of the nodes, and determine the display priority order of each cluster grouping set according to the importance scores corresponding to each node.
[0038] Perform layout allocation based on the time overlap edges. According to the result of the material priority sorting and the weights of the time overlap edges between nodes, allocate the optimal display time sequence and layout area for the nodes.
[0039] Perform local conflict detection based on the conflict detection edge weights and material size and ratio information.
[0040] Construct a fusion optimization objective function based on the weights of the nodes and associated edges of the weighted attribute graph, and automatically generate a display decision plan that meets the requirements of display priority, layout timing, and conflict avoidance based on the fusion optimization objective function.
[0041] The display decision plan includes:
[0042] The display order of multi-source materials, sorted according to the importance scores and semantic clustering results in the fusion objective function.
[0043] The spatial layout plan of multi-source materials, based on the greedy algorithm to search and allocate spatial areas on the basis of no time conflict.
[0044] The interaction strategy of multi-source materials, formulating interaction display or replacement rules according to the conflict edge penalty value and material size and ratio information.
[0045] In a possible design, the construction of the associated edges includes:
[0046] Calculate the weights of the semantic association edges and construct the semantic association edges.
[0047] Calculate the weights of the time overlap edges and construct the time overlap edges.
[0048] Calculate the weights of the conflict detection edges and construct the conflict detection edges.
[0049] In a possible design, the viewer terminal adaptation and layout rendering include:
[0050] Obtain the adaptation information of the viewer terminal, and the adaptation information includes device characteristics, display capabilities, interaction methods, network environment, and system preference settings.
[0051] Based on the adaptation information of each viewer terminal and combined with the display decision plan, dynamically adjust the spatial layout, rendering method, and interaction strategy of multi-source materials.
[0052] Perform rendering and interaction optimization according to the adjusted plan.
[0053] In a second aspect, the present application provides a video content generation system for multi-source material fusion, including:
[0054] A multi-source material collection and preprocessing module that real-time collects multiple terminal data on the host side, performs modal classification, global timestamp annotation, format standardization processing, and structure unification processing for each piece of the terminal data, and performs display intention label annotation on the terminal data based on a display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality;
[0055] A multi-modal semantic analysis and importance scoring module that inputs the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extracts semantic features, and performs importance scoring;
[0056] A multi-source material fusion and layout decision-making module that constructs a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring, and automatically generates a display decision-making plan according to the weighted attribute graph using a fusion strategy;
[0057] An audience terminal adaptation and layout rendering module that dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision-making plan in combination with the screen attributes and interaction capabilities of the audience terminal.
[0058] In a third aspect, the present application provides an electronic device, including:
[0059] A processor; and,
[0060] A memory for storing executable instructions of the processor;
[0061] Wherein, the processor is configured to execute any possible method described in the first aspect by executing the executable instructions.
[0062] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement any possible method described in the first aspect.
[0063] The video content generation method, system, device and medium for multi-source material fusion provided by this application can realize the synchronization, intelligent analysis and adaptive display of multi-modal live data. By annotating with global timestamps, the timing consistency of data on multiple terminals is ensured. Through multi-modal semantic analysis, key information is extracted. Based on the fusion strategy of the weighted attribute graph, the sorting and layout of multi-source materials are optimized, and the display effect is dynamically adjusted in combination with the adaptation information of the viewer terminals. This application breaks through the limitations of the existing multi-source live system with fixed fusion methods and poor real-time performance, improves the fusion ability and display flexibility of the multi-source material live system, enables the live system to dynamically adapt to complex scenario changes, and realizes personalized optimization for different terminals. Moreover, this application has high real-time performance and strong scalability, and can be widely applied to various scenarios such as e-commerce live broadcasts, sports event live broadcasts, and news reports, significantly improving the live broadcast quality and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0065] Figure 1 is a schematic flowchart of the video content generation method for multi-source material fusion shown according to an exemplary embodiment of this application;
[0066] Figure 2 is a schematic flowchart of the multi-source material collection and preprocessing shown according to an exemplary embodiment of this application;
[0067] Figure 3 is a schematic flowchart of the multi-modal semantic analysis and importance scoring shown according to an exemplary embodiment of this application;
[0068] Figure 4 is a schematic flowchart of the multi-source material fusion and layout decision shown according to an exemplary embodiment of this application;
[0069] Figure 5 is a schematic flowchart of the viewer terminal adaptation and layout rendering shown according to an exemplary embodiment of this application;
[0070] Figure 6 is a schematic structural diagram of the video content generation system for multi-source material fusion shown according to an exemplary embodiment of this application;
[0071] Figure 7 is a schematic structural diagram of the electronic device shown according to an exemplary embodiment of this application.
[0072] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the textual description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0073] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0074] Figure 1 is a flowchart showing a method for generating video content by fusing multi-source materials according to an exemplary embodiment of the present application. As Figure 1 shown, the method provided in this embodiment includes:
[0075] Step S101: Multi-source material collection and preprocessing. Real-time collection of multiple terminal data on the host side, performing modal classification, global timestamp annotation, format standardization processing, and structure unification processing for each piece of the terminal data, and performing display intention label annotation on the terminal data based on the display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality.
[0076] In this step, a server is set up to collect terminal data uploaded by multiple collection terminals on the host side in real time to adapt to the current complex live broadcast scenario requirements, including large-scale events, live sales, sports events, virtual hosts, and multi-scenario concatenated live broadcasts, etc.
[0077] For the above-mentioned complex live broadcast scenarios, the terminal data collected on the host side includes the following five modal types:
[0078] Video modality: Multi-view synchronous collection is achieved through multiple camera terminals (such as the main camera, side view, top view, mobile follow-up device) to build a panoramic live broadcast experience and enhance the audience's immersion. This type of collection method is applicable to scenarios such as sports commentary, live sales, and talent shows.
[0079] Audio modality: The audio source is usually independent of the video collection terminal and includes multiple audio contents such as background music, voice assistant output, and remote guest voices. It is applicable to scenarios such as concert live broadcasts, virtual host live broadcasts, game commentaries, and remote connections.
[0080] Image modality: Refers to the static image content uploaded, intercepted, or generated by the system on the host side, commonly used to display background images, courseware PPTs, graphic lecture notes, illustrations, cover diagrams, or product detail images, etc. It is applicable to text-and-graphic combination scenarios such as educational live broadcasts, e-commerce live broadcasts, and corporate press conferences.
[0081] Text modality: Includes real-time interactive information on the viewer side such as barrage information, viewer comments, and chat room messages, which is used to enhance interactivity and semantic supplementation of the live broadcast content.
[0082] Sensor modality: Derived from the operation data of sensing devices and peripherals worn or used by the host, including somatosensory information, heart rate monitoring, pose / action capture, mouse trajectory, keyboard keystrokes, gamepad input, etc. It is widely used in scenarios such as virtual host live broadcasts, game live broadcasts, fitness live broadcasts, and educational interactive live broadcasts.
[0083] Step S102: Multimodal semantic analysis and importance scoring. Input the preprocessed terminal data into a multimodal semantic analysis engine according to the modality type, extract semantic features, and perform importance scoring.
[0084] In this step, the multimodal semantic analysis engine consists of multiple modular artificial intelligence models, and each artificial intelligence model performs customized semantic feature extraction for the terminal data of one modality type. The modality types are as follows:
[0085] Video modality: Through visual artificial intelligence models such as 3D convolutional neural networks and Transformer-based frame-level analysis models, extract semantic features such as scene types, host actions, interaction behaviors, and visual attention areas in video clips, which are used to judge the display value, interaction density, and focus content of the current video content.
[0086] Audio modality: Through audio artificial intelligence models such as speech recognition, speech emotion recognition, and sound source classification, such as speech recognition models like Wav2Vec 2.0 and Whisper, extract features such as speech content text, speaker identity, tone emotion state, and audio type (speech / music / noise), which are used to evaluate the emotional intensity, semantic density, and display relevance of the speech content.
[0087] Image modality: Through image classification, object detection, OCR recognition and other image artificial intelligence models, such as ResNet image classification model, YOLOv5 object detection model, etc., analyze the visual subject, text content, style features, and structural information in the image, extract the category of the image (such as courseware image, product image, background image, etc.) and its key semantic content, which is used to judge its display intention and adapted layout.
[0088] Text modality: Through natural language processing models such as the BERT semantic understanding model and the TextCNN text classification model, semantic understanding is performed on text data such as bullet screens and comments, and features such as keywords, topic directions, sentiment tendencies, and frequency weights are extracted to assist in hot topic recognition, display intent matching, and real-time interaction guidance.
[0089] Sensor modality: Through time series modeling, action recognition, or behavior prediction models such as the HAR-LSTM behavior recognition model and the ST-GCN spatio-temporal graph convolutional network model, behavior state recognition, operation intent prediction, physiological signal analysis, etc. are performed on data collected from somatosensory devices, heart rate sensors, keyboard and mouse trajectories, controllers, etc., and the current action / state information of the anchor is extracted to judge key interaction events or virtual character driving signals.
[0090] Step S103: Multi-source material fusion and layout decision-making. Based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance score, a weighted attribute graph is constructed, and a display decision-making scheme is automatically generated using a fusion strategy according to the weighted attribute graph.
[0091] Step S104: Audience terminal adaptation and layout rendering. Based on the display decision-making scheme, combined with the screen attributes and interaction capabilities of the audience terminal, a visual layout is dynamically generated and the display rendering of multi-source materials is completed.
[0092] Figure 2 It is a schematic flowchart of multi-source material collection and preprocessing shown by this application according to an exemplary embodiment. As Figure 2 shown, the method provided in this embodiment includes:
[0093] Step S201: The server receives terminal data uploaded from multiple collection terminals and labels modal type tags according to the modal type.
[0094] In this step, the terminal data is classified into modalities according to its type, and a corresponding modal type tag is labeled for each piece of data. The modal type tags include video modality, audio modality, text modality, picture modality, and sensor modality.
[0095] Step S202: The server labels a global timestamp for each piece of terminal data according to the modal timestamp labeling strategy.
[0096] In this step, the modal timestamp labeling strategy includes:
[0097] The server maintains a unified global timestamp benchmark to eliminate clock deviations between different collection terminals, achieve time synchronization and unified alignment of multi-modal data, and ensure the timing accuracy and consistency of subsequent fusion and display.
[0098] When the server receives the terminal data uploaded by multiple collection terminals, it records the reception time of each piece of terminal data respectively for calculating the one-way network delay.
[0099] Maintain the one-way network delay of each collection terminal. The calculation formula for the one-way network delay is:
[0100]
[0101] Where, is the one-way network delay of the i-th collection terminal, is the round-trip delay measured for the k-th time by the i-th collection terminal, is the weight of the k-th measurement, is the weighted average of the sliding window of the most recent
[0102] Calculate the global timestamp of each piece of terminal data based on the one-way network delay. The calculation formula for the global timestamp is:
[0103]
[0104] Where, is the current global timestamp of the i-th terminal data, is the current local timestamp of the i-th terminal data, is the clock offset between the server and the collection terminal;
[0105] The calculation formula for
[0106]
[0107] Where, is the reception time of the i-th terminal data recorded by the server;
[0108] Perform global timestamp annotation for each piece of terminal data according to the stamping rule. The stamping rule includes:
[0109] When the terminal data is in video mode, annotate the global timestamp for each frame of video data.
[0110] When the terminal data is in audio mode, annotate the global timestamp for the audio data according to the first preset period.
[0111] When the terminal data is in text mode, annotate the global timestamp for each piece of text.
[0112] When the terminal data is in picture mode, annotate the global timestamp for each picture.
[0113] When the terminal data is in sensor mode, a global timestamp is marked at each update.
[0114] Step S203: For the terminal data with the global timestamp marked, perform format standardization according to its modality type, and convert it into an internally unified data representation format.
[0115] In this step, since the data formats uploaded by different acquisition terminals are different, as shown in Table 1:
[0116] Table 1 Multiple data format types uploaded by the acquisition terminal
[0117]
[0118] Therefore, all terminal data needs to be converted into an internally unified data representation format so that the terminal data can be input into the multi-modal semantic analysis engine for subsequent parsing, decoding, and feature extraction.
[0119] Step 204: For the terminal data with format standardized, perform structure unification according to its modality type, so that the terminal data of the same modality has a unified structure specification.
[0120] In this step, since different acquisition terminals may use different encodings, encapsulations, resolutions, or meta-information formats for the terminal data of the same modality type, as shown in Table 2:
[0121] Table 2 Multiple data structures uploaded by the acquisition terminal
[0122]
[0123] Therefore, the terminal data with format standardized needs to be standardized in terms of its field structure, label information, and data organization method according to its modality type, so that the terminal data of the same modality has a consistent data structure, so that the terminal data can be input into the multi-modal semantic analysis engine for subsequent parsing, decoding, and feature extraction.
[0124] Step S205: Mark the display intent label for the terminal data with structure unified based on the display intent annotation strategy.
[0125] In this step, the marking of the display intent label includes:
[0126] Extract the display intent feature information from the terminal data, and the display intent feature information includes modality type, the source identifier of the host-side acquisition terminal, metadata information, data structure fields, and material size and ratio information. Specifically as follows:
[0127] Modal type: used to determine the analysis path of terminal data, including video modality, audio modality, image modality, text modality, and sensor modality.
[0128] Host-side acquisition terminal source identifier: the unique identification information of the terminal data source, used to distinguish the roles of material content, including the number of the acquisition terminal, the upload channel ID, and the file path.
[0129] Metadata information: used for rule matching and structural unification judgment, including file name, path name, media information (format, duration, frame rate, sampling rate, resolution).
[0130] Data structure field: used to distinguish the type of material and its content structure, including keywords and sensor types.
[0131] Material size and ratio information: used for display layout judgment in the fusion strategy, including parameters such as width, height, and aspect ratio.
[0132] Compare the display intention feature information with the preset matching rules.
[0133] Label the corresponding display intention label for the terminal data according to the comparison result.
[0134] In this step, the display intention label is used to indicate the functional role of the terminal data in the fusion display of multi-source materials. The display intention labels include:
[0135] Display intention type: used to identify the content role of the terminal data in the fusion display. For example, the main speaker's picture is marked as main_speaker, the product picture is marked as product_image, and the sensor layer is marked as sensor_overlay, etc.
[0136] Display priority: used to indicate the importance of the terminal data in the overall display, and is used to assist in content sorting and display conflict judgment in the fusion strategy.
[0137] Display method hint: used to prompt the expected display form of the terminal data on the viewer's terminal interface, such as small window display, large screen display, icon overlay, etc., to support the rendering decision of the terminal adaptation and layout rendering module.
[0138] Figure 3 This is a schematic flowchart of multi-modal semantic analysis and importance scoring shown according to an exemplary embodiment of the present application. As Figure 3 shown, the method provided in this embodiment includes:
[0139] Step S301: Construct a modal data sequence, which includes a unique identifier of the data sequence, a modal type, a global timestamp, an identifier of the source of the host-side acquisition terminal, a display intention label, metadata information, material size and ratio information, and modal-specific content.
[0140] In this step, the modal data sequence serves as the standardized input data structure of the multi-modal semantic analysis engine, which is used to uniformly process and distribute model data of terminal data of different modal types during the semantic analysis phase. Since the terminal data structures and processing requirements of different modal types are different, it is necessary to construct corresponding standardized input formats according to the modal types and send them to the corresponding artificial intelligence models for semantic feature extraction. The modal data sequence includes the following fields:
[0141] Unique identifier of the data sequence: The unique identifier generated for each piece of modal data, which is used for data tracking and result association throughout the processing flow.
[0142] Modal type, global timestamp, identifier of the source of the host-side acquisition terminal, display intention label, metadata information, and material size and ratio information: All are derived from the pre-processed terminal data and display intention labels.
[0143] Modal-specific content: It is a uniformly encapsulated original material data packet for the analysis and processing of artificial intelligence models, including index identification information and original data content. The original data content is the terminal data content after format standardization processing and structure unification processing.
[0144] Step S302: Send the modal data sequence to the multi-modal semantic analysis engine for processing according to the modal type respectively.
[0145] In this step, since the multi-modal semantic analysis engine includes multiple artificial intelligence models corresponding to the modal types, it is necessary to route and distribute the modal data sequence according to the modal type and send it as the standardized input to the corresponding models respectively for subsequent semantic analysis and processing.
[0146] Step S303: The multi-modal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors.
[0147] In this step, the semantic feature vectors will serve as the key basis for subsequent importance score calculation, weighted attribute graph construction, and multi-source material fusion display, providing core support for the fusion strategy and display decision-making plan. According to the different modal types, the semantic feature vectors obtained after being processed by the multi-modal semantic analysis engine are also different. The semantic feature vectors of each modal type are as follows:
[0148] Video modality: The semantic feature vector includes the scene label, action encoding, and visual emotion vector of the video modality.
[0149] Audio modality: The semantic feature vector includes the speech transcription, emotional intensity score, speech rate, and loudness features of the audio modality.
[0150] Text modality: The semantic feature vector includes the keyword embedding, emotional polarity label, and hot word distribution features of the text modality.
[0151] Image modality: The semantic feature vector includes the main object label, image quality score, and visual style description of the image modality.
[0152] Sensor modality: The semantic feature vector includes the change trend vector, anomaly probability, and status label of the sensor modality.
[0153] Step S304: Perform importance scoring based on the collected terminal data, display intention label, and semantic feature vector.
[0154] In this step, the calculation formula for importance scoring is:
[0155]
[0156] where is the weight coefficient of the semantic feature, is the semantic feature score corresponding to the data unit with the data sequence unique identifier x, is the weight coefficient of the display intention label, is the display intention score corresponding to the data unit with the data sequence unique identifier x, is the weight coefficient of the real-time impact, is the real-time score corresponding to the data unit with the data sequence unique identifier x, is the weight coefficient of the host-side terminal source, is the host-side source terminal score corresponding to the data unit with the data sequence unique identifier x;
[0157] The calculation formula for the semantic feature score is:
[0158]
[0159] where is the semantic feature vector corresponding to the data unit with the data sequence unique identifier x, is the standard reference feature vector, is the inner product of the vectors, the product of the vector norms;
[0160] The calculation formula for the display intention score is:
[0161]
[0162] Among them, is the display intention scoring function, is the display intention label corresponding to the data unit with the unique identifier x of the data sequence;
[0163] The calculation formula for real-time scoring is:
[0164]
[0165] Among them, is the time decay coefficient, is the current system time, is the global timestamp corresponding to the data unit with the unique identifier x of the data sequence;
[0166] The formula for scoring the source terminal on the host side is:
[0167]
[0168] Among them, is the weighted scoring function of the source identifier of the collection terminal on the host side, is the source terminal identifier on the host side corresponding to the data unit with the unique identifier x of the data sequence.
[0169] Figure 4 is a schematic flowchart of a method for generating video content by fusing multi-source materials according to an exemplary embodiment of the present application. As Figure 4 shown, the method provided in this embodiment includes:
[0170] Step S401: Construct a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance score. The weighted attribute graph includes nodes and associated edges. Each node represents a terminal data unit. The attributes of the node include node unique identifier, data sequence unique identifier, modal type, global timestamp, source identifier of the collection terminal on the host side, display intention label, metadata information, material size and ratio information, semantic feature vector, and importance score. The associated edges include semantic association edges, time overlap edges, and conflict detection edges.
[0171] In this step, the construction of the associated edges includes:
[0172] Calculate the weight of the semantic association edge and construct the semantic association edge. The calculation formula for the weight of the semantic association edge is:
[0173]
[0174] Among them, is the node with the node unique identifier i, is the node with the node unique identifier j, is the corresponding semantic feature vector, is the corresponding semantic feature vector, is the inner product of vectors, is the product of the vector norms.
[0175] Semantic association edges are used to measure the semantic similarity between nodes and nodes The value range of is .
[0176] When tends to 1, it indicates that the semantic features of nodes and nodes have a high degree of similarity, and the two are relatively close in content, theme or description, and tend to be grouped into the same clustering group.
[0177] When tends to 0, it indicates that the semantic feature similarity between nodes and nodes is relatively low, and the two are quite different in content, theme or description, and usually will not be grouped into the same clustering group.
[0178] Calculate the weight of the time overlap edge and construct the time overlap edge. The calculation formula of the time overlap edge weight is:
[0179]
[0180] where, is the global timestamp of node , is the global timestamp of node , is the time proximity factor, is the time interval of node , is the time interval of node , is the time range overlap factor.
[0181] is used to measure the proximity of the global timestamps of nodes and nodes , is the time decay coefficient, which determines the influence degree of time difference on the weight of the time overlap edge.
[0182] If and are close in time, then tends to 1;
[0183] If and the time difference increases, then tends to 0.
[0184] It is used to measure the degree of overlap between node and node at the time end. When they are completely overlapped, then the value of is 1; when there is no overlap, then the value of is 0; when there is partial overlap, the value of is between 0 and 1.
[0185] In order to construct the time overlap edge between node and node a time correlation threshold needs to be preset in advance. When , then a time overlap edge is constructed between node and node .
[0186] Calculate the weight of the conflict detection edge and construct the conflict detection edge. The calculation formula of the weight of the conflict detection edge is:
[0187]
[0188] where and are weight parameters, is the score for the degree of intention conflict display, is the score for the degree of conflict of the material size ratio, is the score for the degree of conflict of the acquisition terminal source on the host side.
[0189] The conflict detection edge is used to describe the possibility that there is resource competition, spatial occlusion or semantic repetition in the display between node and node . It does not represent semantic relevance, but indicates whether it is possible to be displayed simultaneously. Constructing the conflict edge can help the fusion strategy avoid conflicts when generating the display decision scheme. For example:
[0190] The same product video and picture should not be displayed mainly at the same time;
[0191] Two large-window materials should not appear in adjacent areas;
[0192] The same modal type and the same terminal content should be avoided from being presented repeatedly.
[0193] Therefore, the existence of the conflict detection edge and the size of the conflict detection edge weight directly affect the penalty term or exclusion strategy in the fusion strategy and the spatial layout scheme.
[0194] Step S402: Based on the semantic association edges, perform material priority sorting. Use a clustering algorithm to cluster and group each of the nodes according to the weights of the semantic association edges of the nodes, and determine the display priority order of each clustering group set according to the importance scores corresponding to each node.
[0195] In this step, according to the weights of the semantic association edges, use a clustering algorithm to cluster and group each terminal data, and sort the materials within each clustering group according to the importance scores of the terminal data.
[0196] Step S403: Based on the time overlap edges, perform layout allocation. According to the result of the material priority sorting and the weights of the time overlap edges between the nodes, allocate the optimal display timing sequence and layout area for the nodes.
[0197] In this step, calculate the time correlation degree between the terminal data according to the weights of the time overlap edges, and combine the priority sorting result. Use a greedy algorithm to allocate appropriate display time periods and preliminary layout positions for each node to ensure the temporal connection and layout rationality of multi-source materials.
[0198] Step S404: Based on the conflict detection edge weights and the material size and ratio information, perform local conflict detection.
[0199] In this step, combine the conflict detection edge weights between the terminal data and the size and ratio information of each node to detect local conflict situations such as display intention conflicts, size ratio conflicts, or source terminal conflicts in the display layout, and provide a basis for layout optimization and interaction definition in subsequent display decisions.
[0200] Step S405: Based on the weights of the nodes and associated edges of the weighted attribute graph, construct a fusion optimization objective function, and automatically generate a display decision scheme that meets the requirements of display priority, layout timing, and conflict avoidance based on the fusion optimization objective function.
[0201] In this step, the calculation formula of the fusion optimization objective function is:
[0202]
[0203] where, is the combination of the display intention weight and the importance score of the node , is the set of nodes, is the set of time overlap edges, is the node and the node 's time overlap degree, is the set of conflict detection edges, is to display the conflict score, and is the penalty term coefficient.
[0204] The described display decision scheme includes:
[0205] The display order of multi-source materials is sorted according to the importance score and semantic clustering result in the fusion objective function.
[0206] The spatial layout scheme of multi-source materials searches and allocates spatial regions based on the greedy algorithm on the basis of no time conflict.
[0207] The interaction strategy of multi-source materials formulates interaction display or replacement rules according to the conflict edge penalty value and material size and ratio information.
[0208] Figure 5 is a schematic flow chart of multi-audience terminal adaptation and layout rendering shown by this application according to an exemplary embodiment. As Figure 5 shown, the method provided by this embodiment includes:
[0209] Step S501: Obtain the adaptation information of the audience terminal, where the adaptation information includes device characteristics, display capabilities, interaction methods, network environment, and system preference settings.
[0210] Step S502: Dynamically adjust the spatial layout, rendering method, and interaction strategy of multi-source materials based on the adaptation information of each audience terminal in combination with the display decision scheme.
[0211] Step S503: Perform rendering and interaction optimization according to the adjusted scheme.
[0212] Figure 6 is a schematic structural diagram of a video content generation system for multi-source material fusion shown by this application according to an exemplary embodiment. As Figure 6 shown, the video content generation system 600 for multi-source material fusion provided by this embodiment includes:
[0213] The multi-source material collection and preprocessing module 610 collects multiple terminal data on the host side in real time, performs modal classification, global timestamp annotation, format standardization processing, and structure unification processing on each piece of terminal data, and performs display intention label annotation on the terminal data based on the display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality.
[0214] The multi-modal semantic analysis and importance scoring module 620 inputs the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extracts semantic features, and performs importance scoring.
[0215] The multi-source material fusion and layout decision-making module 630 constructs a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance score, and automatically generates a display decision-making scheme according to the weighted attribute graph by adopting a fusion strategy.
[0216] The viewer terminal adaptation and layout rendering module 640 dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision-making scheme in combination with the screen attributes and interaction capabilities of the viewer terminal.
[0217] Figure 7 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application. As Figure 7 shown, an electronic device 700 provided in this embodiment includes: a processor 701 and a memory 702; wherein:
[0218] The memory 702 is used to store a computer program, and this memory can also be flash (flash memory).
[0219] The processor 701 is used to execute the execution instructions stored in the memory to implement each step in the above method. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0220] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0221] When the memory 702 is a device independent of the processor 701, the electronic device 700 may further include:
[0222] A bus 703 for connecting the memory 702 and the processor 701.
[0223] This embodiment also provides a readable storage medium, in which a computer program is stored. When at least one processor of the electronic device executes this computer program, the electronic device executes the methods provided by the above various embodiments.
[0224] This embodiment also provides a program product, which includes a computer program, and this computer program is stored in a readable storage medium. At least one processor of the electronic device can read this computer program from the readable storage medium, and at least one processor executes this computer program to enable the electronic device to implement the methods provided by the above various embodiments.
[0225] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.
[0226] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for generating video content by fusing multi-source materials, characterized in that, Including: Multi-source material collection and preprocessing, which involves real-time collection of multiple terminal data on the host side, performing modal classification, global timestamp annotation, format standardization processing, and structure unification processing for each piece of the terminal data, and annotating display intention tags for the terminal data based on a display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality; Multi-modal semantic analysis and importance scoring, which input the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extract semantic features, and perform importance scoring; Multi-source material fusion and layout decision-making, which constructs a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring, and automatically generates a display decision-making scheme using a fusion strategy according to the weighted attribute graph, including: Constructing a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring. The weighted attribute graph includes nodes and associated edges. Each node represents a terminal data unit, and the attributes of the node include a node unique identifier, a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intention tag, metadata information, material size and ratio information, a semantic feature vector, and an importance scoring. The associated edges include semantic association edges, time overlap edges, and conflict detection edges; Performing material priority sorting based on the semantic association edges, using a clustering algorithm to cluster and group each node according to the weight of the semantic association edges of the nodes, and determining the display priority order of each clustering group set according to the importance scoring corresponding to each node; Performing layout allocation based on the time overlap edges, and allocating the optimal display timing and layout area for the nodes according to the result of the material priority sorting and the weight of the time overlap edges between the nodes; Performing local conflict detection based on the weight of the conflict detection edges and the material size and ratio information; Constructing a fusion optimization objective function based on the weights of the nodes and associated edges of the weighted attribute graph, and automatically generating a display decision-making scheme that meets the requirements of display priority, layout timing, and conflict avoidance based on the fusion optimization objective function; The display decision-making scheme includes: The display order of multi-source materials, which is sorted according to the importance score and semantic clustering result in the fusion objective function; The spatial layout scheme of multi-source materials, which searches and allocates spatial regions based on a greedy algorithm on the basis of no time conflict; The interaction strategy of multi-source materials, which formulates interaction display or replacement rules according to the conflict edge penalty value and the material size and ratio information; Audience terminal adaptation and layout rendering, which dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision-making scheme in combination with the screen attributes and interaction capabilities of the audience terminal.
2. The method for generating video content by fusing multi-source materials according to claim 1, wherein The multi-source material collection and preprocessing includes: The server receives the terminal data uploaded from multiple acquisition terminals and annotates the modal type tags according to the modal type; The server marks a global timestamp for each piece of the terminal data according to the modal timestamp marking strategy; For the terminal data with the global timestamp marked, perform format standardization processing according to its modal type, and convert it into an internally unified data representation format; For the terminal data with the format standardized, perform structure unification processing according to its modal type, so that the terminal data of the same modality has a unified structure specification; Based on the display intention marking strategy, mark display intention tags for the terminal data with the structure unified, including: Extract display intention feature information from the terminal data, and the display intention feature information includes modal type, the source identifier of the host side acquisition terminal, metadata information, data structure fields, and material size and ratio information; Compare the display intention feature information with the preset matching rules; Mark the corresponding display intention tag for the terminal data according to the comparison result.
3. The video content generation method for multi-source material fusion according to claim 2, wherein The modal timestamp marking strategy includes: The server maintains a unified global timestamp benchmark; When the server receives terminal data uploaded by multiple acquisition terminals, record the reception time of each piece of the terminal data respectively; Maintain the one-way network delay of each acquisition terminal; Calculate the global timestamp of each piece of the terminal data based on the one-way network delay; Mark the global timestamp for each piece of the terminal data according to the stamping rules, and the stamping rules include: When the terminal data is in video modality, mark the global timestamp for each frame of video data; When the terminal data is in audio modality, mark the global timestamp for the audio data according to the first preset period; When the terminal data is in text modality, mark the global timestamp for each piece of text; When the terminal data is in picture modality, mark the global timestamp for each picture; When the terminal data is in sensor modality, mark the global timestamp at each update.
4. The method for generating video content by fusing multi-source materials according to claim 3, characterized in that, The multi-modal semantic analysis and importance scoring include: Construct a modal data sequence, and the modal data sequence includes a unique data sequence identifier, modal type, global timestamp, the source identifier of the host side acquisition terminal, display intention tag, metadata information, material size and ratio information, and modality-specific content; Send the modal data sequence to the multi-modal semantic analysis engine for processing according to the modal type respectively; The multi-modal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors; Perform importance scoring according to the collected terminal data, display intention tags, and semantic feature vectors.
5. The method for generating video content by fusing multi-source materials according to claim 1, characterized in that The construction of the associated edges includes: Calculate the weight of the semantic association edge and construct the semantic association edge; Calculate the weight of the time overlap edge and construct the time overlap edge; Calculate the weight of the conflict detection edge and construct the conflict detection edge.
6. The method for generating video content by fusing multi-source materials according to claim 1, characterized in that, The viewer terminal adaptation and layout rendering include: Obtain the adaptation information of the viewer terminal, and the adaptation information includes device characteristics, display capabilities, interaction methods, network environment, and system preference settings; Based on the adaptation information of each viewer terminal and in combination with the display decision scheme, dynamically adjust the spatial layout, rendering method, and interaction strategy of multi-source materials; Perform rendering and interaction optimization according to the adjusted plan.
7. A video content generation system for multi-source material fusion, characterized in that, Including: A multi-source material collection and preprocessing module that collects multiple terminal data on the host side in real time, performs modal classification, global timestamp annotation, format standardization processing, and structure unification processing for each piece of the terminal data, and annotates display intention tags for the terminal data based on a display intention annotation strategy. The modal types of the terminal data include video modality, audio modality, text modality, picture modality, and sensor modality; A multi-modal semantic analysis and importance scoring module that inputs the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extracts semantic features, and performs importance scoring; A multi-source material fusion and layout decision-making module that constructs a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring, and automatically generates a display decision-making plan using a fusion strategy according to the weighted attribute graph, including: Construct a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector, and importance scoring. The weighted attribute graph includes nodes and associated edges. Each node represents a terminal data unit. The attributes of the node include a node unique identifier, a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intention tag, metadata information, material size and ratio information, a semantic feature vector, and an importance scoring. The associated edges include semantic association edges, time overlap edges, and conflict detection edges; Perform material priority sorting based on the semantic association edges, use a clustering algorithm to group each node according to the weight of the semantic association edges of the node, and determine the display priority order of each clustering group set according to the importance scoring corresponding to each node; Perform layout allocation based on the time overlap edges, and allocate the optimal display timing and layout area for the nodes according to the result of the material priority sorting and the weight of the time overlap edges between the nodes; Perform local conflict detection based on the weight of the conflict detection edges and the material size and ratio information; Construct a fusion optimization objective function based on the weights of the nodes and associated edges of the weighted attribute graph, and automatically generate a display decision-making plan that meets the requirements of display priority, layout timing, and conflict avoidance based on the fusion optimization objective function; The display decision-making plan includes: The display order of multi-source materials, sorted according to the importance score and semantic clustering result in the fusion objective function; A spatial layout plan for multi-source materials, searching and allocating spatial regions based on a greedy algorithm on the basis of no time conflict; An interaction strategy for multi-source materials, formulating interaction display or replacement rules according to the conflict edge penalty value and the material size and ratio information; A viewer terminal adaptation and layout rendering module that dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision-making plan in combination with the screen attributes and interaction capabilities of the viewer terminal.
8. An electronic device, characterized in that, Including: A processor; And, A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Visual monitoring system and method for electric power facilities
CN119628221A