Multi-source material fused video content generation method, system, equipment and medium

By collecting, preprocessing, semantic analysis and fusion decisions on multi-source materials, the problem of lack of deep understanding and adaptive adjustment of multi-source material live broadcast systems in the existing technology is solved, and efficient and flexible multi-source material display is achieved.

CN119996786AActive Publication Date: 2025-05-13SHANGHAI YUNTI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510450251.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing multi-source live broadcast system lacks deep understanding of content semantics and adaptive adjustment capabilities, resulting in a fixed fusion output method, making it difficult to adapt to dynamic changes in complex scenarios, and has poor real-time performance.

Method used

Through multi-source material collection and preprocessing, multiple terminal data on the anchor side are collected in real time, and modal classification, global timestamp annotation, format and structure are uniformly processed. Then, the multimodal semantic analysis engine is used to extract semantic features and perform importance scoring, and a weighted attribute diagram is constructed to automatically generate a display decision plan, and finally adapt and layout rendering is applied to the audience terminal.

Benefits of technology

It realizes synchronization, intelligent analysis and adaptive display of multi-modal live broadcast data, improves the integration ability and display flexibility of the multi-source material live broadcast system, and can dynamically adapt to changes in complex scenarios and realize personalized optimization for different terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996786A_ABST
    Figure CN119996786A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source material fused video content generation method, system and device and a medium. The method comprises the steps of collecting multiple pieces of terminal data of an anchor side in real time, performing modal classification, global timestamp labeling, format standardization processing and structure unification processing on each piece of terminal data, and performing display intention label labeling on the terminal data based on a display intention labeling strategy. And inputting the preprocessed terminal data into a multi-modal semantic analysis engine according to the modal type, extracting semantic features, and performing importance scoring. And constructing a weighted attribute graph based on the preprocessed terminal data, the modal data sequence, the semantic feature vector and the importance score, and automatically generating a display decision scheme by adopting a fusion strategy according to the weighted attribute graph. And dynamically generating a visual layout and completing display rendering of the multi-source materials based on the display decision scheme in combination with screen attributes and interaction capabilities of the audience terminals. According to the invention, synchronization, intelligent analysis and adaptive display of multi-modal live broadcast data can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to data processing technology, and in particular to a method, system, device and medium for generating video content by fusing multiple source materials. Background Art

[0002] With the rapid development of live video technology, live content is becoming increasingly rich and the sources of materials are becoming increasingly diverse. Especially in scenarios such as e-commerce live broadcast, event live broadcast, news reporting, virtual studios, etc., it is often necessary to simultaneously integrate and display video, audio, text, images and sensor data from multiple acquisition terminals. How to achieve real-time integration and intelligent display of multimodal materials while ensuring synchronization has become a core challenge to improve live broadcast quality and user experience.

[0003] However, the existing multi-source live broadcast system mainly relies on preset templates or manual intervention to arrange the materials, lacking the deep understanding of content semantics and adaptive adjustment capabilities, resulting in a fixed fusion output method that is difficult to adapt to the dynamic changes of complex scenes, and poor real-time performance, affecting the live broadcast experience. In addition, the diversity of terminal devices (such as mobile phones, tablets, VR devices) and the variability of user interaction methods make it difficult for traditional static fusion methods to meet personalized optimization needs. Summary of the invention

[0004] The present application provides a method, system, device and medium for generating video content by fusing multiple source materials, so as to solve the problems in the existing multi-source material live broadcast system, which lacks the deep understanding of content semantics and the adaptive adjustment capability, resulting in the fixed fusion output mode, making it difficult to adapt to the dynamic changes of complex scenes, and having poor real-time performance.

[0005] In a first aspect, the present application provides a method for generating video content by fusing multiple source materials, comprising: Multi-source material collection and preprocessing, real-time collection of multiple terminal data on the anchor side, modality classification, global timestamp annotation, format standardization and structural unification for each terminal data, and annotation of display intent labels for the terminal data based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality; Multimodal semantic analysis and importance scoring: inputting the pre-processed terminal data into a multimodal semantic analysis engine according to the modality type, extracting semantic features, and performing importance scoring; Multi-source material fusion and layout decision, constructing a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generating a display decision plan based on the weighted attribute graph using a fusion strategy; Audience terminal adaptation and layout rendering, based on the display decision plan, combined with the screen attributes and interactive capabilities of the audience terminal, dynamically generates a visual layout and completes the display rendering of multi-source materials.

[0006] In a possible design, the multi-source material collection and preprocessing includes: The server receives terminal data uploaded from multiple acquisition terminals, and marks the modality type label according to the modality type; The server annotates each of the terminal data with a global timestamp according to a modality timestamp annotation strategy; For terminal data with global timestamps, the format is standardized according to its modality type and converted into an internal unified data representation format; For the terminal data whose format has been standardized, the structure is unified according to its modality type, so that the terminal data of the same modality has a unified structural specification; Based on the display intent labeling strategy, the terminal data that has been structured uniformly is labeled with display intent labels, including: Extracting display intention feature information from the terminal data, the display intention feature information including modality type, source identifier of the acquisition terminal on the anchor side, metadata information, data structure field, and material size and ratio information; Comparing the display intention feature information with a preset matching rule; The terminal data is labeled with a corresponding display intent label according to the comparison result.

[0007] In a possible design, the modality timestamp annotation strategy includes: The server maintains a unified global timestamp reference; When the server receives terminal data uploaded by multiple acquisition terminals, it records the receiving time of each terminal data respectively; Maintaining the one-way network delay of each of the acquisition terminals; Calculate the global timestamp of each terminal data based on the one-way network delay; Performing a global timestamp for each terminal data according to a stamping rule, wherein the stamping rule includes: When the terminal data is in video mode, a global timestamp is marked for each frame of video data; When the terminal data is in audio mode, marking the audio data with a global timestamp according to a first preset period; When the terminal data is in text mode, a global timestamp is marked for each text; When the terminal data is in picture mode, a global timestamp is marked for each picture; When the terminal data is in sensor mode, a global timestamp is added at each update.

[0008] In a possible design, the multimodal semantic analysis and importance scoring includes: Constructing a modal data sequence, wherein the modal data sequence includes a data sequence unique identifier, a modal type, a global timestamp, a source identifier of a collection terminal on the host side, a display intent label, metadata information, material size and ratio information, and modality-specific content; Sending the modal data sequences to a multimodal semantic analysis engine for processing according to the modality type; The multimodal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors; Importance scoring is performed based on the collected terminal data, display intent labels, and semantic feature vectors.

[0009] In a possible design, the multi-source material fusion and layout decision include: A weighted attribute graph is constructed based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, wherein the weighted attribute graph includes nodes and associated edges, each of the nodes represents a terminal data unit, and the attributes of the node include a node unique identifier, a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intent label, metadata information, material size and ratio information, a semantic feature vector and an importance score, and the associated edges include a semantic association edge, a time overlap edge and a conflict detection edge; The materials are prioritized based on the semantically associated edges, and each node is clustered and grouped based on the semantically associated edge weights of the nodes using a clustering algorithm, and the display priority order of each cluster grouping set is determined based on the importance score corresponding to each node; Perform layout allocation based on the time overlapping edges, and allocate optimal display timing and layout area to the nodes according to the result of material priority sorting and the time overlapping edge weights between nodes; Performing local conflict detection based on the conflict detection edge weights and the material size and ratio information; Based on the weights of the nodes and associated edges of the weighted attribute graph, a fusion optimization objective function is constructed, and based on the fusion optimization objective function, a display decision plan that meets the display priority, layout timing, and conflict avoidance requirements is automatically generated; The display decision plan includes: The display order of multi-source materials is sorted according to the importance score and semantic clustering results in the fusion objective function; The spatial layout scheme of multi-source materials is based on the search and allocation of spatial areas based on the greedy algorithm on the basis of non-conflict in time; The interactive strategy for multi-source materials formulates interactive display or replacement rules based on the conflict edge penalty value and material size and ratio information.

[0010] In a possible design, the construction of the association edge includes: Calculating the weight of the semantic association edge and constructing the semantic association edge; Calculating the time overlapping edge weights and constructing the time overlapping edges; The conflict detection edge weight is calculated, and the conflict detection edge is constructed.

[0011] In a possible design, the viewer terminal adaptation and layout rendering include: Acquire the adaptation information of the viewer terminal, the adaptation information including device characteristics, display capabilities and interaction methods, network environment and system preference settings; Based on the adaptation information of each of the terminals combined with the display decision plan, dynamically adjust the spatial layout, rendering method and interaction strategy of the multi-source materials; Perform rendering and interaction optimization according to the adjusted plan.

[0012] In a second aspect, the present application provides a video content generation system for multi-source material fusion, comprising: The multi-source material collection and preprocessing module collects multiple terminal data from the anchor side in real time, performs modality classification, global timestamp annotation, format standardization and structural unification for each terminal data, and annotates the terminal data with display intent labels based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality. A multimodal semantic analysis and importance scoring module, which inputs the pre-processed terminal data into a multimodal semantic analysis engine according to the modality type, extracts semantic features, and performs importance scoring; The multi-source material fusion and layout decision module constructs a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generates a display decision plan based on the weighted attribute graph using a fusion strategy; The audience terminal adaptation and layout rendering module dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision plan and in combination with the screen attributes and interactive capabilities of the audience terminal.

[0013] In a third aspect, the present application provides an electronic device, including: processor; and, A memory, configured to store executable instructions of the processor; The processor is configured to perform any possible method described in the first aspect by executing the executable instructions.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement any possible method described in the first aspect.

[0015] The video content generation method, system, device and medium of multi-source material fusion provided in this application can realize the synchronization, intelligent analysis and adaptive display of multimodal live broadcast data. The temporal consistency of multiple terminal data is ensured by global timestamp annotation, key information is extracted by multimodal semantic analysis, the sorting and layout of multi-source materials are optimized based on the fusion strategy with weighted attribute graph, and the display effect is dynamically adjusted in combination with the adaptation information of the audience terminal. This application breaks through the limitations of the existing multi-source live broadcast system with fixed fusion mode and poor real-time performance, improves the fusion ability and display flexibility of the multi-source material live broadcast system, enables the live broadcast system to dynamically adapt to complex scene changes, and realizes personalized optimization for different terminals. And this application has high real-time performance and strong scalability, and can be widely used in various scenarios such as e-commerce live broadcast, event live broadcast, news reporting, etc., significantly improving the live broadcast quality and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] Figure 1 It is a flowchart of a method for generating video content by fusing multiple source materials according to an exemplary embodiment of the present application; Figure 2 is a schematic diagram of a process of multi-source material collection and preprocessing according to an exemplary embodiment of the present application; Figure 3 is a flowchart of multimodal semantic analysis and importance scoring according to an exemplary embodiment of the present application; Figure 4 is a flowchart of multi-source material fusion and layout decision-making according to an exemplary embodiment of the present application; Figure 5 is a flow chart of viewer terminal adaptation and layout rendering according to an exemplary embodiment of the present application; Figure 6 is a structural schematic diagram of a video content generation system for multi-source material fusion according to an exemplary embodiment of the present application; Figure 7It is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application.

[0018] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0019] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0020] Figure 1 FIG. 1 is a flow chart of a method for generating video content by fusing multiple source materials according to an exemplary embodiment of the present application. Figure 1 As shown, the method provided in this embodiment includes: Step S101: Multi-source material collection and preprocessing, real-time collection of multiple terminal data on the anchor side, modality classification, global timestamp annotation, format standardization and structure unification for each terminal data, and annotation of display intent labels for the terminal data based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality.

[0021] In this step, the server is set up to collect terminal data uploaded by multiple collection terminals on the anchor side in real time to adapt to the current complex live broadcast scene requirements, including large-scale events, live broadcasts with goods, sports events, virtual anchors, and multi-scene serial live broadcasts.

[0022] For the above complex live broadcast scenarios, the terminal data collected by the anchor side includes the following five modal types: Video mode: Through multiple camera terminals (such as main camera, side view, top view, mobile tracking device), multi-view synchronous acquisition is realized, so as to build a panoramic live broadcast experience and enhance the audience's immersion. This type of acquisition method is suitable for scenes such as sports commentary, live broadcast with goods, and talent show.

[0023] Audio mode: The audio source is usually independent of the video acquisition terminal, including background music, voice assistant output, remote guest voice and other multi-channel audio content. It is suitable for live concerts, live broadcasts by virtual anchors, game commentary and remote connections.

[0024] Image mode: refers to static image content uploaded, captured or generated by the host, which is often used to display background images, PPT courseware, graphic handouts, illustrations, cover illustrations or product details, etc. It is suitable for educational live broadcasts, e-commerce live broadcasts, corporate press conferences and other graphic and text-based scenarios.

[0025] Text mode: includes real-time interactive information from the audience, such as bullet screen information, audience comments, and chat room messages, which is used to enhance interactivity and semantically supplement the live broadcast content.

[0026] Sensor modality: It comes from the sensor devices and peripheral operation data worn or used by the anchor, including body information, heart rate monitoring, posture / motion capture, mouse trajectory, keyboard buttons, game controller input, etc. It is widely used in scenarios such as virtual anchor live broadcast, game live broadcast, fitness live broadcast and educational interactive live broadcast.

[0027] Step S102: multimodal semantic analysis and importance scoring, inputting the pre-processed terminal data into a multimodal semantic analysis engine according to the modality type, extracting semantic features, and performing importance scoring.

[0028] In this step, the multimodal semantic analysis engine is composed of multiple modular artificial intelligence models, each of which performs customized semantic feature extraction for terminal data of a modality type. The modality types are as follows: Video modality: Through visual artificial intelligence models, such as 3D convolutional neural networks and Transformer-based frame-level analysis models, semantic features such as scene type, host actions, interactive behaviors, and visual focus areas in video clips are extracted to determine the display value, interactive density, and focus content of the current video content.

[0029] Audio modality: Through audio AI models such as speech recognition, speech emotion recognition, and sound source classification, such as Wav2Vec 2.0 and Whisper, speech recognition models are used to extract features such as speech content text, speaker identity, tone and emotional state, and audio type (speech / music / noise) to evaluate the emotional intensity, semantic density, and presentation relevance of the speech content.

[0030] Image modality: Through image classification, object detection, OCR recognition and other image AI models, such as the ResNet image classification model and the YOLOv5 object detection model, the visual subject, text content, style features and structural information in the image are analyzed, and the image category (such as courseware images, product images, background images, etc.) and its key semantic content are extracted to determine its display intent and adapt the layout.

[0031] Text modality: Through natural language processing models, such as the BERT semantic understanding model and the TextCNN text classification model, semantic understanding of text data such as barrage and comments is performed to extract features such as keywords, topic orientation, emotional tendency, frequency weight, etc., which are used to assist in hot topic identification, display intent matching and real-time interaction guidance.

[0032] Sensor modality: Through time series modeling, action recognition or behavior prediction models, such as the HAR-LSTM behavior recognition model and the ST-GCN spatiotemporal graph convolutional network model, the data collected from somatosensory devices, heart rate sensors, keyboard and mouse trajectories, controllers, etc. are used to perform behavior state recognition, operation intention prediction, physiological signal analysis, etc., and extract the host's current action / state information to determine key interactive events or virtual image driving signals.

[0033] Step S103: Multi-source material fusion and layout decision, constructing a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generating a display decision plan using a fusion strategy based on the weighted attribute graph.

[0034] Step S104: audience terminal adaptation and layout rendering, based on the display decision plan, combined with the screen attributes and interactive capabilities of the audience terminal, dynamically generate a visual layout and complete the display rendering of multi-source materials.

[0035] Figure 2 FIG. 1 is a flow chart of multi-source material collection and preprocessing according to an exemplary embodiment of the present application. Figure 2 As shown, the method provided in this embodiment includes: Step S201: The server receives terminal data uploaded from multiple acquisition terminals, and marks the modality type label according to the modality type.

[0036] In this step, the terminal data is modally classified according to its type, and each piece of data is labeled with a corresponding modality type label, which includes video modality, audio modality, text modality, image modality, and sensor modality.

[0037] Step S202: the server annotates each terminal data with a global timestamp according to a modality timestamp annotation strategy.

[0038] In this step, the modal timestamp annotation strategy includes: The server maintains a unified global timestamp reference to eliminate clock deviations between different acquisition terminals, achieve time synchronization and unified alignment of multimodal data, and ensure the timing accuracy and consistency of subsequent fusion and display.

[0039] When the server receives terminal data uploaded by multiple acquisition terminals, it records the receiving time of each terminal data respectively for calculating the one-way network delay.

[0040] Maintain the one-way network delay of each acquisition terminal, and the calculation formula of the one-way network delay is:

[0041] in, is the one-way network delay of the i-th acquisition terminal, is the round-trip delay measured by the i-th acquisition terminal for the kth time, is the weight of the kth measurement, For the recent Sliding window weighted average of measurements.

[0042] The global timestamp of each terminal data is calculated based on the one-way network delay, and the calculation formula of the global timestamp is:

[0043] in, is the current global timestamp of the i-th terminal data, is the current local timestamp of the i-th terminal data, It is the clock offset between the server and the acquisition terminal; The calculation formula is:

[0044] in, The receiving time of the i-th terminal data recorded by the server; Performing a global timestamp for each terminal data according to a stamping rule, wherein the stamping rule includes: When the terminal data is in video mode, a global timestamp is marked for each frame of video data.

[0045] When the terminal data is in audio mode, a global timestamp is marked for the audio data according to a first preset period.

[0046] When the terminal data is in text mode, a global timestamp is marked for each text.

[0047] When the terminal data is in picture mode, a global timestamp is marked for each picture.

[0048] When the terminal data is in sensor mode, a global timestamp is added at each update.

[0049] Step S203: The terminal data marked with the global timestamp is format-standardized according to its modality type and converted into an internal unified data representation format.

[0050] In this step, the data formats uploaded by different acquisition terminals are different, as shown in Table 1: Table 1 Various data formats uploaded by the acquisition terminal

[0051] Therefore, all terminal data need to be converted into an internal unified data representation format so that the terminal data can be input into the multimodal semantic analysis engine for subsequent parsing, decoding, and feature extraction.

[0052] Step 204: The terminal data whose formats have been standardized are subjected to structural unification processing according to their modality types, so that the terminal data of the same modality have a unified structural specification.

[0053] In this step, different acquisition terminals may use different encoding, packaging, resolution or meta-information formats for the terminal data of the same modality type, as shown in Table 2: Table 2 Various data structures uploaded by the acquisition terminal

[0054] Therefore, it is necessary to standardize the field structure, label information and data organization of the standardized terminal data according to its modality type, so that the terminal data of the same modality has a consistent data structure, so that the terminal data can be input into the multimodal semantic analysis engine for subsequent parsing, decoding and feature extraction.

[0055] Step S205: labeling the terminal data that has been structured uniformly with display intent labels based on the display intent labeling strategy.

[0056] In this step, the intent labels are displayed, including: Extract the display intention feature information from the terminal data, the display intention feature information includes the modality type, the source identifier of the acquisition terminal on the anchor side, metadata information, data structure fields and material size and ratio information. Specifically as follows: Modality type: used to determine the analysis path of terminal data, including video modality, audio modality, image modality, text modality, and sensor modality.

[0057] The source identifier of the acquisition terminal on the anchor side: as the unique identification information of the terminal data source, it is used to distinguish the roles of the material content, including the number of the acquisition terminal, the upload channel ID, and the file path.

[0058] Metadata information: used for rule matching and unified structure judgment, including file name, path name, media information (format, duration, frame rate, sampling rate, resolution).

[0059] Data structure field: used to identify the material type and its content structure, including keywords and sensor types.

[0060] Material size and ratio information: used for display layout judgment in fusion strategy, including parameters such as width, height, and aspect ratio.

[0061] The display intention feature information is compared with the preset matching rules.

[0062] The terminal data is labeled with a corresponding display intent label according to the comparison result.

[0063] In this step, the display intent tag is used to indicate the functional role of the terminal data in the fusion display of multi-source materials. The display intent tag includes: Display intent type: used to identify the content role of terminal data in the integrated display, for example, the speaker image is marked as main_speaker, the product image is marked as product_image, and the sensor layer is marked as sensor_overlay.

[0064] Display priority: used to indicate the importance of the terminal data in the overall display, and to assist in content sorting and display conflict judgment in the fusion strategy.

[0065] Display method prompt: used to prompt the expected display form of the terminal data on the audience terminal interface, such as small window display, large screen display, icon overlay, etc., to support the rendering decision of the terminal adaptation and layout rendering module.

[0066] Figure 3 FIG. 1 is a flow chart of multimodal semantic analysis and importance scoring according to an exemplary embodiment of the present application. Figure 3 As shown, the method provided in this embodiment includes: Step S301: construct a modal data sequence, wherein the modal data sequence includes a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intent label, metadata information, material size and ratio information, and modality-specific content.

[0067] In this step, the modal data sequence is used as the standardized input data structure of the multimodal semantic analysis engine to achieve unified processing and model distribution of terminal data of different modal types in the semantic analysis stage. Due to the different terminal data structures and processing requirements of each modal type, it is necessary to construct corresponding standardized input formats according to the modal type and send them to the corresponding artificial intelligence model for semantic feature extraction. The modal data sequence includes the following fields: Data sequence unique identifier: A unique identifier generated for each modal data, used for data tracking and result association throughout the entire processing flow.

[0068] Modality type, global timestamp, source identifier of the acquisition terminal on the anchor side, display intent tag, metadata information, and material size and ratio information: all come from pre-processed terminal data and display intent tags.

[0069] Modality-specific content: a uniformly encapsulated package of raw material data used for analysis and processing of artificial intelligence models, including index identification information and raw data content. The raw data content is the terminal data content that has been processed through format standardization and structural unification.

[0070] Step S302: sending the modal data sequence to a multimodal semantic analysis engine for processing according to the modality type.

[0071] In this step, since the multimodal semantic analysis engine includes multiple artificial intelligence models corresponding to the modal types, the modal data sequences need to be routed and distributed according to the modal types, and sent to the corresponding models as standardized inputs for subsequent semantic analysis and processing.

[0072] Step S303: The multimodal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors.

[0073] In this step, the semantic feature vector will serve as the key basis for subsequent importance score calculation, weighted attribute graph construction and multi-source material fusion display, providing core support for fusion strategy and display decision-making plan. Depending on the modality type, the semantic feature vectors obtained after processing by the multimodal semantic analysis engine are also different. The semantic feature vectors of each modality type are as follows: Video modality: The semantic feature vector includes the scene label, action encoding, and visual emotion vector of the video modality.

[0074] Audio modality: The semantic feature vector includes the speech transcription, sentiment intensity score, speaking rate and loudness characteristics of the audio modality.

[0075] Text modality: The semantic feature vector includes keyword embedding, sentiment polarity labels, and hot word distribution features of the text modality.

[0076] Image modality: The semantic feature vector includes the subject object label, image quality score and visual style description of the image modality.

[0077] Sensor modality: The semantic feature vector includes the change trend vector, abnormal probability and state label of the sensor modality.

[0078] Step S304: performing importance scoring based on the collected terminal data, display intent labels, and semantic feature vectors.

[0079] In this step, the importance score is calculated as:

[0080] in, is the weight coefficient of the semantic feature, Score the semantic feature corresponding to the data unit uniquely identified as x in the data sequence, To show the weight coefficient of the intent label, The display intention score corresponding to the data unit uniquely identified as x in the data sequence, is the weight coefficient of real-time impact, is the real-time score corresponding to the data unit uniquely identified as x in the data sequence, is the weight coefficient of the source of the anchor terminal. Score the source terminal on the anchor side corresponding to the data unit with the unique identifier x in the data sequence; The calculation formula for semantic feature score is:

[0081] in, is the semantic feature vector corresponding to the data unit uniquely identified as x in the data sequence, is the standard reference eigenvector, is the vector inner product, Product of vector norms; The formula for calculating the impression intent score is:

[0082] in, To demonstrate the intent scoring function, is the display intent label corresponding to the data unit uniquely identified as x in the data sequence; The calculation formula for real-time score is:

[0083] in, is the time attenuation coefficient, is the current system time, The global timestamp corresponding to the data unit uniquely identified as x in the data sequence; The formula for the source terminal score on the anchor side is:

[0084] in, The weight scoring function for the source identification of the acquisition terminal on the broadcast side, The host-side source terminal identifier corresponding to the data unit whose unique identifier in the data sequence is x.

[0085] Figure 4 FIG. 1 is a flow chart of a method for generating video content by fusing multiple source materials according to an exemplary embodiment of the present application. Figure 4 As shown, the method provided in this embodiment includes: Step S401: construct a weighted attribute graph based on the preprocessed terminal data, modal data sequence, semantic feature vector and importance score, the weighted attribute graph includes nodes and associated edges, each of the nodes represents a terminal data unit, the attributes of the node include node unique identifier, data sequence unique identifier, modal type, global timestamp, anchor side acquisition terminal source identifier, display intent label, metadata information, material size and ratio information, semantic feature vector and importance score, the associated edges include semantic association edges, time overlap edges and conflict detection edges.

[0086] In this step, the construction of the associated edge includes: The semantic association edge weight is calculated to construct the semantic association edge. The calculation formula of the semantic association edge weight is:

[0087] in, is the node with the unique identifier i. is the node with the unique identifier j, for The corresponding semantic feature vector, for The corresponding semantic feature vector, is the vector inner product, is the product of the vector norms.

[0088] Semantic association edges are used to measure node and nodes The semantic similarity between The value range is .

[0089] when When it approaches 1, it means the node and nodes The semantic features of the two are highly similar, and they are close in content, theme or description, and tend to be classified into the same cluster group.

[0090] when When it approaches 0, it means the node and nodes The semantic features of the two are relatively similar, and they are quite different in content, subject or description, and are usually not grouped into the same cluster.

[0091] The time overlapping edge weight is calculated and the time overlapping edge is constructed. The calculation formula of the time overlapping edge weight is:

[0092] in, For Node The global timestamp of For Node The global timestamp of is the temporal proximity factor, For Node time interval, For Node time interval, is the time range overlap factor.

[0093] To measure the node and nodes The global timestamp proximity of is the time decay coefficient, which determines the influence of time difference on the weight of time overlapping edges.

[0094] if and As time approaches, Approaching 1; if and As the time difference increases, Approaching 0.

[0095] To measure the node and nodes The degree of overlap at the time end, when completely overlapped, then The value of is 1; when there is no overlap, The value of is 0; when there is partial overlap, The value of is a balance between 0 and 1.

[0096] To build the node and nodes Time overlap edges need to set a time correlation threshold in advance ,when When and nodes Construct temporal overlapping edges between them.

[0097] The conflict detection edge weight is calculated to construct the conflict detection edge. The calculation formula of the conflict detection edge weight is:

[0098] in, and is the weight parameter, Score the degree of display intention conflict, Score the degree of conflict in the size ratio of the material. Score the conflict level of the acquisition terminal sources on the host side.

[0099] Conflict detection edges are used to describe nodes and nodes There is a possibility of resource competition, spatial occlusion, or semantic duplication in display. It does not indicate semantic relevance, but whether it is possible that they cannot be displayed at the same time. Constructing conflict edges can help the fusion strategy avoid conflicts when generating display decision plans, for example: The same product video and picture should not be displayed at the same time; Two large window clips should not appear in adjacent areas; Avoid repeated presentation of the same modal type and the same terminal content.

[0100] Therefore, the existence of conflict detection edges and the weight of conflict detection edges directly affect the penalty term or exclusion strategy in the fusion strategy and spatial layout scheme.

[0101] Step S402: Prioritize the materials based on the semantically associated edges, use a clustering algorithm to cluster each node based on the semantically associated edge weights of the nodes, and determine the display priority order of each cluster grouping set based on the importance score corresponding to each node.

[0102] In this step, a clustering algorithm is used to cluster and group each terminal data according to the semantic association edge weight, and the materials within each cluster group are sorted according to the importance score of the terminal data.

[0103] Step S403: Perform layout allocation based on the time overlapping edges, and allocate optimal display timing and layout area to the nodes according to the result of material priority sorting and the time overlapping edge weights between nodes.

[0104] In this step, the temporal correlation between terminal data is calculated according to the weight of the time overlapping edges. Combined with the priority sorting results, a greedy algorithm is used to assign a suitable display period and preliminary layout position to each node to ensure the temporal connectivity of multi-source materials and the rationality of the layout.

[0105] Step S404: Perform local conflict detection based on the conflict detection edge weights and the material size and ratio information.

[0106] In this step, the conflict detection edge weights between terminal data and the size and proportion information of each node are combined to detect whether there are local conflicts such as display intention conflicts, size proportion conflicts or source terminal conflicts in the display layout, and provide a basis for layout optimization and interaction definition in subsequent display decisions.

[0107] Step S405: Based on the weights of the nodes and associated edges of the weighted attribute graph, a fusion optimization objective function is constructed, and based on the fusion optimization objective function, a display decision plan that meets the display priority, layout timing and conflict avoidance requirements is automatically generated.

[0108] In this step, the calculation formula of the fusion optimization objective function is:

[0109] in, To combine display intent weight and nodes The importance rating of is a collection of nodes, is the set of time overlapping edges, For Node and nodes The degree of temporal overlap, is the conflict detection edge set, To display the conflict score, and is the penalty coefficient.

[0110] The display decision plan includes: The display order of multi-source materials is sorted according to the importance score in the fusion objective function and the semantic clustering results.

[0111] The spatial layout scheme of multi-source materials allocates spatial areas based on greedy algorithm search on the basis of no time conflict.

[0112] The interactive strategy for multi-source materials formulates interactive display or replacement rules based on the conflict edge penalty value and material size and ratio information.

[0113] Figure 5FIG. 1 is a flow chart of multi-viewer terminal adaptation and layout rendering according to an exemplary embodiment of the present application. Figure 5 As shown, the method provided in this embodiment includes: Step S501: Acquire the adaptation information of the viewer terminal, wherein the adaptation information includes device characteristics, display capabilities and interaction methods, network environment and system preference settings.

[0114] Step S502: dynamically adjusting the spatial layout, rendering method and interaction strategy of the multi-source materials based on the adaptation information of each of the terminals in combination with the display decision solution.

[0115] Step S503: Execute rendering and interaction optimization according to the adjusted solution.

[0116] Figure 6 FIG. 1 is a schematic diagram of a video content generation system for fusion of multiple source materials according to an exemplary embodiment of the present application. Figure 6 As shown, the video content generation system 600 for multi-source material fusion provided in this embodiment includes: The multi-source material collection and preprocessing module 610 collects multiple terminal data on the anchor side in real time, performs modality classification, global timestamp annotation, format standardization and structure unification for each terminal data, and annotates the terminal data with display intent labels based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality.

[0117] The multimodal semantic analysis and importance scoring module 620 inputs the pre-processed terminal data into the multimodal semantic analysis engine according to the modality type, extracts semantic features, and performs importance scoring.

[0118] The multi-source material fusion and layout decision module 630 constructs a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generates a display decision plan using a fusion strategy based on the weighted attribute graph.

[0119] The viewer terminal adaptation and layout rendering module 640 dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision plan and in combination with the screen attributes and interactive capabilities of the viewer terminal.

[0120] Figure 7 is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application. Figure 7 As shown, an electronic device 700 provided in this embodiment includes: a processor 701 and a memory 702; wherein: The memory 702 is used to store computer programs. The memory may also be a flash memory.

[0121] The processor 701 is used to execute the execution instructions stored in the memory to implement each step in the above method. For details, please refer to the relevant description in the above method embodiment.

[0122] Optionally, the memory 702 may be independent or integrated with the processor 701 .

[0123] When the memory 702 is a device independent of the processor 701, the electronic device 700 may further include: The bus 703 is used to connect the memory 702 and the processor 701 .

[0124] This embodiment further provides a readable storage medium, in which a computer program is stored. When at least one processor of an electronic device executes the computer program, the electronic device executes the methods provided in the above-mentioned various implementation modes.

[0125] This embodiment also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device implements the methods provided in the above various embodiments.

[0126] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0127] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for generating video content by fusing multiple source materials, characterized in that: include: Multi-source material collection and preprocessing, real-time collection of multiple terminal data on the anchor side, modality classification, global timestamp annotation, format standardization and structural unification for each terminal data, and annotation of display intent labels for the terminal data based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality; Multimodal semantic analysis and importance scoring: inputting the pre-processed terminal data into a multimodal semantic analysis engine according to the modality type, extracting semantic features, and performing importance scoring; Multi-source material fusion and layout decision, constructing a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generating a display decision plan based on the weighted attribute graph using a fusion strategy; Audience terminal adaptation and layout rendering, based on the display decision plan, combined with the screen attributes and interactive capabilities of the audience terminal, dynamically generates a visual layout and completes the display rendering of multi-source materials.

2. The method for generating video content by fusing multiple source materials according to claim 1, characterized in that: The multi-source material collection and preprocessing includes: The server receives terminal data uploaded from multiple acquisition terminals, and marks the modality type label according to the modality type; The server annotates each of the terminal data with a global timestamp according to a modality timestamp annotation strategy; For terminal data with global timestamps, the format is standardized according to its modality type and converted into an internal unified data representation format; For the terminal data whose format has been standardized, the structure is unified according to its modality type, so that the terminal data of the same modality has a unified structural specification; Based on the display intent labeling strategy, the terminal data that has been structured uniformly is labeled with display intent labels, including: Extracting display intention feature information from the terminal data, the display intention feature information including modality type, source identifier of the acquisition terminal on the anchor side, metadata information, data structure field, and material size and ratio information; Comparing the display intention feature information with a preset matching rule; The terminal data is labeled with a corresponding display intent label according to the comparison result.

3. The method for generating video content by fusing multiple source materials according to claim 2, characterized in that: The modal timestamp annotation strategy includes: The server maintains a unified global timestamp reference; When the server receives terminal data uploaded by multiple acquisition terminals, it records the receiving time of each terminal data respectively; Maintaining the one-way network delay of each of the acquisition terminals; Calculate the global timestamp of each terminal data based on the one-way network delay; Performing a global timestamp for each terminal data according to a stamping rule, wherein the stamping rule includes: When the terminal data is in video mode, a global timestamp is marked for each frame of video data; When the terminal data is in audio mode, marking the audio data with a global timestamp according to a first preset period; When the terminal data is in text mode, a global timestamp is marked for each text; When the terminal data is in picture mode, a global timestamp is marked for each picture; When the terminal data is in sensor mode, a global timestamp is added at each update.

4. The method for generating video content by fusing multiple source materials according to claim 3, characterized in that: The multimodal semantic analysis and importance scoring include: Constructing a modal data sequence, wherein the modal data sequence includes a data sequence unique identifier, a modal type, a global timestamp, a source identifier of a collection terminal on the host side, a display intent label, metadata information, material size and ratio information, and modality-specific content; Sending the modal data sequences to a multimodal semantic analysis engine for processing according to the modality type; The multimodal semantic analysis engine analyzes different modal data sequences respectively and outputs semantic feature vectors; Importance scoring is performed based on the collected terminal data, display intent labels, and semantic feature vectors.

5. The method for generating video content by fusing multiple source materials according to claim 4, characterized in that: The multi-source material fusion and layout decision-making include: A weighted attribute graph is constructed based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, wherein the weighted attribute graph includes nodes and associated edges, each of the nodes represents a terminal data unit, and the attributes of the node include a node unique identifier, a data sequence unique identifier, a modal type, a global timestamp, a host-side acquisition terminal source identifier, a display intent label, metadata information, material size and ratio information, a semantic feature vector and an importance score, and the associated edges include a semantic association edge, a time overlap edge and a conflict detection edge; The materials are prioritized based on the semantically associated edges, and each node is clustered and grouped based on the semantically associated edge weights of the nodes using a clustering algorithm, and the display priority order of each cluster grouping set is determined based on the importance score corresponding to each node; Perform layout allocation based on the time overlapping edges, and allocate optimal display timing and layout area to the nodes according to the result of material priority sorting and the time overlapping edge weights between nodes; Performing local conflict detection based on the conflict detection edge weights and the material size and ratio information; Based on the weights of the nodes and associated edges of the weighted attribute graph, a fusion optimization objective function is constructed, and based on the fusion optimization objective function, a display decision plan that meets the display priority, layout timing, and conflict avoidance requirements is automatically generated; The display decision plan includes: The display order of multi-source materials is sorted according to the importance score and semantic clustering results in the fusion objective function; The spatial layout scheme of multi-source materials is based on the search and allocation of spatial areas based on the greedy algorithm on the basis of non-conflict in time; The interactive strategy for multi-source materials formulates interactive display or replacement rules based on the conflict edge penalty value and material size and ratio information.

6. The method for generating video content by fusing multiple source materials according to claim 5, characterized in that: The construction of the association edge includes: Calculating the weight of the semantic association edge and constructing the semantic association edge; Calculating the time overlapping edge weights and constructing the time overlapping edges; The conflict detection edge weight is calculated, and the conflict detection edge is constructed.

7. The method for generating video content by fusing multiple source materials according to claim 5, characterized in that: The audience terminal adaptation and layout rendering include: Acquire the adaptation information of the viewer terminal, the adaptation information including device characteristics, display capabilities and interaction methods, network environment and system preference settings; Based on the adaptation information of each of the terminals combined with the display decision plan, dynamically adjust the spatial layout, rendering method and interaction strategy of the multi-source materials; Perform rendering and interaction optimization according to the adjusted plan.

8. A video content generation system integrating multiple source materials, characterized in that: include: The multi-source material collection and preprocessing module collects multiple terminal data from the anchor side in real time, performs modality classification, global timestamp annotation, format standardization and structural unification for each terminal data, and annotates the terminal data with display intent labels based on the display intent annotation strategy. The modality types of the terminal data include video modality, audio modality, text modality, image modality and sensor modality. A multimodal semantic analysis and importance scoring module, which inputs the pre-processed terminal data into a multimodal semantic analysis engine according to the modality type, extracts semantic features, and performs importance scoring; The multi-source material fusion and layout decision module constructs a weighted attribute graph based on the pre-processed terminal data, modal data sequence, semantic feature vector and importance score, and automatically generates a display decision plan based on the weighted attribute graph using a fusion strategy; The audience terminal adaptation and layout rendering module dynamically generates a visual layout and completes the display rendering of multi-source materials based on the display decision plan and in combination with the screen attributes and interactive capabilities of the audience terminal.

9. An electronic device, characterized in that: include: processor; as well as, A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Video post-editing and video synthesis optimization method

    CN116847123A

  • Controllable video generation method and system based on multi-modal fusion

    CN119091362A

  • Intelligent decision-making system for multi-modal data fusion

    CN119337313A

  • Text-to-video conversion method and device based on deep semantic analysis, equipment and medium

    CN119399330A

  • Intelligent visualization and text association method for multi-modal knowledge graph

    CN119441281A

Cited By

  • Voice intention recognition method, electronic equipment and storage medium

    CN120279912A

  • Method for making video special effect cover based on artificial intelligence

    CN120512573A

  • Automatic video generation system and method based on AI Agent multi-mode cooperative control

    CN121126084A